Artificial nucleic acid for regulating expression of target gene, and system for regulating expression of multiple genes by using same
The SMART component addresses the challenge of controlling gene expression ratios in operons by using translational synchronization, stabilizing gene expression and improving metabolic pathway efficiency in microbial cell factories.
Patent Information
- Application Number
- PCT/KR2025/008926
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2025-06-25
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies face challenges in precisely controlling the expression of multiple genes in operons due to limitations in genetic components, leading to unstable gene maintenance, mutations, and difficulties in achieving the desired expression ratios, especially when using non-model strains for industrial applications.
Development of an artificial nucleic acid, referred to as the SMART component, which utilizes translational synchronization sequences to link two target genes, ensuring that the translation of a downstream gene occurs only when the upstream gene is translated, thereby stabilizing the expression ratio between operon constituent genes.
The SMART component effectively controls the expression ratio between genes, minimizing homologous recombination and maintaining consistent protein production, even when different host strains or promoters are used, enhancing the production of high-value metabolites and proteins in microbial cell factories.
Smart Images

Figure KR2025008926_02012026_PF_FP_ABST
Abstract
Description
Artificial nucleic acid for regulating expression of target gene and system for regulating expression of multiple genes using the same
[0001] The present invention relates to an artificial nucleic acid (genetic component) capable of precisely controlling the expression ratio between constituent genes in an operon artificially constructed in bacteria, and more specifically, to an artificial translation control system that uses different overlapping gene pair sequences that cause a translation coupling mechanism as the artificial nucleic acid of the present invention to control operon expression.
[0002] The present invention was developed with the support of the Ministry of Science and ICT's Bio-Medical Technology Development Project, "Development of an Ultra-Fast-Growing Microorganism Platform Chassis for Synthetic Biology-Based Cell Factory for Bio-Manufacturing Innovation" (Project Unique Number: 1711200058, Project Number: 2021M3A9I4024737), the Ministry of Oceans and Fisheries' Marine Bio-Industry Materials Domestication Technology Development Project, "Development of Eco-Friendly Marine Bio-Plastic Materials to Replace Petrochemical Materials" (Project Unique Number: 1525014268, Project Number: 20220258), the Ministry of Science and ICT's Synthetic Biology Core Technology Development Project, "Development of Synthetic Biology Source Technology for Advanced Non-Pathogenic Vibrio Cell Factory and Advanced Chassis Production" (Project Number: RS-2024-00352569), and the Ministry of Science and ICT's Excellent New Researcher Project, "Development of Synthetic Biology-Based Microbial Biotherapeutics" This work was supported by the Research Project (Project No. RS-2024-00345885).
[0003] Microbial cell factories are a technology that can replace petrochemical-based manufacturing processes by utilizing microorganisms with rapid growth and metabolism to mass-produce high-value metabolites and proteins. However, practical applications of cell factories are limited due to issues such as the unstable maintenance of foreign genes introduced into microorganisms, limited substrate availability, and difficulties in controlling complex intracellular metabolic pathways. To address these issues, ongoing efforts are underway to develop advanced synthetic biology tools capable of precisely and predictably controlling biological responses, including gene expression.
[0004] Microorganisms known as industrial strains, such as Escherichia coli and yeast, lack or have incomplete metabolic pathways for producing specific target products, making the introduction of foreign genes essential. If the host strain's metabolic pathway for the target product is incomplete, or if the goal is to produce non-natural biochemicals, two or more foreign genes, such as biosynthetic gene clusters (BGCs), must be introduced to establish a new pathway within the strain. Furthermore, even when non-model strains, which possess characteristics such as rapid growth rates, acid tolerance, and overproduction of intermediates compared to industrial strains, are used as host strains, appropriate expression control of each gene is necessary for mass production of the target product. Regulating the expression of each gene requires genetic components such as promoters, 5' untranslated regions (5' UTRs), and terminators. However, the number of available genetic components may be limited in some cases. Furthermore, the duplication of genetic components with identical sequences can lead to mutations or loss of the introduced foreign genes through homologous recombination mechanisms. To address these issues, many studies have utilized operons, one of the transcriptional units of microorganisms, to introduce multiple foreign genes into host strains.
[0005] Operons can express multiple genes using a single promoter, making them relatively free from limitations on the number of genetic components. Furthermore, because translation of multiple genes is initiated from a single mRNA, they offer the advantage of efficiently utilizing cellular resources required for gene expression, such as nucleotides and polymerases. However, when multiple genes are expressed in an operon format, regulating the expression of individual genes is difficult. To address this, previous studies have explored methods such as altering the order of operon genes, post-transcriptional modification of mRNA, and modeling-based operon gene expression control (operon calculators). However, existing technologies still face challenges, such as discrepancies in the expression of individual genes or limitations in the application of genetic components, which limit the scope of expression control. Therefore, the development of new synthetic biology tools for operon expression control continues.
[0006] Translational synchronization, a translation mechanism within operons, allows translation of an upstream gene to influence the translation of a downstream gene. If the stop codon of an upstream gene overlaps with the start codon of a downstream gene (e.g., TAATG, ATGA) or the distance between the two genes is very close, the ribosome that translated the upstream gene can also participate in the translation of the downstream gene. While some previous studies have attempted to develop genetic components utilizing translational synchronization, these studies utilized the sequences that induce translational synchronization as leader peptides to precisely control the expression of a single target gene, and did not utilize them to regulate the expression of multiple genes. Furthermore, depending on the secondary structure created by the sequences that induce translational synchronization, new ribosomes that have not yet translated the upstream gene can independently initiate translation of the downstream gene. However, previous studies did not screen for sequences that suppress this independent translation initiation. This approach is unsuitable for precise gene expression control, such as for cytotoxic genes, as it can lead to expression leakage. Therefore, the development of genetic components capable of simultaneously regulating the expression of multiple genes while precisely achieving the predicted expression ratio for each gene is necessary. In the present invention, we identified overlapping gene pair sequences capable of inducing translational synchronization without independently initiating translation from downstream genes, and developed these genetic components.
[0007] The present invention aims to develop an artificial translation control system based on translational coupling that stably controls operon expression and maintains the ratio between produced proteins at a constant, unique value.
[0008] To achieve the above object, the present invention provides an artificial nucleic acid that is inserted between two or more target genes to regulate gene expression. The target genes of the present invention can be divided into a first target gene located upstream of the artificial nucleic acid and a second target gene located downstream. The artificial nucleic acid of the present invention operably links the 3' end of the first target gene and the 5' end of the second target gene. In the present invention, the artificial nucleic acid can link each target gene by removing the 3' end stop codon of the first target gene and the 5' end start codon of the second target gene.
[0009] In this specification, the artificial nucleic acid for regulating the expression of a target gene may be referred to as a 'gene component' or a 'SMART component'.
[0010] The artificial nucleic acid of the present invention comprises ATGA, GTGA, or TTGA, in which the initiation codon and the termination codon overlap, and comprises a Shine-Dalgarno sequence upstream of the overlapping sequence. The artificial nucleic acid of the present invention has a sequence of 90 to 110 bp (preferably 99 bp) upstream of the overlapping sequence, and a sequence of 10 to 40 bp (preferably 30 bp) downstream of the overlapping sequence. Accordingly, the artificial nucleic acid of the present invention has a length of about 100 to 150 bp, preferably 120 to 130 bp, and most preferably 125 bp.
[0011] In a preferred embodiment of the present invention, the artificial nucleic acid comprises a sequence of any one of SEQ ID NOs: 1 to 10 or a sequence having at least 90% homology thereto.
[0012] The present invention also provides a gene expression control system utilizing the artificial nucleic acid. The gene expression control system of the present invention comprises an artificial nucleic acid inserted between a first target gene and a second target gene, thereby connecting the first target gene and the second target gene. The artificial nucleic acid of the present invention is operably linked to each target gene via a linker.
[0013] In the present invention, the artificial nucleic acid may be linked to an upstream target gene via a linker. In this case, the linker gene comprises a gene sequence encoding the sequence of SEQ ID NO: 18.
[0014] In the present invention, the artificial nucleic acid can be connected to a downstream target gene through a linker, and the linker gene connected to the downstream target gene is composed of a gene sequence encoding 'amino acid AG'.
[0015] The present invention also provides a method for controlling the expression level of a target gene using the gene expression control system.
[0016] The artificial nucleic acid and system of the present invention utilize a translation synchronization mechanism to ensure that a downstream gene is translated only when an upstream gene is translated, and to ensure that ribosomes involved in the translation of the upstream gene are used in the translation of the downstream gene, thereby accurately controlling the expression ratio between operon constituent genes as predicted.
[0017] The artificial nucleic acid of the present invention is composed of different sequences extracted from the genome of E. coli MG1655, so that even when each genetic component is used together to construct an artificial operon, the phenomenon of homologous recombination that may occur within the host strain can be minimized.
[0018] The artificial nucleic acid of the present invention can control the unique gene expression ratio exhibited by the translation synchronization sequence by changing the Shine-Dalgarno sequence present in each translation synchronization sequence.
[0019] Figure 1 is a schematic diagram illustrating the gene expression control system of the present invention. In Figure 1, mcherry is used as the upstream target gene, and sgfp is used as the downstream target gene. Expression control nucleic acids (SMART Biopart in Figure 1) are inserted between each target gene to control the expression level of the target gene.
[0020] Figure 2 is an image showing the structure and mechanism of action of the gene expression control system of the present invention.
[0021] Figure 3 is a flow chart showing a screening process performed to determine a gene expression regulating nucleic acid of the present invention.
[0022] FIG. 4 is an image showing the locations of 17 candidate gene expression regulatory nucleic acids (SMART) selected from the E. coli K-12MG1655 genome according to one embodiment of the present invention.
[0023] Figure 5 shows the TAA linked to the 5' end of the candidate nucleic acid (TAA + This graph shows the ratio of expressed mcherry and sgfp proteins using SMART cassette production. The expression levels of mcherry and sgfp proteins were both standardized to the amount of mcherry protein in each candidate group.
[0024] Figure 6 shows the results of comparing the amount of target protein produced by inserting a nucleic acid (SMART cassette) synthesized according to the present invention. In Figure 6, the SMART activation ratio refers to the expression ratio between two adjacent genes, mcherry and sgfp. Furthermore, the y-axis in Figures 5 and 6 represents the relative amount of each protein, and each data point is expressed as the mean ± standard deviation (SD) of three biologically independent sample measurements, with white dots representing actual data points.
[0025] Figure 7 shows the TAA binding (TAA) of nucleic acids selected according to the present invention. +The results are a comparison of the transcription rate of the target gene and the SMART activation rate according to the cassette. SMART and TAA + When comparing transcripts derived from cassettes, no significant relationship was observed between TC and transcription rate.
[0026] Figure 8 shows the results of comparing SMART activation ratios when T7 and Tac were used as promoters in an operon into which the nucleic acid (SMART component) of the present invention was introduced. Referring to Figure 8, it can be confirmed that the SMART activation ratio is maintained within the margin of error depending on the type of promoter.
[0027] Figure 9 shows the results of confirming the effect of changes in the expression level of an upstream gene in an operon into which the nucleic acid (SMART component) of the present invention was introduced on changes in the expression level of a downstream gene. In this example, the translation initiation rate of the mCherry gene was controlled by changing the 5' untranslated region (5' UTR) sequence of mCherry. In Figure 9, the x-axis and y-axis represent the amounts of mCherry and sGFP, respectively, normalized to cell biomass (OD600).
[0028] The left side of Figure 10 shows the results confirming the influence of the coding sequence of the target gene on the SMART activation rate. Various synthetic cassettes with various reporter combinations were used in this example, and the mcherry, sgfp, and luc genes were representatively used as reporters. The right side of Figure 10 shows the results confirming the influence of the linker sequence. Three linker sequences were introduced upstream of the nucleic acid of the present invention and the results were verified. The results of Figure 10 indicate that the differences between the linkers are negligible. Data are expressed as the mean ± standard deviation (sd) of measurements from three biologically independent samples (n = 3), and each dot represents an actual data point.
[0029] Figure 11 shows the results of examining whether the SMART activation ratio changed according to the change in the host strain. In each case of Figure 11, the SMART cassette was transcribed by the Tac promoter. In Figure 11, ECD represents BL21 (DE3); ECB represents BL21; ECN represents Nissle 1917; ECM represents K-12 MG1655; and ECW represents the W strain. Each data point was expressed as the mean ± standard deviation (sd) of three biologically independent samples (n = 3) measured.
[0030] Figure 12 shows the results of confirming whether the SMART activation ratio by each nucleic acid can be changed by changing the Shine-Dalgarno sequence (SD sequence) present in the nucleic acid (SMART component) of the present invention.
[0031] Figure 13 is an image showing a simplified biosynthetic pathway leading to the production of 3-HP, P(3HB), and lycopene. In Figure 13, dhaB123 is a gene encoding three protein subunits of glycerol dehydratase (GDHt); kgsadh is a gene encoding an improved ALDH enzyme mutant; gdrAB is a gene encoding GDHt revertase; phaA is a gene encoding 3-ketothiolase; phaB is a gene encoding acetoacetyl coenzyme A reductase; phaC is a polyhydroxyalkanoate (PHA) synthase; crtE is a gene encoding geranylgeranyl pyrophosphate synthase; crtB is a gene encoding phytoene synthase; crtI is a gene encoding phytoene desaturase.
[0032] FIG. 14 shows the results of confirming the 3-HP production using a synthetic operon composed of kgsadh, dhaB1, and eight different nucleic acid sequences according to one embodiment of the present invention.
[0033] Figure 15 shows a schematic diagram of P(3HB) production using a synthetic operon composed of phaABC and a combination of three different nucleic acids (SMART components) and the resulting results.
[0034] Figure 16 shows the results of lycopene production using synthetic operons composed of various combinations of crtEBI and five SMART sequences. The SMART activation ratios used for each operon in Figures 14, 15, and 16 are indicated by shaded circles or semicircles below the bar graphs. Each data point is expressed as the mean ± standard deviation (sd) of measurements from three biologically independent samples (n = 3), with each dot representing an actual data point.
[0035] Below, various examples are presented to aid understanding of the invention. These examples are provided solely to facilitate understanding of the invention and are not intended to limit the scope of protection of the invention.
[0036] The symbols A, T, C, G, and U used herein are to be interpreted as meanings understood by a person skilled in the art. They may be appropriately interpreted as bases, nucleosides, or nucleotides on DNA or RNA depending on the context and technology. For example, when referring to a base, they may be interpreted as adenine (A), thymine (T), cytosine (C), guanine (G), or uracil (U) themselves, respectively, and when referring to a nucleoside, they may be interpreted as adenosine (A), thymidine (T), cytidine (C), guanosine (G), or uridine (U), respectively, and when referring to a nucleotide in a sequence, they should be interpreted to mean a nucleotide including each of the above nucleosides.
[0037] The term "operably linked" as used herein means, in gene expression technology, that a specific structure is linked to another structure so that the specific structure can function in the intended manner. For example, when a promoter sequence is said to be operably linked to a coding sequence, it means that the promoter is linked so as to affect the transcription and / or expression of the coding sequence in a cell. In addition, the term includes all meanings that can be recognized by a person skilled in the art, and can be appropriately interpreted according to the context.
[0038] As used herein, the term "target gene" or "target nucleic acid" basically refers to a cellular gene or nucleic acid that is the target of gene expression regulation. The term "target gene" or "target nucleic acid" may be used interchangeably and may refer to the same subject. Unless otherwise specified, the term "target gene" or "target nucleic acid" may refer to either a gene or nucleic acid inherent to the target cell or an exogenous gene or nucleic acid, and is not particularly limited as long as it can be the target of gene expression regulation. The term "target gene" or "target nucleic acid" may be single-stranded DNA, double-stranded DNA, and / or RNA. Furthermore, the term encompasses all meanings that can be recognized by those skilled in the art and may be appropriately interpreted depending on the context.
[0039] The target genes of the present invention may be located upstream and downstream of the "artificial nucleic acid (gene component or SMART component of the present invention)", respectively. For example, the first target gene may be located upstream of the artificial nucleic acid, and the second target gene may be located downstream of the artificial nucleic acid. The first target gene and the second target gene may be composed of different sequences. The target genes of the present invention are not limited in type, and any gene sequence capable of expression through the translation mechanism of an operon may be included.
[0040] The "artificial nucleic acid" of the present invention can be inserted between different target genes (e.g., a first target gene and a second target gene) to control the expression ratio of each target gene. The "artificial nucleic acid" of the present invention includes a Shine-Dalgarno sequence within a translational synchronization sequence, and the expression level of the target gene can be controlled by changing the Shine-Dalgarno sequence. When the binding energy of mRNA and rRNA is strengthened by changing the Shine-Dalgarno sequence, the expression ratio of the downstream (second) target gene can be increased.
[0041] For example, the expression ratio of the downstream gene can be increased in the order of SEQ ID NO: 1 (cpxRA), SEQ ID NO: 2 (pcnB / folK), SEQ ID NO: 3 (fecCD), SEQ ID NO: 4 (pqiBC), SEQ ID NO: 5 (nanEK), SEQ ID NO: 6 (ybgPO), SEQ ID NO: 7 (hyfHI), SEQ ID NO: 8 (hisDC), SEQ ID NO: 9 (astDB), SEQ ID NO: 10 (fumE / yggC), which are artificial nucleic acids of the present invention. That is, when the artificial nucleic acid of SEQ ID NO: 1 is used, the expression ratio of the upstream gene can be increased compared to the artificial nucleic acid of SEQ ID NO: 10.
[0042] Referring to Table 1 below, the expression ratios of upstream genes (e.g., first target gene) and downstream genes (e.g., second target gene) that appear when the artificial nucleic acid of the present invention is used can be determined. In Table 1, the SMART activation ratio represents the expression ratio of the downstream gene to the upstream gene (downstream gene expression amount / upstream gene expression amount). When using cpxRA as an artificial nucleic acid, the expression level can be controlled at a ratio of 100:8~10 (upstream gene expression level: downstream gene expression level, hereinafter the same), 100:12~14 when using pcnB / folK, 100:16~20 when using fecCD, 100:26~29 when using pqiBC, 100:32~37 when using nanEK, 100:40~49 when using ybgPO, 100:45~51 when using hyfH, 100:66~70 when using hisDC, 100:81~89 when using astDB, and 100:84~99 when using fumE / yggC.
[0043] Sequence name SMART activation ratio ratio range min max cpxRA 0.09 0.08 0.1 pcnB / folK 0.120.120.14 fecCD 0.180.160.2 pqiBC 0.270.260.29 nanEK 0.340.320.37 ybgPO 0.440.400.49 hyfHI 0.470.450.51 hisDC 0.680.660.7 astDB 0.840.810.89 fumE / yggC 0.920.840.99 yidKJ 3.173.003.40 livHM 12.1811.6012.60
[0044] The following examples are presented to help understand the invention and do not limit the scope of protection of the invention.
[0045]
[0046] [Example 1]
[0047] Screening of operon expression control modules based on translational synchronization sequences using fluorescent proteins
[0048] To verify whether operon gene expression is regulated by translational synchronization, an artificial translation cassette was constructed using a reporter protein. The artificial translation cassette contains two genes encoding the fluorescent proteins mCherry and sGFP, which are sequentially located, and a T7 promoter and T7 terminator were introduced to enable transcription by T7 RNA polymerase. The 3' end of the mCherry coding sequence was replaced with a stop codon to insert a linker sequence encoding 10 amino acids (DSAGSAGSAG) to prevent the translational synchronization sequence from affecting the folding of the upstream gene. The 5' end of the sGFP coding sequence was replaced with a start codon to insert a linker sequence encoding 2 amino acids (AG). By introducing a Synthetic Module for Accurate Expression Ratio by Translational Coupling (SMART component) based on a translational synchronization sequence between the mCherry and sGFP coding sequences, translation of the downstream gene can be initiated by the ribosome translating the upstream gene (Fig. 2).
[0049] First, to secure sequences that induce translational synchronization, we searched for gene pairs that overlapped in the same direction in the genome of E. coli K-12 MG1655, and extracted only specific sequences through a series of processes (Fig. 3). First, among the 660 total overlapping gene pairs existing in E. coli, we extracted 261 most frequent 4-bp overlapping gene pairs. A 4-bp overlapping gene pair is a form in which the stop codon of the upstream gene overlaps with the start codon of the downstream gene, and in the present invention, these correspond to ATGA, GTGA, and TTGA. From the list of 4-bp overlapping gene pairs, 17 gene pairs distributed at various locations on the E. coli genome were randomly selected (Fig. 4). Afterwards, only the 99 bp sequences from the 3' end of the upstream gene and the 30 bp sequences from the 5' end of the downstream gene were extracted to secure a group of translational synchronization sequence candidates. In this case, the length of the extracted translational synchronization sequence was 125 bp due to the 4 bp overlapping sequence of the upstream and downstream genes (see Table 1).
[0050]
[0051] In Table 2 above, the bold text indicates genes with overlapping start and stop codons (ATGA, GTGA), and the underlined sequence is the Shine-Dalgarno predicted sequence.
[0052]
[0053] In order to ensure that operon translation occurs solely through the translation synchronization mechanism, we excluded gene pairs in which translation of downstream genes was initiated when the translation synchronization mechanism did not occur in the selected gene pairs. To achieve this, an artificial stop codon, TAA, was added to the 5' end of the translation synchronization sequence, preventing the ribosome translating the upstream gene from reaching the 4 bp overlapping sequence, thereby preventing translation synchronization. A group of translation synchronization sequence candidates linked to the artificial stop codon was inserted between the mCherry and sGFP coding sequences of the artificial translation cassette constructed above (TAA + SMART cassette). For translationally synchronized sequences in which the amount of sGFP protein calculated from the fluorescence measurement results exceeded 2% of the amount of mCherry protein, translation initiation of the downstream gene was judged to occur independently (see Table 3; sequences with an sGFP / mCherry ratio > 0.02 were eliminated). By excluding sequences that caused independent translation initiation, only 12 translationally synchronized sequences were selected (Fig. 5).
[0054] Selection of nucleic acids that cause translational affinity (TAA) +Using SMART cassette)mCherry (pmol / OD600)sGFP (pmol / OD600)Nametriplicate 1triplicate 2triplicate 3triplicate 1triplicate 2triplicate 3sGFP / mCherry ratiocpxRA91.284.783.80.00.00.00.00pcnB / folK79.173.974.50.00.00.00. 00fecCD80.377.475.30.00.00.00.00pqiBC82.083.079.21.81.91.90.02nanEK 88.683.581.40.00.00.00.00ybgPO81.279.380.50.00.00.00.00hyfHI83.379. 579.60.60.60.50.01hisDC64.772.569.90.40.50.50.01astDB101.9100.396.90 .60.50.50.01fumE / yggC80.879.174.91.11.10.90.01yidKJ67.961.362.31.00 .90.90.01livHM77.378.881.51.01.21.30.01lpxB / rnhB70.273.271.02.72.92. 70.04ydbH / ynbE78.574.676.58.98.28.80.11sgcXB57.856.158.08.38.38.60. 15napHB60.163.857.914.414.713.30.23lldRD69.871.466.437.536.534.90.52
[0055]
[0056] In Table 3 above, lpxB / rnhB, ydbH / ynbE, sgcXB, napHB, and lldRD, which had sGFP / mCherry ratios > 0.02, were eliminated from the candidate group.
[0057]
[0058] The selected translational synchronization sequences were inserted into a translation cassette without an artificial stop codon, and fluorescence was measured, and 10 SMART components were selected whose expression ratio of the downstream gene to the upstream gene (sGFP protein / mCherry protein, hereinafter referred to as SMART activation ratio) did not exceed 1 (Fig. 6). This selection criterion was established based on the assumption that if translation of the downstream gene is initiated solely by the ribosomes used to translate the upstream gene without independent translation initiation of the downstream gene, the number of ribosomes used to translate the downstream gene cannot exceed the number of ribosomes used to translate the upstream gene. The 10 SMART components finally selected showed different SMART activation ratios.
[0059] Selection of nucleic acids using the expression levels of upstream target genes / downstream target genes mCherry (pmol / OD600) sGFP (pmol / OD600) Name triplicate 1 triplicate 2 triplicate 3 triplicate 1 triplicate 2 triplicate 3 SMART activation ratio (sGFP / mCherry)cpxRA37.636.533.63.43.23.10.09pcnB / folK52.551.350.66.46.76.10.12fecCD16.115.718.72.83. 13.20.18pqiBC54.248.547.714.413.713.20.27nanEK24.924.021.58.07.87.90.34ybgPO22.323.625.310.810.310.2 0.44hyfHI43.043.542.119.620.021.20.47hisDC30.027.327.620.918.218.90.68astDB38.839.940.434.233.332.90 .84fumE / yggC15.014.513.112.713.512.90.92yidKJ8.38.58.927.826.327.33.17livHM3.23.23.439.639.839.912.18
[0060]
[0061] Fluorescence measurement experiments to confirm whether the selected sequences exhibited a translational synchronization mechanism were conducted as follows. After introducing the artificial translation cassette into the E. coli BL21 (DE3) strain, the strain was cultured overnight in LB medium supplemented with 34 μg / mL chloroamphenicol to maintain the plasmid containing the artificial translation cassette. The strain was then diluted 1 / 100, inoculated into 5 mL of fresh medium, and cultured in a shaking incubator at 37°C and 250 rpm. When the absorbance (OD600) was around 0.8-1.0, 0.2 mM IPTG was added to induce expression of the artificial translation cassette. After the addition of IPTG, the culture was incubated for an additional 6 hours, and the final absorbance was measured. The fluorescence from each fluorescent protein (mCherry, sGFP) and the luminescence from nanoluciferase were measured using a Hidex Sense Microplate Reader. The wavelength filters used for fluorescence measurements in the present invention are as follows. mCherry, 575-nm excitation filter, 616-nm emission filter; sGFP, 485-nm excitation filter, 535-nm emission filter. Luminescence by nanoluciferase was detected using Nano-Glo from Promega. ®Luciferase Assay System was used to measure the amount of fluorescent protein according to the user manual using Hidex Sense Microplate Reader. The amount of fluorescent protein was estimated from each fluorescence to calculate the gene expression ratio by SMART components. To this end, each protein was purified using Ni-NTA method by linking a 6X His-tag to the 3' end of the fluorescent protein, and the amount of purified protein was measured using Bradford assay. The fluorescence of the quantified fluorescent protein was measured while diluting it to different amounts, and an equation for the relationship between the fluorescence value and the protein amount was obtained. Considering the molecular weight of each fluorescent protein, the mole number was calculated from the fluorescence value. Similarly, the mole number of protein was calculated from the luminescence value of nanoluciferase. All subsequent fluorescence measurement experiments were conducted according to this method.
[0062]
[0063] [Example 2]
[0064] SMART component characteristic analysis
[0065] Droplet digital PCR was performed to determine the effect of SMART components on the transcription amount of operon constituent genes. When an artificial stop codon (TAA) was added to the 5' end of the SMART component + No particular trend was observed between the transcript amounts of (SMART) and (non-SMART) cases, and no trend was observed between the SMART activation ratio and the transcript amount (Fig. 7). Therefore, it was confirmed that the nucleic acid (SMART component) selected according to the present invention does not affect the transcript amounts of operon constituent genes.
[0066] Conversely, to determine whether changes in transcript abundance affect the SMART activation ratio, we replaced the T7 promoter with the Tac promoter in a conventional artificial translation cassette and performed fluorescence measurements. Even when different promoters were used, the SMART activation ratio remained largely intact (Fig. 8).
[0067] We investigated the effect of changes in the expression level of upstream genes on changes in the expression level of downstream genes in an operon introducing SMART components. To this end, the translation initiation rate of the mCherry gene was controlled by changing the 5' untranslated region (5' UTR) sequence of mCherry. Fluorescence measurements were performed on SMART components pcnB / folK, hyfHI, hisDC, and astDB. The results confirmed that the amount of sGFP protein changed in proportion to the change in the amount of mCherry protein, maintaining the SMART activation ratio (Fig. 9).
[0068] In the present invention, we also examined the effects of altering the gene sequence and linker constituting the artificial operon on the SMART activation ratio. In addition to the genes encoding mCherry and sGFP, nanoluciferase was additionally utilized as a reporter. The linker was altered by changing only the nucleotide sequence while maintaining the 10 amino acid sequence present at the 3' end of the mCherry coding sequence in the existing artificial translation cassette (using the linker sequence in Table 5). It was confirmed that the SMART activation ratio of each SMART component was well maintained even when the gene sequence and linker were altered (Fig. 10).
[0069] Linker Sequence NameSequence (5' -> 3')Sequence NumberLinker (amino acid)DSAGSAGSAG18LinkerGAT AGT GCT GGT AGT GCT GGT AGT GCT GGT19Linker_SGAT TCG GCC GGA AGC GCT G6G TCC GCA GGC20Linker_SSGAT AGC GCT GGG TCC GCA GGA AGT GCT GGC21
[0070]
[0071] In the present invention, we also confirmed whether the SMART activation ratio changed when the host strain was changed. An artificial translation cassette based on the Tac promoter introducing 10 SMART components was introduced into 4 other E. coli strains (E. coli K-12 MG1655, W, BL21, Nissle 1917) in addition to E. coli BL21 (DE3). As a result of comparing the SMART activation ratio calculated from the amount of fluorescent protein measured in each strain, E. coli BL21 strain showed the highest correlation (R2 = 0.94) with the existing host strain, and the other three strains also showed high correlations (R2 ≥ 0.85), confirming the applicability of SMART components to other strains (Fig. 11).
[0072]
[0073] [Example 3]
[0074] Expanding the component types by changing the Shine-Dalgarno sequence within SMART components.
[0075] We verified whether the SMART activation ratio by each SMART component could be changed by changing the Shine-Dalgarno sequence (SD sequence) present in the SMART component. We randomly selected bases that did not significantly change the mRNA secondary structure among the SD sequences of SMART components extracted from pcnB / folK, hyfHI, hisDC, and astDB and changed them, and calculated the binding energy (ΔGmRNA_rRNA) of mRNA and rRNA in the changed SD sequence. The RNAfold web server (http: / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) was used to predict the mRNA secondary structure of the SMART component part, and the RBS calculator (https: / salislab.net / software / design_rbs_calculator) was used to calculate the binding energy. When using a SMART component with a modified SD sequence, it was confirmed that the expression ratio of downstream and upstream genes (sGFP / mCherry protein ratio) increased as the binding energy of mRNA and rRNA was strengthened (Fig. 12).
[0076]
[0077] [Example 4]
[0078] Analysis of Productivity Changes When Applying SMART Components to Metabolite Production Pathways
[0079] According to the present invention, we investigated whether the metabolic flow involved in the production of biochemicals could be controlled by introducing selected nucleic acids (SMART components) into metabolic pathways. Through this, we aimed to confirm the possibility of optimizing metabolite production without the order of the constructed artificial operon and additional strain engineering. High-value metabolites such as 3-hydroxypropionic acid (3-HP), P(3HB) (poly(3-hydroxybutyrate), a type of polyhydroxybutyrate), and lycopene were selected as target metabolites, and their metabolic pathways were constructed in E. coli MG1655 or E. coli W strains (Fig. 13).
[0080] The 3-HP metabolic pathway consists of glycerol dehydratase (GDHt), which produces 3-hydroxypropionaldehyde (3-HPA) from glycerol, aldehyde dehydrogenase (KGSADH), which produces 3-HP from 3-HPA, and glycerol dehydratase reactivase (GDHt reactivase), which reactivates GDHt. GDHt is a complex composed of three protein units (DhaB1, DhaB2, and DhaB3), of which DhaB1 protein is known to be the main unit that causes the enzymatic reaction. In the present invention, the amount of 3-HP produced from an artificial operon constructed by inserting eight different SMART components between the genes encoding KGSADH and DhaB1 was compared. Five SMART components derived from Escherichia coli (pcnB / folK, pqiBC, hyfHI, hisDC, astDB) and three SMART components constructed by modifying the SD sequence (hisSD1, astSD3, hyfSD14) were used. The metabolic pathway for 3-HP production was constructed by modifying 3-HP production plasmids 1 and 2 used in two previous studies. 3-HP production plasmid 1 expresses DhaB1 as an independent gene cassette, and DhaB2, DhaB3, and glycerol dehydratase reactivator as an operon (Lim et al., ACS Synth. Biol., 2016). 3-HP production plasmid 2 expresses KGSADH (Seok et al., Metab. Eng., 2018). In the present invention, the DhaB1 gene cassette present in 3-HP production plasmid 1 was removed, and a 10-amino acid linker-SMART component-2-amino acid linker-DhaB1 gene was placed downstream of the KGSADH gene of 3-HP production plasmid 2.After introducing the newly reconstructed 3-HP production plasmids 1 and 2 into the E. coli W strain and conducting a 3-HP production experiment, it was confirmed that a difference of up to 7.7 times in 3-HP concentration occurred, and that 3-HP production decreased as the expression level of the downstream gene increased compared to the upstream gene (Fig. 14).
[0081] For 3-HP production, minimal medium [0.5 g / L MgSO4·7H2O, 2 g / L NH4Cl, 2 g / L NaCl, 1 g / L yeast extract, 100 mM potassium phosphate buffer (pH 7.0)] with 100 mM glycerol as a carbon source was used. To maintain the plasmid, 34 μg / mL chloramphenicol and 80 μg / mL ampicillin were added. Cells cultured overnight were subcultured into 5 mL medium when the initial absorbance reached 0.05 and cultured in a shaking incubator at 37°C and 250 rpm. When the absorbance reached 0.9–1.0, 0.1 mM IPTG and 0.2 μM B12 were added. After 8 and 16 hours of incubation, pH was measured, and if the pH was below 7.0, 10 M NaOH was added to correct the pH to 7.0. After 24 hours of incubation, the culture solution was analyzed using HPLC (high-performance liquid chromatography). The analysis equipment used was a Shimadzu LC-40 HPLC system, and the analysis conditions were an Aminex HPX-87H column and 5 mM H2SO4 flowing at 0.6 mL / min as the mobile phase.
[0082] P(3HB) can be produced from acetyl-CoA in E. coli, and the metabolic pathway additionally required to produce P(3HB) consists of three enzymes (PhaA, PhaB, PhaC). An artificial operon was constructed with the PhaABC genes, and a total of nine artificial operons were constructed by combining and inserting three types of SMART parts (pcnB / folK, hyfHI, astDB) between each gene. At this time, the artificial operons were expressed by the Tac promoter. After introducing the artificial operon into the E. coli MG1655 strain, the strain was cultured overnight in LB medium supplemented with 34 μg / mL Chloramphenicol to maintain the plasmid containing the artificial operon. Afterwards, the strain was inoculated into 100 mL of fresh medium so that the initial absorbance was 0.02 and cultured in a shaking incubator at 30°C and 200 rpm. When the absorbance (OD600) was 1.0-1.2, 0.2 mM IPTG was added to induce expression of the artificial operon. P(3HB) was extracted and quantified, and the production yield (the weight of P(3HB) produced per dry weight of cells) was calculated. It was confirmed that as the expression of PhaC increased due to changes in the SMART component, the yield of P(3HB) also increased (Fig. 15).
[0083] For P(3HB) extraction, the strain was cultured for 72 hours, then the cells were harvested by centrifugation and washed once with 1X PBS solution. After the cells were lyophilized for 48 hours, methanolysis for P(3HB) extraction was performed using approximately 50 mg of lyophilized cells. For methanolysis, 2 mL of chloroform, 1.9 mL of methanol containing 15% sulfuric acid, and 0.1 mL of methanol containing 50 mg / L benzoic acid were mixed with the lyophilized cells and reacted at 65°C for more than 48 hours. After adding 2 mL of distilled water and vortexing, the mixture was allowed to separate into layers, and only the lower layer solution was extracted and used for GC analysis. An Agilent 7890B instrument was used, and P(3HB) and benzoic acid were detected using FID. P(3HB) quantification was measured by calculating the signal area of P(3HB) relative to benzoic acid and comparing the signal area with that of the previously quantified P(3HB) standard.
[0084] The lycopene metabolic pathway is composed of three foreign enzymes (CrtE, CrtB, and CrtI) added to the DXP metabolic pathway existing in E. coli. After constructing an artificial operon with the CrtEBI gene, five SMART parts derived from E. coli (pcnB / folK, pqiBC, hyfHI, hisDC, and astDB) were randomly inserted between each gene, resulting in a total of 25 SMART part combinations. The artificial operon was expressed by the Tac promoter. After introducing the artificial operon into E. coli MG1655 strain, the strain was cultured overnight in LB medium supplemented with 34 μg / mL Chloramphenicol to maintain the plasmid containing the artificial operon. The strain was then inoculated into 3 mL of fresh medium to an initial absorbance of 0.05 and cultured in a shaking incubator at 37°C and 250 rpm. When the optical density (OD600) was 1.0-1.2, 0.2 mM IPTG was added to induce expression of the artificial operon. After 24 h of culture, the final optical density of the culture was measured, and lycopene was extracted and quantified to calculate the production yield (weight of lycopene produced per dry weight of cells). By varying the expression ratio of the operon through combinations of different SMART components, the production yield of the target product with a difference of up to 2.2-fold was controlled (Fig. 16).
[0085] To extract lycopene, 1 mL of the culture medium was washed once with 1X PBS solution, and then 200 μL of acetone was added and reacted at 55°C for 15 minutes. After centrifugation to settle cell debris, the supernatant was mixed with an equal volume of DMSO, and the lycopene content was measured from the absorbance value at a wavelength of 475 nm. A NanoDrop spectrophotometer (ThermoFisher) was used for absorbance measurement. To convert the amount of lycopene from the absorbance value, the quantified lycopene was diluted to various concentrations and the absorbance was measured, and the proportional relationship between the absorbance and the amount of lycopene was determined. Referring to Figure 16, it can be seen that the productivity of metabolites can be controlled by applying the present technology to the pathway genes that produce lycopene. Therefore, when using the present invention, multiple genes can be rapidly introduced into a host strain, while the expression of each gene can be accurately and predictably controlled.
[0086]
[0087] The present invention has been described above, focusing on preferred embodiments thereof. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.
[0088] Sequence number 1
[0089] cpxRA
[0090] ATTTCCAACCTGCGTCGTAAACTGCCGGATCGTAAAGATGGTCACCCGTGGTTTAAAACCTTGCGTGGTCGCGGCTATCTGATGGTTTCTGCTTCATGATAGGCAGCTTAACCGCGCGCATCTTC
[0091]
[0092] 서열번호 2
[0093] pcnB / folK
[0094] CAAAAAGGGATGCTCAACGAGCTGGATGAAGAACCGTCACCGCGTCGTCGTACTCGTCGTCCACGCAAACGCGCACCACGTCGTGAGGGTACCGCATGACAGTGGCGTATATTGCCATAGGCAGC
[0095]
[0096] 서열번호 3
[0097] fecCD
[0098] GCACGCGCGCTGGCCTTCCCCGGAGATCTGCCCGCAGGCGCAGTGCTGGCGCTGATTGGCAGCCCTTGCTTTGTCTGGCTTGTGAGGAGGCGAGGATGAAAATTGCGCTGGTTATTTTCATCACC
[0099]
[0100] 서열번호 4
[0101] pqiBC
[0102] CTGCAACCGGTGCTGAAAACGCTCAATGAGAAGAGTAACGCGCTGGTATTTGAAGCGAAGGACAAAAAAGATCCAGAGCCGAAGAGGGCGAAACAATGAAAAAGTGGCTAGTGACGATTGCAGCA
[0103]
[0104] 서열번호 5
[0105] nanEK
[0106] CGCCACGGCGCGTGGGCGGTGACGGTCGGTTCTGCAATCACGCGTCTTGAGCACATTTGTCAGTGGTACAACACAGCGATGAAAAAGGCGGTGCTATGACCACACTGGCGATTGATATCGGCGGT
[0107]
[0108] 서열번호 6
[0109] ybgPO
[0110] CAGTTTTATCTCGGTTACATGGATGATTACGGCGCATTACGCATGACGACGCTCAACTGCAGCGGACAATGCCGTTTACAAGCAGTGGAGGCGAAATGAGTGCTGGCAAGGGATTGTTGCTCGTC
[0111]
[0112] 서열번호 7
[0113] hyfHI
[0114] TTGTGGGCGCAAGCGAGCGTCTGCCCGGAATGCAAACAACGCGCGACGCTGATCAACGACGATACAGATGTACTGCTGGTGGCTAAGGAGCAGCTATGAGTCCAGTGCTTACACAACATGTCAGC
[0115]
[0116] 서열번호 8
[0117] hisDC
[0118] CTGGCTTCAACCATAGAAACACTGGCCGCCGCCGAGCGCCTGACCGCCCACAAAAATGCCGTTACTTTGCGTGTTAACGCCCTTAAGGAGCAAGCATGAGCACCGTGACTATTACCGATTTAGCG
[0119]
[0120] 서열번호 9
[0121] astDB
[0122] TACTGCGCATGGCCGATGGCGAGCCTGGAGTCGGACTCGTTAACATTGCCCGCCACGCTTAACCCCGGGCTGGATTTTTCCGATGAGGTGGTGCGATGAACGCCTGGGAAGTCAATTTCGACGGG
[0123]
[0124] 서열번호 10
[0125] fumE / yggC
[0126] GCTGATTTCTCCGTCTGGTCAGAGGCGCGTTTTAGCGGAATGGTCAAAACGGCGCTGACGCTGGCAGTAACGACCACCTTAAAGGAATTAACGCCGTGAAAATTGAATTAACGGTGAATGGGCTG
[0127]
[0128] 서열번호 11
[0129] yidKJ
[0130] ATTGCCGCCGTAGTGATTGTCTACCTGATTTTTGACAGCTGGCGGCATCGTCACGACCCAGCCGTAACCTTTACTCCCGACGGGAAGGATAGCCTATGAAACGCCCCAATTTTCTGTTCGTCATG
[0131]
[0132] 서열번호 12
[0133] livHM
[0134] ACGGAATATAAAGATGTGGTCTCATTCGCCCTGCTGATTCTGGTGCTGCTGGTGATGCCGACCGGTATTCTGGGTCGCCCGGAGGTAGAGAAAGTATGAAACCGATGCATATTGCAATGGCGCTG
[0135]
[0136] 서열번호 13
[0137] lpxB / rnhB
[0138] AGCCACGCGATGCACGATACCTTCCGTGAACTGCATCAGCAGATCCGCTGCAATGCCGATGAGCAGGCGGCACAAGCCGTTCTGGAGTTAGCACAATGATCGAATTTGTTTATCCGCACACGCAG
[0139]
[0140] 서열번호 14
[0141] ydbH / ynbE
[0142] TTACGCTTTGGCGATAATCTCCAGGCATGGCTGGAGCAGAACGCACGTCTGCCGGGAAATGACTGTCCGCAAGGAAAAGAGTGTGAGGAAAAACAATGAAAATTTTACTGGCTGCGTTGACGTCA
[0143]
[0144] 서열번호 15
[0145] sgcXB
[0146] GATTGCATCCGCCTATTGACCGCTCTGGCAGGTATGTCAGCAGCACATTTCCCCGTTGAGCCTGATTCAGGCACTACACAAGAGGCACATCCATTATGAAAAAGATCCTTGTGGCATGCGGTACC
[0147]
[0148] 서열번호 16
[0149] napHB
[0150] ACCAGCCGCGATTGCATGACGTGCGGTCGCTGCGTGGATGTCTGTTCTGAGGATGTATTTACAATAACTACACGATGGAGTTCGGGAGCGAAATCATGAAAAGCCATGACCTGAAGAAAGCGCTG
[0151]
[0152] 서열번호 17
[0153] lldRD
[0154] ACCACCATGAAACGATTCGATGAAGATCAGGCTCGCCACGCACGGATTACCCGCCTGCCCGGTGAGCATAATGAGCATTCGAGGGAGAAAAACGCATGATTATTTCCGCAGCCAGCGATTATCGC
[0155]
[0156] Sequence number 18
[0157] Linker
[0158] DSAGSAGSAG
[0159]
[0160] Sequence number 19
[0161] Linker
[0162] GAT AGT GCT GGT AGT GCT GGT AGT GCT GGT
[0163]
[0164] Sequence number 20
[0165] Linker_S
[0166] GAT TCG GCC GGA AGC GCT G6G TCC GCA GGC
[0167]
[0168] Sequence number 21
[0169] Linker_SS
[0170] GAT AGC GCT GGG TCC GCA GGA AGT GCT GGC
Claims
1. An artificial nucleic acid comprising any one of the sequences of SEQ ID NOs. 1 to 10 or a sequence having 90% or more homology therewith, and inserted between a first target gene and a second target gene to regulate expression of the gene.
2. In paragraph 1, The above artificial nucleic acid is an artificial nucleic acid that connects the 3' end of the first target gene and the 5' end of the second target gene.
3. In paragraph 1, An artificial nucleic acid wherein the first target gene has the stop codon at the 3' end removed, and the second target gene has the start codon at the 5' end removed.
4. In paragraph 1, The artificial nucleic acid has a length of 100 to 150 bp.
5. In paragraph 1, An artificial nucleic acid, wherein the nucleic acid comprises any one of the sequences of SEQ ID NOs: 1 to 10.
6. As a gene expression control system, An artificial nucleic acid sequence comprising any one of the sequences of SEQ ID NOs. 1 to 10 or a sequence having 90% or more homology thereto, and inserted between a first target gene and a second target gene to connect the first target gene and the second target gene, A system wherein the above artificial nucleic acid is operably linked to each target gene via a linker.
7. In paragraph 6, A system in which the above artificial nucleic acid sequence is connected to an upstream target gene by a linker consisting of a gene sequence encoding the sequence of SEQ ID NO:
18.
8. In paragraph 6, A system in which the above artificial nucleic acid sequence is connected to a downstream target gene by a linker consisting of a gene sequence encoding amino acid AG.
9. In paragraph 6, A system in which the artificial nucleic acid sequence includes a Shine-Dalgarno sequence, and the expression levels of an upstream target gene and a downstream target gene are controlled by changing the sequence of the Shine-Dalgarno sequence.
10. In paragraph 6, A system wherein the first target gene has the stop codon at the 3' end removed, and the second target gene has the start codon at the 5' end removed.
11. In paragraph 6, The above artificial nucleic acid sequence has a length of 100 to 150 bp, the system.
12. In paragraph 6, A system wherein the artificial nucleic acid sequence is composed of any one of sequence numbers 1 to 10.
13. In paragraph 12, The expression ratio of the upstream target gene and the downstream target gene is If the above artificial nucleic acid sequence consists of sequence number 1, 100:8~10, If it consists of sequence number 2, 100:12~14, If it consists of sequence number 3, 100:16~20, If it consists of sequence number 4, 100:26~29, If it consists of sequence number 5, 100:32~37, If it consists of sequence number 6, 100:40~49, If it consists of sequence number 7, 100:45~51, If it consists of sequence number 8, 100:66~70, If it consists of sequence number 9, 100:81~89, If the sequence number is 10, the system is 100:84~99.
14. A method for controlling the expression level of a target gene using the gene expression control system of any one of claims 6 to 13.
Citation Information
Patent Citations
Regulatory nucleic acid elements
KR101485853B1
Two-way, portable riboswitch mediated gene expression control device
WO2013000278A1