A promoter intelligent design method and device based on local sequence constraint and application thereof
By employing a local sequence-constrained intelligent design method, deep learning and genetic algorithms are used to optimize the generation of synthetic gene regulatory elements, thus solving the problem of insufficient performance of natural elements, achieving efficient design of synthetic gene regulatory elements, and improving gene expression performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2023-01-18
- Publication Date
- 2026-08-04
AI Technical Summary
In existing technologies, natural gene regulatory elements cannot meet the performance requirements of synthetic gene circuit construction, and the use of natural elements may lead to genome instability, affecting the yield and quality of products.
We employ a local sequence constraint-based intelligent design method, utilizing conditional generative adversarial networks and deep learning techniques, combined with genetic algorithms to optimize the generation of synthetic gene regulatory elements. This method preserves prior biological knowledge of the sequence and optimizes flanking regions. By predicting regulatory performance using DenseNet and LSTM networks, we design efficient synthetic promoters.
We designed high-performance synthetic gene regulatory elements, achieving constitutive and inducible promoters with high regulatory expression performance, exceeding the performance of known natural and synthetic elements.
Smart Images

Figure HDA0004073155230000011
Abstract
Description
Technical Field
[0001] This invention relates to the fields of synthetic biology and medicine, specifically to a method, system device, and application for intelligent design of gene regulatory element promoters based on local sequence constraints. Background Technology
[0002] Gene regulatory elements are DNA sequences containing promoters and enhancers that regulate gene expression levels in both time and space, thereby further regulating various life activities such as cell growth, division, and differentiation. Gene regulatory elements can be modularly integrated into gene circuits, independently or collaboratively exercising their function of regulating gene expression. They are fundamental units of synthetic life systems and have wide applications in gene therapy, metabolic pathway optimization, and vaccine production.
[0003] Natural gene regulatory elements are DNA fragments extracted from the natural genome that can regulate gene expression and have been used in the construction of synthetic gene circuits. However, natural elements are insufficient in terms of performance and quantity to meet the needs of gene circuit construction: on the one hand, natural elements evolved by life to meet its own growth and development needs, and cannot meet the ever-increasing performance requirements of humans; more importantly, the use of natural elements can lead to homologous recombination, resulting in genomic instability and affecting the yield and quality of products. Therefore, designing diverse and high-performance synthetic gene elements is of great significance for the construction of synthetic gene circuits and meeting the needs of application.
[0004] Traditional gene regulatory element design methods rely heavily on strong prior biological knowledge, such as transcription factor binding sites (TFBS) and nucleosome arrangement. These sequences are generally considered the core of cis-regulatory logic and a key part of gene regulation. However, recent studies have shown that the regulatory patterns of gene regulatory elements largely depend implicitly on the interactions between TFBS and their flanking regions. These common weak regulatory patterns include the potential dependence between TFBS and their flanking regions, long-distance regulation between regions, or limitations imposed by physicochemical properties. Because these weak regulatory patterns are implicitly represented within the sequence, they cannot be generalized into concise design criteria; however, ignoring these implicit regulatory patterns will reduce the success rate of gene regulation design.
[0005] In recent years, intelligent design strategies have demonstrated a powerful ability to capture complex patterns and have been successfully applied in natural language modeling and image representation learning. Summary of the Invention
[0006] The technical problem to be solved by this invention is how to design and synthesize gene regulatory elements and / or how to design and synthesize gene promoters and / or how to design and synthesize diverse gene regulatory elements with excellent regulatory performance and / or how to design and synthesize gene promoters that regulate high gene expression levels.
[0007] To address the aforementioned technical problems, this invention first provides a method for intelligent design of gene regulatory elements, which may include the following steps:
[0008] A1) Based on the known common sequence and position information of the common sequence of gene regulatory elements, a conditional gene-adversarial network is used to generate a gene regulatory element generation model; the gene regulatory element generation model is used to generate initial gene regulatory elements that conform to the natural distribution; the initial gene regulatory elements contain the common sequence and the flanking sequences of the common sequence;
[0009] A2) Based on the known gene regulatory element sequences and their corresponding gene expression data as the training set, a gene regulatory element regulatory performance prediction model is constructed using the DenseNet neural network and the Long Short-Term Neural Network LSTM (DenseNet-LSTM network); the gene regulatory element regulatory performance prediction model is used to predict the regulatory performance of the initial gene regulatory element;
[0010] A3) Use a genetic algorithm based on population crossover and mutation to iteratively optimize the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model to obtain a gene regulatory element intelligent design model that includes the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model; use the gene regulatory element intelligent design model to design and obtain gene regulatory elements.
[0011] In the above method, the predicted regulatory performance may include predicting the expression level of the regulated gene. The predicted regulatory performance of the gene regulatory element may be higher than that of the known gene regulatory elements.
[0012] The gene regulatory element is derived from the initial gene regulatory element sequence; the gene regulatory element is predicted to have the highest regulatory performance in the initial gene regulatory element sequence.
[0013] The flanking sequence may consist of N, which may be any nucleotide among A, T, C, and G.
[0014] In the above method, A2) may also include the step of optimizing the regulatory performance prediction of the constructed gene regulatory element regulatory performance prediction model: capturing the long-range regulatory function of the known gene regulatory element based on the conditional gene-adversarial network model and attention mechanism to achieve optimization of the gene regulatory element regulatory performance prediction.
[0015] The optimization can be established by a method including the following steps: using an attention mechanism in the generator and discriminator of a conditional generative adversarial network (cGAN) model to learn the long-range regulatory relationships of the known gene regulatory elements, and incorporating the long-range regulatory relationships into the prediction of the regulatory performance of the gene regulatory elements.
[0016] In the above method, the common sequence can be a common sequence of a known inducible gene regulatory element, the gene regulatory element can be an inducible gene regulatory element, and the predicted regulatory performance can be the predicted expression level of the regulated gene.
[0017] The shared sequence may also be a shared sequence of a known constitutive gene regulatory element, and the gene regulatory element may be a constitutive gene regulatory element. The regulatory performance may be the regulation of gene expression level.
[0018] In the above method, the gene regulatory element may be a promoter.
[0019] In the above method, the known inducible gene regulatory element may be the E. coli IPTG inducible promoter and / or the mammalian dox inducible promoter. The constitutive gene regulatory element may be a commonly expressed E. coli promoter.
[0020] To address the aforementioned technical problems, the present invention also provides a device for intelligent design of gene regulatory elements, the device comprising the following modules:
[0021] B1) Gene Regulatory Element Generation Model Construction Module: Used to generate a gene regulatory element generation model based on the common sequence and position information of the known gene regulatory elements using a conditional gene-adversarial network; the gene regulatory element generation model is used to generate an initial gene regulatory element sequence; the initial gene regulatory element sequence contains the common sequence and the flanking sequences of the common sequence;
[0022] B2) Gene Regulatory Element Function Prediction Model Construction Module: This module is used to construct a gene regulatory element regulatory performance prediction model based on the known gene regulatory elements and their corresponding regulatory gene expression data as a training set, using the DenseNet neural network and the Long Short-Term Neural Network LSTM (DenseNet-LSTM network); the gene regulatory element regulatory performance prediction model is used to predict the regulatory performance of the initial gene regulatory element sequence;
[0023] B3) Gene Regulatory Element Intelligent Design System Generation Module: Used to iteratively optimize the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model based on a genetic algorithm of population crossover and variation, to obtain a gene regulatory element intelligent design model that includes the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model; and to design and obtain gene regulatory elements using the gene regulatory element intelligent design model.
[0024] In the above-described device, the predicted regulatory performance can be the predicted regulatory gene expression level. The predicted regulatory performance of the gene regulatory element can be higher than that of the known gene regulatory element sequence.
[0025] The gene regulatory element is derived from the initial gene regulatory element sequence; the gene regulatory element may have the highest predictive regulatory performance in the initial gene regulatory element sequence.
[0026] In the above-described device, the flanking sequence may be composed of N, which may be any nucleotide among A, T, C, and G.
[0027] In the aforementioned device, the gene regulatory element regulatory performance prediction model construction module may further include a regulatory performance prediction optimization module. The regulatory performance prediction optimization module is used to capture the long-range regulatory function of the gene regulatory element based on a conditional generative adversarial network model and an attention mechanism to optimize the prediction of the gene regulatory element's regulatory performance.
[0028] The optimization can be established by a method including the following steps: using an attention mechanism in the generator and discriminator of a conditional generative adversarial network (cGAN) model to learn the long-range regulatory relationships of the known gene regulatory elements, and incorporating the long-range regulatory relationships into the predicted regulatory performance of the gene regulatory elements.
[0029] In the above-mentioned device, the common sequence may be a common sequence of a known inducible gene regulatory element, the gene regulatory element is an inducible gene regulatory element, and the regulatory performance may be the regulation of gene expression level.
[0030] In the above-mentioned device, the common sequence may also be a common sequence of a known constitutive gene regulatory element, the gene regulatory element may be a constitutive gene regulatory element, and the regulatory performance may be the regulation of gene expression level.
[0031] In the above-described device, the gene regulatory element may be a promoter.
[0032] In the above-described apparatus, the known inducible gene regulatory element may be the E. coli IPTG inducible promoter and / or the mammalian dox inducible promoter. The constitutive gene regulatory element may be a commonly expressed E. coli promoter.
[0033] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium for intelligent design of gene regulatory elements, wherein the computer-readable storage medium enables a computer to perform the steps of the method described above.
[0034] The following applications of the methods described above and / or the apparatus described above and / or the computer-readable storage medium described above also fall within the scope of protection of this invention:
[0035] C1) Applications in the development or preparation of synthetic gene products;
[0036] C2) Applications in the development or preparation of products that regulate gene expression;
[0037] C3) Applications in the development or preparation of gene therapy products;
[0038] C4) Applications in the development or preparation of products that optimize biological metabolic pathways;
[0039] Application of C5 in vaccine production.
[0040] This invention provides a system that integrates biological expert knowledge into the design of gene regulatory elements based on deep learning technology. This system has been successfully used to design constitutive and inducible promoters for Escherichia coli and mammals.
[0041] Here, this invention proposes a knowledge-model co-driven learning strategy of "flanking sequence filling" to address this problem. This strategy preserves sequences derived from prior biological knowledge and optimizes flanking regions based on these prior sequences. For example, in optimizing the commonly expressed promoter in *E. coli*, the -10 region sequence "TATAAT" and the -35 region sequence "TAGACT" are retained as prior sequences; in optimizing the IPTG-inducible promoter in *E. coli*, in addition to retaining the -10 and -35 regions, the lacO motif sequence "AATTGTTATCGGATAACAATT" is additionally retained as a prior sequence. In optimizing the mammalian Dox-inducible promoter, the TetO motif sequence "TCCCTATCAGTGATAGAGA" is retained as a prior sequence. Meanwhile, attention mechanisms were added to the generator and discriminator in the conditional generative adversarial network (cGAN) model to learn long-range regulatory relationships. The sequences generated by the generator were iteratively optimized using genetic algorithms and prediction models based on DenseNet structures. Finally, a conditional gene expression regulation molecular machine optimization system based on prior knowledge constraints was obtained.
[0042] The beneficial effects of this invention are as follows:
[0043] The design strategy provided by this invention can learn from data based on domain expert knowledge, enabling the automatic generation of flanking sequences for complex regulatory rules that are difficult to describe with precise rules, and automatically generating expert prior sequences. This invention provides a tool for efficient and rapid component design. It can be applied to the design of constitutive and IPTG-inducible promoters in *E. coli* and Dox-inducible promoters in mammals, successfully designing constitutive promoter elements with high regulatory expression performance and inducible promoter elements with high induction rates, whose performance indicators exceed those of known natural and synthetic elements. Attached Figure Description
[0044] Figure 1 This describes the model establishment and optimization process in the intelligent design (generation) system for gene regulatory elements of the present invention. (a) is the training step of the generated model; (b) is the training step of the predictive model; (c) is the optimization step of the genetic algorithm; and (d) is the optimization path of the gene regulatory element in the functional space. Detailed Implementation
[0045] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0046] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0047] Example 1: Establishment of a Smart Design (Generation) System for Gene Regulatory Elements
[0048] In principle, this invention proposes a knowledge- and data-driven approach. Specifically, it presents a constraint-based optimization framework that integrates prior biological knowledge and data learning. The method incorporates specific functional information, such as TFBS, regulatory RNA binding sites, and ribosome binding sites, to meet the design requirements of regulatory elements.
[0049] Previous attempts at data-driven design methods can be formulated as finding compatible sequence-fitness pairs to maximize the joint probability P(s,prop), while considering both the sequence s and the target attribute prop. Here, PccGEO introduces prior knowledge by preserving the motif sequence m as a constraint in the designed gene sequence s, and models the flanking region as f, i.e.:
[0050] s = (m, f)
[0051] The objective can be written as maximizing the joint probability P(m, f, prop). This can be obtained by applying the chain rule:
[0052] P(m, f, prop)
[0053] =P(m|prop)P(f|m,prop)
[0054] The first stage, P(m|prop), refers to the process of allocating constraint information m compatible with the target attribute based on prior knowledge. The second stage, (f|m, prop), represents the optimization flanking region f conditioned by the constraint region m and the attribute prop. The first stage, prior knowledge integration, refers to specifying the region of optimization constrained by the model based on prior knowledge; and the second stage, sequence optimization, refers to conditional distribution modeling and optimization of the sequence based on the objective function.
[0055] Prior knowledge integration: In the first stage, the sequence and position are determined by maximizing P(m|prop) through the selection of sequence motifs proven crucial for optimizing the target prop. However, there are cases where prior constraints are ambiguous due to insufficient biological knowledge. Therefore, additional weak constraints can be set as a finite variation in the number of nucleotides for the functional element, ensuring the preservation of the main structure of the functional element.
[0056] Sequence optimization: Considering that m*, determined in the first stage, serves as the prior sequence, the second stage is responsible for maximizing P(f|m*)P(prop|f, m*). According to the chain rule,
[0057] P(f|m*,prop)=P(f|m*)P(m*)P(prop|f,m*)
[0058] P(f|m*)P(prop|f,m*)
[0059] P(f|m*) indicates that the optimization should follow the natural regulation rule of the constrained prior sequence m*, thus conditionally approximating the functional element distribution to sample the synthetic elements. Simultaneously, a predictor based on a dense network-LSTM (densenet-LSTM) is trained to evaluate the properties of the input elements, i.e., P(prop|f, m*).
[0060] Finally, a genetic algorithm combined with a generative model and a densenet-LSTM predictor is used to maximize the probability of P(f|m*, prop) to design functional elements with target attributes.
[0061] 1. First stage: Integration of prior knowledge: Construction of a gene regulatory element generation model
[0062] First, biological expert knowledge is formalized into machine-understandable signals. This invention selects known gene regulatory element sequences (such as the -10 sequence 5'-TATAAT-3' and -35 sequence 5'-TTGACA-3' of the E. coli promoter, the lacO motif sequence 5'-AATTGTTATCGGATAACAATT-3' of the lactose operon, and the tetO motif sequence in mammals) from expert knowledge (reported in E. coli and mammalian species). The consensus sequences among these known gene regulatory element sequences and their functional positions in the sequence (the -10 and -35 sequences must be within 50 bases before the transcription start site, the lacO motif must be within 165 bp before the transcription start site, and the tetO motif must be before the minicmv region) are used as conditional input data for the algorithm. Specifically, the common sequence and its corresponding position on the gene need to be extracted from the known gene regulatory element sequences and formalized into a fixed-length sequence (165 bp in E. coli and 150 bp in mammals). In this fixed-length sequence containing the gene regulatory element sequence, the common sequence represents the binding site sequence, and the remaining positions (flanking sequences of the common sequence) are marked with the letter 'N'. N represents any one of the nucleotides A, T, C, and G.
[0063] Subsequently, using the extracted common sequences and their positional information as conditional inputs, a convolutional neural network is used to train a conditional deep generative model to estimate the functional distribution of regulatory elements, transforming the conditional inputs into functional sequences conforming to the distribution. When training the conditional generative model (Conditional Generative Adversarial Network, cGAN), a portion of the known functional sequence (gene regulatory element sequence) is masked with 'N' placeholders to simulate the conditional input, and the complete sequence is used as the output to generate gene regulatory element sequences of the initial design conforming to the natural distribution. The deep generative model consists of convolutional layers, attention layers, and fully connected layers.
[0064] 2. Second-stage sequence optimization: Construction of a predictive model for the regulatory performance of gene regulatory elements.
[0065] Subsequently, a gene regulatory element performance prediction model was trained, using the known gene regulatory element sequences from step 1 as input and the predicted expression activity values of the gene regulatory elements (obtained through transcriptional or translational intensity from public datasets) as output. Here, gene regulatory element sequences and experimentally quantitatively measured expression levels of the corresponding sequences were selected as the training set (downloadable from E.coli: https: / / www.nature.com / articles / nmeth.4633, Mammalian: https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE71279). A gene regulatory element performance prediction model (densenet-LSTM) was constructed using the DenseNet network structure in convolutional neural networks to capture long-range regulatory relationships within gene regulatory element sequences, achieving more accurate performance predictions for gene regulatory elements. Specifically, an attention mechanism was added to the generator and discriminator of a conditional generative adversarial network (cGAN) model to learn long-range regulatory relationships.
[0066] The gene regulatory element performance prediction model can predict the regulatory performance of the initial gene regulatory element sequence generated in step 1 (predicting the expression level of the regulated gene), and then screen the regulatory element sequence with the optimal functional parameters based on the regulatory performance prediction model and genetic algorithm.
[0067] 3. Intelligent Design System for Gene Regulatory Elements Based on Genetic Algorithms
[0068] Finally, the genetic algorithm based on population crossover and mutation is used for iterative optimization. Combined with the conditional gene regulatory element generation model obtained in step 1 and the gene regulatory element regulatory performance prediction model (densenet-LSTM) obtained in step 2, the gene regulatory elements are fully synthesized and designed.
[0069] In the intelligent design system for gene regulatory elements, conditional prior inputs (for different species, such as the -10 sequence 5'-TATAAT-3' and the -35 sequence 5'-TTGACA-3' of the E. coli promoter, the lacO motif sequence 5'-AATTGTTATCGGATAACAATT-3' of the lactose operon, and the tetO motif sequence in mammals) are fed into the gene regulatory element generation model to generate initial gene regulatory element sequences that conform to the natural distribution. Subsequently, these generated initial gene regulatory element sequences are fed into the densenet-LSTM prediction model to predict their regulatory expression characteristics. Finally, the genetic algorithm is used for iterative optimization to obtain gene regulatory elements with good regulatory performance that conform to the natural distribution.
[0070] Example 2: Application of the Intelligent Design System for Gene Regulatory Elements
[0071] 1. Optimization design of commonly expressed promoters in E. coli
[0072] The gene regulatory element intelligent design system established in Example 1 was used to optimize the design of commonly expressed promoters of Escherichia coli.
[0073] In the gene regulatory element generation model of the intelligent design system for gene regulatory elements, based on prior knowledge, the -10 and -35 sequences of the E. coli promoter are used as common sequences of known gene regulatory element sequences. The flanking sequences of the common sequences are designed to generate three optimized E. coli commonly expressed promoters.
[0074] For three commonly used E. coli promoters with different regulatory strengths in the iGEM database (http: / / parts.igem.org / Promoters / Catalog / Constitutive), namely BBa_J23119, BBa_J23118, and BBa_J23114, the intelligent gene regulatory element design system established in Example 1 was used for optimization design. Experimental results showed that the expression levels of the regulatory genes in the three optimized promoters designed using the intelligent gene regulatory element design system were all higher than those in the original three natural promoters.
[0075] The original sequences of the three constant promoters are as follows:
[0076] BBa_J23119: 5'-TTGACAGCTAGCTCAGTCCTAGGTATAATGCTAGC-3' (Regulated gene expression level: 2.19);
[0077] BBa_J23118: 5'-TTGACGGCTAGCTCAGTCCTAGGTATTGTGCT-3' (Regulated gene expression level: 0.75);
[0078] BBa_J23114:5'-TTTATGGCTAGCTCAGTCCTAGGTACAATGCTAGC-3'(Regulated gene expression level: 0.04).
[0079] The sequences of three common promoters designed and optimized using the intelligent design system for gene regulatory elements established in Example 1 are as follows:
[0080] Optimized BBa_J23119:
[0081] Conditional prior input sequence:
[0082] 5'-NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTGACANNNNNNNNNNNNNNNNNNNTATAATNNNNNNNNNNNNNNNNNNNNNNNNNNNNN-3';
[0083] Optimized sequence:
[0084] 5'-CAAAAAAAAAAAAAATGTTGCGCTACTTCGCCTTTTATCTTAAATTGACGACAGGGAACCCCCCGAGGAATGCCGAAGTATGCACGTGTTTCTCTTTTTTATGGTGTTGACAACTTACACTAAATCTGTTATAATGATATATCAAAAAATTAAAGGAGATTATTG-3' (Regulated gene expression level: 2.29);
[0085] Optimized BBa_J23118:
[0086] Conditional prior input sequence:
[0087] 5’-NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTTGACGNNNNNNNNNNNNNNNNNNTATTGTNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN-3’
[0088] Optimized sequence:
[0089] 5’-CAAAAAAAACTTGAGTTCTGGGTTATCTGTATATTATGTATAACTTGATATGCGGTAAAAAGGCGCAATAGAGCGAGACGATTATCTTCACATAAATGAAAAAGTGTTGACGACTCTCACTTGTTATGTTATTGTGGTAAGCCATAATACTAAATGAGATACAGA-3’(Regulatory gene expression level: 2.56);
[0090] Optimized BBa_J23114:
[0091] Conditional prior input sequence:
[0092] 5’-NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTTTATGNNNNNNNNNNNNNNNNNNTACAATNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN-3’
[0093] Optimized sequence:
[0094] 5’-CAAAAAAGATTGACTCAGGGCGTTTCGCGGTTACAATAGCCCATAAAAGGAGCCGAGTCACACAGGGGTATTAGGAACGGATAAGGCAAAAACTAATTAAAAAGTGTTTATGAAATTCTCTAAATATGGTACAATGATTAATAAAAAAAGTAAAGGAGAGATTGC-3’(Regulatory gene expression level: 2.61).
[0095] The method for determining the expression level of promoter-regulated genes is as follows: Three original constant promoters and three optimized constant promoter sequences were synthesized. Using the NEBuilder HiFi DNA assembly reaction system, the promoter sequence (nucleotides 2001-2035 of sequence 1) of the constitutive promoter test plasmid (sequence 1 in the sequence listing) was replaced by Gibson assembly to obtain six recombinant plasmid vectors. The six recombinant plasmid vectors were transformed into Escherichia coli trans5α strain (Full Gold, CD201-02) to obtain six recombinant strains (J23119-E.coli, J23118-E.coli, J23114-E.coli, D-J23119-E.coli, D-J23118-E.coli and D-J23114-E.coli).
[0096] Six recombinant bacterial strains and the original trans5α strain without transformed plasmid were cultured overnight (16 hours) in 5 mL LB broth containing 50 μg / mL chloramphenicol at 37°C and 220 rpm. Seven overnight cultures were then diluted 1:100 in fresh LB broth containing 50 μg / mL chloramphenicol, and the experiments were repeated three times. After 6 hours of incubation, promoter activity was measured. 150 μL of culture was plated in a 96-well plate (Corning 3603), and OD600 was measured using a Varioskan Flash (Thermo) microplate reader. nm Absorbance and sfGFP fluorescence intensity (excitation at 485 nm and emission at 520 nm; nucleotides 2137-2672 of sequence 1 are the sfGFP coding sequence). OD600 of fresh LB medium was also measured. nm The absorbance value served as a blank control for bacterial concentration, and the fluorescence intensity of the original strain sfGFP without plasmid transformation served as a blank control for fluorescence expression level.
[0097] Promoter-regulated gene expression level is defined as (OD600 of recombinant bacterial cultures containing the promoter) nm Absorbance (LB fresh medium blank control) / (fluorescence intensity of sfGFP in recombinant strain culture containing promoter - fluorescence intensity of sfGFP in original strain without plasmid transformation) is the average value of different replicates.
[0098] 2. Optimization of E. coli inducible expression promoter
[0099] The gene regulatory element intelligent design system described in Example 1 was used to design and optimize the inducible promoter of E. coli.
[0100] In the gene regulatory element generation model of the intelligent design system for gene regulatory elements, based on prior knowledge, the -10 and -35 sequences of the E. coli promoter are used as common sequences of known gene regulatory element sequences. At the same time, 2, 3 or 4 tandem repeat lacO sites are used as prior sequences (templates for designing and generating gene regulatory element sequences). The flanking sequences of the common sequences are designed to generate three optimized E. coli inducible expression promoters.
[0101] The results showed that the optimized E. coli inducible expression promoter was a promoter sequence with a high induction rate, which was higher than that of the IPTG inducible promoter.
[0102] Methods for obtaining induction rate data:
[0103] The original IPTG inducible promoter and three optimized E. coli inducible expression promoter sequences were synthesized. Using the NEBuilder HiFi DNA assembly reaction system, the promoter sequence (nucleotides 2293-2327 of sequence 2) of the inducible promoter test plasmid (sequence 2 in the sequence listing) was replaced by Gibson assembly to obtain six recombinant plasmid vectors. The six recombinant plasmid vectors were chemically transformed into trans5α strain (Full Gold, CD201-02) to obtain four recombinant strains (IPTG-E.coli, D-2lac-IPTG-E.coli, D-3lac-IPTG-E.coli and D-4lac-IPTG-E.coli).
[0104] Four recombinant bacterial strains and the original trans5α strain without plasmid transformation were cultured overnight (16 hours) in 5 mL LB broth containing 50 μg / mL chloramphenicol at 37°C and 220 rpm. The five overnight cultures were then diluted 1:100 in fresh LB broth containing 50 μg / mL chloramphenicol. During dilution, 0.1 mM IPTG was added to the broth to induce fluorescent protein expression. This experiment was repeated three times, and after 6 hours of incubation, promoter activity was measured. 150 μL of culture was transferred to a 96-well plate (Corning 3603), and OD600 was measured using a Varioskan Flash (Thermo) microplate reader. nm Absorbance and sfGFP fluorescence intensity (excitation at 485 nm and emission at 520 nm; nucleotides 2429-2964 of sequence 2 are the sfGFP coding sequence). OD600 of fresh LB medium was also measured. nmThe absorbance value served as a blank control for bacterial concentration, and the fluorescence intensity of the original strain sfGFP without plasmid transformation served as a blank control for fluorescence expression level.
[0105] Promoter-regulated gene expression level is defined as (OD600 of recombinant bacterial cultures containing the promoter) nm Absorbance (LB fresh medium blank control) / (sfGFP fluorescence intensity of recombinant strain culture containing promoter - sfGFP fluorescence intensity of original strain without plasmid transformation) is the average of different replicates. Induction rate is defined as the ratio of promoter-regulated gene expression level after induction to promoter-regulated gene expression level before induction.
[0106] The sequence of the IPTG-induced promoter is as follows:
[0107] 5'-AATTGTGAGCGGATAACAATTGACATTGTGAGCGGATAACAAGATACTGAGCAC-3' (Induction rate: 22.75987386);
[0108] The optimized sequence of induced promoters is as follows:
[0109] 2lac (obtained by optimizing two tandemly repeated lacO sequences as prior sequences):
[0110] Conditional prior input sequence:
[0111] 5'-NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNAATTGTTATCGGATAACAATTNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNCTTACTNNNNNNNNNNNNNNNNNNNNNNTAAAATNNNNNNAAATTGTGAGCGCTCACAATTNNNNNNNNN-3';
[0112] Optimized sequence:
[0113] 5'-CAAAATTTTGCTTGCTTTTAAGTGTTTCTATATAATTGTTATCGGATAACAATTCGGAAGGTCGGCGGAAGGAGGCGGAATCCAAACATGAAAAGATATTTTAAACTTTTTAAAAACGTGGTATACTGTTATTAAATTGTGAGCGCTCACAATTGGATAGGAG-3' (induction rate: 3.27);
[0114] 3lac (obtained by optimizing a prior sequence of 3 tandemly repeated lacO sequences):
[0115] Conditional prior input sequence:
[0116] 5'-NNNNNNNNNNNNNNAATTGTGAGCGGATAACAATTGGCAGTGAGCGCAACGCAATTNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTTTAAANNNNNNNNNNNNNNTATACTNNNNNNAAATTGTGAGCGCTCACAATTNNNNNNNNN-3';
[0117] Optimized sequence:
[0118] 5'-AAACGCCGTCAATTAATTGTGAGCGGATAACAATTGGCAGTGAGCGCAACGCAATTATGATGAAAAGCATTATTTCAATGCTATTTTTTGGGTTTTTGGTTTTAAAAAATAAAACTCAGGGGTATACTTAGATTAAATTGTGAGCGCTCACAATTGAAAAACAT-3' (induction rate: 18.16);
[0119] 4lac (obtained by optimizing a prior sequence of 4 tandemly repeated lacO sequences):
[0120] Conditional prior input sequence:
[0121] 5'-NNNNAATTGTGAGCGGATAACAATTNNNNNNNNNNAATTGTTATCGGATAACAATTNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNCTTACTNNNNNTTGTGAGCGGATAACAATAAAATNNNNNNAAATTGTGAGCGCTCACAATTNNNNNNNNN-3';
[0122] Optimized sequence:
[0123] 5'-CATTAATTGTGAGCGGATAACAATTATATTTTTTAAATTGTTATCGGATAACAATTCACTACTTATTAATAAAAATTAAGTTATTTTTCTGACATCTTACTGTAGTTTGTGAGCGGATAACAATAAAATGGAACAAAATTGTGAGCGCTCACAATTTGATCTCGA-3' (induction rate: 61.83).
[0124] 3. Optimization of mammalian inducible expression promoters
[0125] The intelligent design system for gene regulatory elements described in Example 1 was used to design and optimize mammalian dox inducible promoters.
[0126] In the gene regulatory element generation model of the intelligent gene regulatory element design system, based on prior knowledge, three spaced tandem repeats of the tetO sequence from mammals are used as prior sequences (templates for designing and generating gene regulatory element sequences). Flanking sequences of the shared sequences are designed to generate optimized mammalian inducible expression promoters with high inducibility rates. The inducibility rate of the optimized mammalian inducible expression promoter is higher than that of the mammalian dox-inducible promoter.
[0127] The original mammalian dox-inducible promoter and the optimized mammalian dox-inducible promoter sequences were synthesized separately. Using the NEBuilder HiFi DNA assembly reaction system, the promoter sequence of the mammalian promoter sequencing plasmid pwx158 (sequence 3 in the sequence listing) was replaced with nucleotides 3616-3867 using the Gibson assembly method to obtain two recombinant mammalian promoter plasmid vectors. The HEK293(293-H) cell line was used as the test host. Cells were cultured in DMEM medium containing 4.5 g / L glucose (GIBCO, 11965118) supplemented with 10% FBS (GIBCO, 16000-044), 1x NEAA (GIBCO, 11140050), and 0.5x penicillin-streptomycin (Solarbio, P1400) at 37°C and 5% CO2.
[0128] Lipofectamine LTX and PLUS reagent (Invitrogen, 15338100) were used for transient transfection of plasmids. In the transfection experiment, approximately 1.8 × 10⁵ HEK293 cells were seeded into 12-well plates containing 1 mL of fresh culture medium and allowed to grow for approximately 24 hours until 70-90% confluence was achieved. Before transfection, Dox was added to the culture medium at a final concentration of 1 μg / mL to induce promoter expression. Transient transfection was performed in each well by adding 600 ng of the recombinant mammalian promoter plasmid vector, 600 ng of the rtTa protein expression plasmid (sequence 4 in the sequence listing), and appropriate transfection reagents. The culture medium contained 1 μg / mL Dox before transfection to induce promoter expression. After transfection, cells were cultured for another 24 hours and then harvested for flow cytometry analysis. Cells were first digested with trypsin and then centrifuged at 300 g for 5 minutes at room temperature. The recombinant cells were then washed once with PBS and resuspended in 1×PBS to a total volume of 400 μL. Cells were then analyzed using an LSR Fortessa (BD Biosciences). The excitation laser (Ex), emission filter (Em), and photomultiplier tube (PMT) voltages used for the corresponding fluorescent protein measurements were as follows: TagBFP (Ex: 405nm laser, Em: 450 / 50 filter, PMT: 350V). EYFP (Ex: 488nm laser, Em: 530 / 30 filter, PMT: 200V). Approximately 1 × 10⁻⁶ cells were collected for each sample. 6 Individual cell events were used for downstream analysis. For data analysis, raw data was filtered using the FlowAI plugin to remove low-quality data. TagBFP intensity was selected at 1×10⁻⁶. 4 ~5×10 4 Cells containing appropriate concentrations of rtTA protein were selected. The average EYFP value was used as the activity of the promoter to be tested. Three independent biological replicates were performed for each promoter. The induction rate was defined as the ratio of the promoter activity after induction to the promoter activity before induction.
[0129] The sequence of the dox-inducible promoter is as follows:
[0130] 5'-GAGTTTACTCCCTATCAGTGATAGAGAACGTATGTCGAGTTTATCCCTATCAGTGATAGAGAACGTATGTCGAGTTTACTCCCTATCAGTGATAGAGAACGTATGTCGA-3' (Induction rate: 304.37);
[0131] The optimized sequences of mammalian inducible promoters are as follows:
[0132] Conditional prior input sequence:
[0133] 5'-NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNTCCCTATCAGTGATAGAGANNNNNNNNNNNNNNNTCCCTATCAGTGATAGAGANNNNNNNNNNNNNNNTCCCTATCAGTGATAGAGANNNNNNNNN-3';
[0134] Optimized sequence:
[0135] 5'-GTTACGACTTCCGGCTTCGGGGGCGCCTCCGCCGCAAGTCACCCACCGGTCCCTATCAGTGATAGAGAAGGCACGTGACGTCCCTCCCTATCAGTGATAGAGACGGACGCACAGCGGAAGTCCCTATCAGTGATAGAGAGTCTCCGATCT-3' (induction rate: 409.83).
[0136] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.
Claims
1. A method for intelligent design of genetic regulatory elements, characterized by: The method includes the following steps: A1) Based on the known common sequences and positional information of common sequences of gene regulatory elements, a conditional gene-adversarial network is used to generate a gene regulatory element generation model; the gene regulatory element generation model is used to generate initial gene regulatory element sequences that conform to the natural distribution. The initial gene regulatory element sequence contains the common sequence and the flanking sequences of the common sequence; A2) Based on the known gene regulatory elements and their corresponding gene expression data as the training set, a gene regulatory element regulatory performance prediction model is constructed using the DenseNet neural network and the Long Short-Term Neural Network (LSTM); the gene regulatory element regulatory performance prediction model is used to predict the regulatory performance of the initial gene regulatory element. A3) Use a genetic algorithm based on population crossover and mutation to iteratively optimize the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model to obtain a gene regulatory element intelligent design model that includes the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model. Gene regulatory elements were designed and obtained using the aforementioned intelligent design model for gene regulatory elements. The gene regulatory element is a promoter.
2. The method of claim 1, wherein: A2) also includes the step of optimizing the regulatory performance prediction of the constructed gene regulatory element regulatory performance prediction model: capturing the long-range regulatory function of the known gene regulatory elements based on the conditional gene-adversarial network model and attention mechanism to achieve the optimization of gene regulatory element regulatory performance prediction.
3. The method according to claim 1 or 2, characterized in that: The shared sequence is a shared sequence of a known inducible gene regulatory element, and the gene regulatory element is an inducible gene regulatory element; Alternatively, the shared sequence may be a shared sequence of known constitutive gene regulatory elements, wherein the gene regulatory element is a constitutive gene regulatory element.
4. An apparatus for intelligent design of genetic regulatory elements, characterized by: The device includes the following modules: B1) Gene regulatory element generation model construction module: used to generate gene regulatory element generation models using conditional gene-growth adversarial networks based on the common sequences and positional information of known gene regulatory elements; the gene regulatory element generation model is used to generate initial gene regulatory element sequences; The initial gene regulatory element sequence contains the common sequence and the flanking sequences of the common sequence; B2) Gene Regulatory Element Function Prediction Model Construction Module: This module is used to construct a gene regulatory element regulatory performance prediction model based on the known gene regulatory elements and their corresponding regulatory gene expression data as a training set, using the DenseNet neural network and the Long Short-Term Neural Network (LSTM); the gene regulatory element regulatory performance prediction model is used to predict the regulatory performance of the initial gene regulatory element sequence. B3) Gene Regulatory Element Intelligent Design System Generation Module: Used to iteratively optimize the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model based on a genetic algorithm of population crossover and variation, to obtain a gene regulatory element intelligent design model that includes the gene regulatory element generation model and the gene regulatory element regulatory performance prediction model. Gene regulatory elements were designed and obtained using the aforementioned intelligent design model for gene regulatory elements. The gene regulatory element is a promoter.
5. The apparatus of claim 4, wherein: B2) also includes a regulation performance prediction and optimization module; The regulatory performance prediction and optimization module is used to capture the long-range regulatory functions of the known gene regulatory elements based on a conditional gene-adversarial network model and an attention mechanism, so as to optimize the prediction of the regulatory performance of gene regulatory elements.
6. The apparatus of claim 4 or 5, wherein: The shared sequence is a shared sequence of a known inducible gene regulatory element, and the gene regulatory element is an inducible gene regulatory element; Alternatively, the shared sequence may be a shared sequence of known constitutive gene regulatory elements, wherein the gene regulatory element is a constitutive gene regulatory element.
7. A computer-readable storage medium with intelligent design of gene regulatory elements, characterized in that: The computer-readable storage medium enables a computer to perform the steps of the method according to any one of claims 1-3.
8. The method of any one of claims 1-3 and / or the apparatus of any one of claims 4-6 and / or the computer-readable storage medium of claim 7, in any of the following applications: C1) Applications in the development or preparation of synthetic gene products; C2) Applications in the development or preparation of products that regulate gene expression; C3) Applications in the development or preparation of gene therapy products; C4) Applications in the development or preparation of products that optimize biological metabolic pathways; Application of C5 in vaccine production.