A generated super-high-activity plant core promoter library and application thereof
By generating highly active plant core promoters using generative adversarial networks and activation maximization techniques, and constructing a promoter library, we have solved the problems of reliance on prior knowledge and high cost in traditional methods, and achieved a significant improvement in promoter activity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAZHONG AGRI UNIV
- Filing Date
- 2026-03-11
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies make it difficult to design plant promoters with a wider range of activities. Traditional methods rely on prior knowledge and are costly to experiment, making it difficult to overcome the activity limitations of natural promoters.
Generative Adversarial Networks (GANs) are used to learn the distribution of natural promoters, and activation maximization techniques are combined to generate ultra-high-activity plant core promoters. A promoter library is constructed through high-throughput screening, and artificial intelligence is used to deeply mine sequence features.
The generated promoter library contains promoters with higher activity than the natural maize ubiquitin promoter, with some promoters exhibiting activity more than 16 times that of the 35S minimum promoter. This solves the problem of limited promoter activity in existing technologies and meets the needs of plant synthetic biology.
Smart Images

Figure CN121801916B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioengineering technology, specifically to a generated ultra-high activity plant core promoter library and its applications. Background Technology
[0002] Currently existing natural promoters have a narrow activity range; in maize protoplasm, the maximum activity of natural promoters is only 16 times that of the smallest 35S promoter. Furthermore, natural promoters occupy only a small portion of the random sequence space (for example, 4 in a 170bp sequence). 170 ≈2.23×10 102 (There are many combinations), so designing novel promoters that outperform natural promoters in this vast sequence space and constructing generative promoter libraries with a wider range of activities has become an urgent need in plant synthetic biology.
[0003] Traditional plant promoter design methods (such as mutation induction and regulatory element combination assembly) have shortcomings: these methods mainly involve minor modifications or limited recombination of natural promoters, and their effectiveness is highly dependent on prior knowledge and design experience of plant promoter regulatory mechanisms; moreover, because this method is limited to exploring the neighboring sequence space of natural promoters, the perturbations caused are relatively small, making it difficult to achieve functional breakthroughs; in addition, a large number of candidate promoter sequences generated lack the expected activity and require large-scale experimental screening, which significantly increases development costs and time investment.
[0004] In recent years, artificial intelligence (AI) technology has made significant breakthroughs in fields such as speech recognition and autonomous driving. Generative AI, based on AI, has shown great potential in synthetic biology. This type of technology enables data-driven sequence space exploration, allowing models to autonomously discover regulatory patterns from massive amounts of data without relying on prior knowledge. Summary of the Invention
[0005] To address the aforementioned technical limitations, this application proposes a highly active plant core promoter library, its construction method, and its application; it overcomes the deficiencies and defects mentioned in the background technology.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] The first inventive point of this application is to provide a generated ultra-high activity plant core promoter library, the promoter library comprising promoters with nucleotide sequences shown in SEQ ID No. 1-SEQ ID No. 7.
[0008] The specific nucleotide sequences of the promoter library are shown in Table 1.
[0009] Table 1
[0010]
[0011] Optionally, in the above-mentioned plant core promoter library, the promoters of the nucleotide sequences shown in SEQ ID No. 1-SEQ ID No. 7 have the ability to sequentially weaken the expression of the target gene.
[0012] All the generative promoters in the library exhibited higher transcriptional activity than the natural maize ubiquitin (UBI) promoter; among them, the activity of the promoters was more than 16 times that of the minimum 35S promoter.
[0013] Optionally, the plant in the above-mentioned plant core promoter library is selected as maize.
[0014] The second inventive point of this application is to provide a vector containing the nucleotide sequences shown in SEQ ID No. 1-SEQ ID No. 7 as promoter elements.
[0015] The third inventive point of this application is to provide a genetically engineered host cell containing the aforementioned vector; or,
[0016] The host cell genome includes nucleic acids with nucleotide sequences shown in SEQ ID No. 1-SEQ ID No. 7.
[0017] The fourth inventive point of this application is to provide the use of the above-mentioned plant core promoter library, which is used for plant core promoters to operably link the plant core promoters to a target gene to regulate the expression of the target gene; the target gene is the firefly luciferase reporter gene fLuc.
[0018] This plant core promoter library is mainly used to drive the expression of exogenous genes in plants; the exogenous genes include, but are not limited to, reporter genes (such as the aforementioned firefly luciferase reporter gene fLuc), resistance genes, or agronomic trait improvement genes.
[0019] The fifth inventive point of this application is to provide a method for constructing a library of plant core promoters with ultra-high activity. This method utilizes generative adversarial networks to learn the distribution of natural plant core promoters and generate generative promoters with similar characteristics to natural promoters. It then uses activation maximization techniques to generate promoters with specified activities. Based on multiple different targets, promoters are generated, and then generative promoters with activities higher than the maximum value of natural promoters are selected, which is the aforementioned library of plant core promoters with ultra-high activity.
[0020] Optionally, the above construction method specifically includes the following steps:
[0021] 1) Utilize Generative Adversarial Networks (GANs) to learn the distribution of natural promoters and generate generative promoters with similar characteristics to natural promoters;
[0022] 2) The activation maximization technique is used to couple the generator of the generative adversarial network (GAN) with a pre-trained predictor that can predict promoter activity, thereby generating a promoter with a specified activity.
[0023] 3) Set multiple different objectives and select the generation promoters that meet the objective requirements;
[0024] 4) The activity of the generated promoters selected in step 3) is detected by a high-throughput method. Then, all generated promoters with activities higher than the maximum activity of natural promoters are selected to obtain plant core promoters and construct a plant core promoter library.
[0025] Optionally, in the above construction method, in step 1), the generative adversarial network (GAN) is selected as the WGAN-GP architecture, which consists of two networks: a generator and a discriminator; in step 3), nine targets are selected, with target values being the maximum value (max), 2, 1, 0, -1, -2, -3, -4, and the minimum value (min). 5250 generator promoters that meet the target requirements are selected, including 1500 from the maximum value (max), 250 from the minimum value (min), and 500 from each of the remaining target values; in step 4), the high-throughput method is selected as the STARR-seq method.
[0026] Optionally, the STARR-seq method described above includes the following steps:
[0027] a) Design the library and construct the promoter-barcode library;
[0028] b) Prepare and transform maize protoplasts;
[0029] c) Perform STARR-seq experiments and Illumina sequencing;
[0030] d) Further analysis of the results data from the STARR-seq experiment.
[0031] Compared with the prior art, this application has the following advantages:
[0032] This application utilizes artificial intelligence technology to deeply mine the sequence characteristics of plant core promoters, designing and screening ultra-high activity plant core promoters to form a promoter library. Validated using a maize protoplast system, the seven plant core promoters protected in this application showed superior activity to the natural strong promoter UBI in STARR-seq assays. In dual-luciferase (LUC) experiments, the activities of all seven promoters reached more than 16 times that of the smallest 35S promoter, with two promoters still showing higher activity than UBI. This indicates that this promoter library can effectively address the problem of limited activity in existing natural elements, meeting the practical needs of plant synthetic biology for ultra-high activity regulatory elements. Attached Figure Description
[0033] Figure 1 This is a diagram of the WGAN-GP architecture.
[0034] Figure 2 To generate the TargetGAN architecture diagram for a specified active promoter by coupling the generator of WGAN-GP with a pre-trained predictor capable of predicting promoter activity using the activation maximization technique.
[0035] Figure 3 The graph shows the statistical results of how the activity of the generated promoter gradually approaches the corresponding target with the number of iterations under the nine set objectives.
[0036] Figure 4 This is a flowchart of the STARR-seq experiment.
[0037] Figure 5 The results of the STARR-seq experiment are shown in the figure. Analysis reveals that among the 3763 generative promoters, 30 generative promoters exhibit activities exceeding the maximum value of the natural promoters; among them, Figure 5 A represents the Pearson correlation coefficient between the experimental values and the predicted values from the promoter model. Figure 5 B represents the statistical distribution of promoter activity on different targets. Figure 5 C represents the generation of a promoter that exceeds the maximum activity range of the natural promoter.
[0038] Figure 6 The Luc detection results are shown in the figure, after comprehensive consideration, for the seven generative promoters with relatively strong activity; among them, Figure 6 A represents the detection procedure for dual-luciferase reporter genes. Figure 6 B shows a comparison of STARR-seq and luciferase activity data for 7 ultra-high activity promoters, UBI natural promoters, and 4 low activity promoters; in the figure, the LUC activity of the 7 promoters in this application is more than 16 times that of the smallest 35S promoter (y-axis greater than 4). Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this application clearer, a more detailed description is provided below. However, it should be understood that the description herein is merely for explaining this application and is not intended to limit its scope.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. All reagents and instruments used herein are commercially available, and the characterization methods involved can be found in relevant descriptions in the prior art, and will not be repeated here.
[0041] To further understand this application, the following detailed description is provided in conjunction with the preferred embodiments.
[0042] Example 1
[0043] I. Utilize Generative Adversarial Networks (GANs) to learn the distribution of natural promoters and generate generative promoters with similar characteristics to natural promoters.
[0044] The training dataset in this embodiment contains natural plant core promoter sequences from Arabidopsis thaliana, maize, and sorghum (Jores, T. et al., Synthetic promoter designs enabled by a comprehensive analysis of plant core promoters). Nat. Plants 7, 842–855 (2021)). Activity data were obtained using STARR-seq technology in maize protoplasts (dark conditions, no enhancers). The dataset contains 76,851 sequences, which were randomly divided into training, validation, and test sets in an 8:1:1 ratio.
[0045] In terms of model construction, the WGAN-GP architecture is specifically adopted, such as... Figure 1 As shown, WGAN-GP consists of two networks: a generator and a discriminator. These two networks have a symmetrical architecture, each comprising one convolutional neural network layer, five residual blocks, and one fully connected layer, arranged in opposite directions. Each residual block contains two CNN layers and one skip connection, using a residual scaling factor of 0.3. Each CNN layer in the residual block has 200 filters, a kernel size of 5, and a stride of 1. Furthermore, the CNN layers in both the generator and discriminator employ a convolutional configuration with a kernel size of 1 and a stride of 1.
[0046] During the data preprocessing stage, the natural promoters are transformed into a 1×170×4 tensor through one-hot encoding, where each base is represented as a four-dimensional binary vector: adenine (A) is [1, 0, 0, 0], cytosine (C) is [0, 1, 0, 0], guanine (G) is [0, 0, 1, 0], and thymine (T) is [0, 0, 0, 1].
[0047] To generate sequences similar to real data structures, the generator takes latent variables sampled from a Gaussian distribution as input. Its fully connected layer outputs a size of 170×200, corresponding to the sequence length and the CNN filter dimension. The last CNN layer in the generator uses a softmax activation function and contains four filters, representing the four bases (A, C, G, T), resulting in a tensor shape of 1×170×4 for the generated sequence. The discriminator then receives both the one-hot encoded natural promoter tensor and the generator's output tensor as input. The discriminator's CNN layer contains 200 filters, the same number as in the residual block. The discriminator's final output is a scalar score used to distinguish whether the input sequence comes from a real natural promoter or is a synthetic sequence generated by the generator. Based on this structure, the generator and discriminator engage in alternating adversarial training. The generator continuously optimizes its parameters to generate more realistic sequences to "deceive" the discriminator, while the discriminator continuously improves its ability to distinguish between real and fake sequences. After extensive iterative training (54,000 iterations in this study), the two reached equilibrium in a dynamic game, enabling the generator to fully extract and master the potential distribution patterns of natural promoters, thereby generating plant core promoters with high biological rationality and novel sequences.
[0048] During model training, Wasserstein distance was used as the loss function, combined with gradient penalty to evaluate the difference between the generated data distribution and the real data distribution. The model used the Lion optimizer and ReLU activation function. The learning rate was set to 1e-5, beta1 to 0.9, beta2 to 0.99, and the batch size was set to 64. To ensure training stability, the parameter update ratio of the discriminator to the generator was strictly set to 5:1.
[0049] 2. The generator of WGAN-GP is coupled with a pre-trained predictor that can predict promoter activity using the activation maximization technique, thereby generating promoters with specified activity. This architecture is named TargetGAN.
[0050] TargetGAN architecture such as Figure 2 As shown.
[0051] The specific implementation steps are as follows: The trained generator uses the latent variable Z to generate synthetic promoters and predicts their activity using a pre-trained predictor. Then, the loss is calculated based on the difference between the predicted promoter activity and the target value. Through backpropagation, the gradient of the loss with respect to the latent variables is calculated. Subsequently, these latent variables are adjusted in directions that bring the promoter activity closer to the target value, guided by the gradient. This process is iterated multiple times until the generated promoter activity approaches the given target. Depending on the specific target, the system defines a piecewise loss function. Let the given objective be... The goal This represents the user-defined desired promoter activity level, specifically the log2 transformation ratio of the generated promoter activity relative to the minimum 35S promoter activity. Let the mean and variance of the promoter activity generated in the current iteration be... and The formula is as follows:
[0052]
[0053] Calculate the gradient using the backpropagation algorithm. Then move along the gradient direction (where (In small increments), gradually adjust the original input. To reduce the value of the loss function:
[0054] .
[0055] Third, based on actual needs, nine different targets were designed, from which 5250 generative promoters were selected for STARR-seq experiments.
[0056] Based on actual needs, and to comprehensively cover and exceed the activity range of natural promoters, this embodiment designed nine different target values. These nine target values are: maximum value (max, corresponding to a predicted activity range ≥ 3), 2 (corresponding to a range of 2 ± 0.1), 1 (corresponding to a range of 1 ± 0.1), 0 (corresponding to a range of 0 ± 0.1), -1 (corresponding to a range of -1 ± 0.1), -2 (corresponding to a range of -2 ± 0.1), -3 (corresponding to a range of -3 ± 0.1), -4 (corresponding to a range of -4 ± 0.1), and minimum value (min, corresponding to a predicted activity range ≤ -6). Figure 3 As shown, under these nine objectives, as the number of iterations increases (10,000 iterations in this example), the population mean of the generated promoter activity rapidly approaches the objective, and the distribution variance gradually decreases, eventually stabilizing and precisely aligning to the given objective interval.
[0057] After generation, this application selected 5250 generated promoters for STARR-seq high-throughput experimental validation based on target-specific screening criteria. The specific selection strategy and quantity allocation are as follows: For the "maximum (max)" target, the 1500 sequences with the highest predicted activity were selected first to obtain promoters with ultra-high activity; for the "minimum (min)" target, to ensure that the sequence activity could still be detected by STARR-seq without being completely drowned out by background noise, 250 sequences with relatively high predicted activity were selected; for the remaining 7 specific numerical targets, the 500 sequences closest to their set target value were selected from each target (a total of 3500 sequences for the 7 targets). Figure 4 ).
[0058] Fourth, 5250 promoters were selected from the promoters generated in step three under 10,000 iterations for 9 objectives for STARR-seq experiments. The experiments verified the activity of 3763 generated promoters, of which 30 generated promoters had activity higher than the maximum activity of natural promoters. After comprehensive consideration, the 10 promoters with the highest activity were selected to build a library.
[0059] The STARR-seq experimental procedure is as follows:
[0060] 1) Library Design and Construction:
[0061] In addition to the 5250 generated promoters, 750 natural promoters were selected for comparison, resulting in a total of 6,000 candidate 170 bp promoter sequences (e.g., ...). Figure 4 (As shown). To facilitate cloning, a 15 bp universal sequence was added to the beginning and end of each sequence, extending its total length to 200 bp. These 200 bp candidate sequences were synthesized using a microarray (TwistBioscience). The pSPO-01 vector was constructed using the Gibson assembly method, replacing the ccdB sequence in the pSTEM02 backbone with a candidate promoter sequence and ligating the luciferase gene (fLuc).
[0062] To link the barcodes to their corresponding promoters, a 67 bp 5' untranslated region (UTR) was first isolated from the maize histone H3.2 gene (Zm00001d041672) and located upstream of a 3 bp ATG start codon and a 15 bp barcode sequence. These sequences were directly synthesized using primer design (Sangon Biotech). As a negative control, the 35S minimal promoter was surrounded by the same 67 bp maize histone H3.2 UTR and paired with a different 12 bp barcode (CTACCGGCCCTA), using the same UTR-ATG-barcode design. To associate the barcodes with their respective promoters, PCR amplification was performed using primers containing the barcode downstream and primers containing the 15 bp universal sequence upstream, typically for 20 to 25 cycles. The experimental library (containing candidate promoter sequences) and the negative control library (containing the 35S minimal promoter) were then constructed using the Gibson assembly method. The assembled DNA fragments were directly transformed into E. coli to generate plasmids, achieving a library coverage of approximately 50×. Colonies were then collected and plasmid DNA was extracted to generate the final promoter-barcode library.
[0063] 2) Preparation and transformation of maize protoplasts:
[0064] Maize seeds (Z. mays L. B73 variety) were germinated under light for 4 days, followed by 9-10 days of growth in soil at 25°C in darkness. Protoplast isolation and transformation were performed according to the previously established PEG4000-Ca² method. + The transfection was performed using a mediated transformation method. The candidate promoter library and the negative control library were mixed at a ratio of 1000:1, resulting in a total of 500 μg of mixed plasmid library. Five reactions were used in each transfection experiment, with a total of 50 μg of mixed plasmid added to cover the entire library. Transfected protoplasts were incubated at 25°C in the dark for 12–16 hours. After incubation, protoplasts were collected by centrifugation at 100 g for 3 minutes for subsequent analysis.
[0065] 3) STARR-seq experiments and Illumina sequencing:
[0066] Approximately 50 μg of plasmid was used to transform 5 × 10⁶ cells / year. 6Each experiment consisted of three independent biological replicates. Protoplasts were harvested after incubation at 25°C in the dark for 12–16 hours and immediately resuspended in 1 mL TRIzol. Total RNA was extracted according to the manufacturer's instructions (Vazyme, China). Reverse transcription was then performed using primers for the pSPO-01 plasmid (SPO-XSRT1) to prepare cDNA (TransGen, China). Barcode-containing sequences were amplified from the cDNA library or the corresponding plasmid library. Finally, these amplified products were sequenced using next-generation sequencing (NGS, Illumina HiSeq, PE150) and used for subsequent analysis.
[0067] 4) STARR-seq data analysis:
[0068] First, use Trimmomatic v.0.39 (Bolger, AM, M. Lohse, and B. Usadel). Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics, 2014, 30(15): p. 2114-20.) performed quality control filtering on the raw data input for cDNA and plasmids, with all parameters set to default. Subsequently, Bowtie v.2.5.2 (Langmead, B. and SL Salzberg, Fast gapped-read alignment with Bowtie 2. (Nat Methods, 2012.9(4): p. 357-9.) Plasmid input data was aligned to candidate promoter reference sequences to establish the correspondence between promoters and barcodes. Next, barcodes corresponding to multiple promoters were filtered. The enrichment of this barcode was defined as the ratio of the barcode count in the cDNA to the corresponding barcode count in the plasmid input. The final enrichment of each promoter was defined as the median enrichment of all associated barcodes for that promoter. To facilitate comparison of the relative activities of different promoters, promoter activity was expressed as the log2 ratio of each promoter enrichment to the 35S minimum promoter enrichment. The final analysis results were calculated based on the average promoter intensity from three independent replicate experiments.
[0069] The results of the STARR-seq experiment are shown as follows: Figure 5 As shown.
[0070] 4) Experimental process and results of Luc activity detection:
[0071] A total of 3,763 generative promoters were analyzed. Among them, 30 generative promoters had activities exceeding the maximum value of natural promoters. Taking all factors into consideration, the 10 promoters with the highest activities were selected for library construction, which are the generative promoters of the nucleotide sequences shown in SEQ ID No. 1 to SEQ ID No. 7.
[0072] To evaluate the activity of 10 highly active synthetic promoters, a dual-luciferase reporter gene assay was performed on maize protoplasts. Briefly, the experiment used mesophyll protoplasts isolated from leaves of 10-day-old etiolated B73 maize seedlings. Reporter plasmids were introduced into the protoplasts via polyethylene glycol-mediated transformation and cultured in the dark for 16 hours. Subsequently, luciferase activity was measured using a dual-luciferase reporter gene assay system, with nanoluciferase (nLuc) activity used as an internal control. The activity assay for each promoter included four biological replicates, with two technical replicates per biological replicate. Finally, the relative LUC activity was calculated by normalizing the firefly luciferase activity to the nLuc internal control activity.
[0073] Depend on Figure 6 The comparison of STARR-seq and luciferase activity assay results shows that ( Figure 6 In the figure, the vertical axis shows the result after normalizing the activity of the 35s promoter. The vertical axis 0 indicates that the activity of the promoter is comparable to that of 35s, the vertical axis 2 indicates that the activity of the promoter is four times that of 35s, and so on. The 10 ultra-high activity promoters screened in this application all have higher activities than 35s, and the activities of the sp358 promoter and sp1482 promoter are even higher than those of the UBI natural promoter.
[0074] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A library of generated ultra-highly active plant core promoters, characterized in that, The promoter library includes promoters with nucleotide sequences shown in SEQ ID No. 1-SEQ ID No.
7.
2. The plant core promoter library according to claim 1, characterized in that, The plant in question is corn.
3. The use of the plant core promoter library according to claim 1 or 2, characterized in that, It is used for plant core promoters, which operatively link the plant core promoter to a target gene to regulate the expression of the target gene; the target gene is the firefly luciferase reporter gene fLuc.