Saccharomycetes core promoter sequence generation method based on potential diffusion model

Through the yeast core promoter sequence generation method based on the potential diffusion model, the problems of instability in training, low quality of generation samples and high calculation costs in the prior art are solved, and high quality and low-cost promoter sequence generation are achieved.

CN120048352AInactive Publication Date: 2025-05-27YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510510719.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing promoter sequence generation methods have problems such as instability in training, low quality of generation samples, and high computational cost.

Method used

The yeast core promoter sequence generation method based on the potential diffusion model is adopted. By collecting and preprocessing data, a potential diffusion model is constructed, and the model parameters are adjusted using optimization algorithms to generate new promoter sequences.

Benefits of technology

The calculation cost is significantly reduced, and the generation process retains the high-quality generation ability of the diffusion model well, solving the problems of high computational cost, slow generation speed and poor generation samples in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048352A_ABST
    Figure CN120048352A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating a yeast core promoter sequence based on a potential diffusion model. The problems that in the prior art, a promoter sequence generation method is unstable in training, the quality of generated samples is low, and the calculation cost is high are solved. The method comprises the following steps: S1, collecting and preprocessing data; s2, sequence coding; s3, constructing a potential diffusion model; s4, performing model training; s5, generating a new promoter sequence; and S6, comparing the generated sequence with an original natural sequence. The method has the advantages that a potential diffusion model based on deep learning is trained by using an existing saccharomycetes core promoter sequence data set, so that the model learns and has the capability of generating a new saccharomycetes core promoter, high-quality and diversified saccharomycetes core promoter sequences are generated, and the defects that a traditional method is high in calculation cost and high in calculation efficiency are overcome. The generation speed is slow, and the quality of generated samples is poor. The biological experiment cost of the saccharomycetes promoter is reduced, the research process is accelerated, and the industrial fermentation optimization of the saccharomycetes is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of bioinformatics and computer science technology, and in particular to a method for generating a yeast core promoter sequence based on a potential diffusion model. Background Art

[0002] Most of the existing methods for generating promoter sequences are based on technologies such as Generative Adversarial Network (GAN), Variational Autoencoder (VAE) and Diffusion Model. Among them, GAN is a deep learning model consisting of two neural networks: Generator and Discriminator. These two networks compete and collaborate with each other through adversarial learning to generate promoter sequence data; VAE is a generative model that combines the advantages of probabilistic graphical models and deep learning, and is mainly used to generate promoter sequence data and learn the potential representation of promoter sequence data; Diffusion Model is a generative model based on probability theory, which generates promoter sequence data samples by simulating the diffusion process of data in time.

[0003] Although existing methods are widely used in promoter sequence generation, they still have some limitations. Although GAN can generate high-quality samples, its training is unstable and prone to mode collapse; the quality of samples generated by VAE is usually low because the distribution assumption of its latent space may be oversimplified; the traditional diffusion model generates high quality but has a high computational cost.

[0004] In summary, the currently used promoter sequence generation models have problems such as unstable training, low quality of generated samples, and high computational cost. Summary of the invention

[0005] The purpose of the present invention is to provide a method for generating a yeast core promoter sequence based on a potential diffusion model in order to solve the above problems.

[0006] To achieve the above object, the present invention adopts the following technical scheme: a method for generating a yeast core promoter sequence based on a potential diffusion model, the method comprising the following steps: S1. Collect and preprocess the public datasets of yeast core promoter sequences and their biological activity evaluation indicators; S2. Perform one-hot encoding on the preprocessed sequence data; S3, construct potential diffusion model; S4, using the training set to train the potential diffusion model, and using the optimization algorithm to adjust the model parameters; S5, using the trained latent diffusion model to randomly generate latent space data that conforms to the Gaussian distribution, and inputting the latent space data as input data into the diffusion process and decoder of the latent space of the model to generate a new promoter sequence; S6. Determine the isodistribution and diversity of the generated sequences with natural sequences.

[0007] In the above-mentioned method for generating a yeast core promoter sequence based on a potential diffusion model, step S1 specifically includes the following steps: S11. Treat each set of data as an independent data unit and perform data preprocessing on the collected data. The sequence data after preprocessing must ensure that the sequence length is consistent; S12. Select data units with higher biological activity of promoter sequences to form a data set, and format the data set into CSV format to ensure consistency of data format for subsequent sequence encoding and model training.

[0008] In step S1, a public data set of yeast core promoter sequences and evaluation indexes of their biological activities is downloaded from a public gene expression comprehensive database, each public data set contains a column of yeast core promoter sequences and corresponding biological activity values, each public data set is marked as a group of units, and the data set has a total of n groups of units; In step S1, the preprocessed sequence data consists of four bases: A: adenine; T: thymine; C: cytosine; G: guanine; The one-hot encoding of the preprocessed sequence data is: The encoding of A is [1,0,0,0], The encoding of T is [0,1,0,0], The encoding of C is [0,0,1,0], The encoding of G is [0,0,0,1].

[0009] In step S3, the construction of the potential diffusion model includes three core parts: ① Encoder, which compresses high-dimensional sequence encoding data into a low-dimensional latent space; ② Latent Diffusion Process in latent space: forward diffusion and reverse generation of the diffusion model in latent data space; ③Decoder: maps the potential data representation to the original data space to generate the final high-quality sequence data.

[0010] Step S3 specifically includes the following steps: S31. Construct the encoder Encoder in the autoencoder Autoencoder; S32, construct the decoder Decoder in the automatic encoder; S33. Design the loss function of the autoencoder. S34, Constructing the Latent Diffusion Process of the Latent Space; S35. Design the total loss function of the latent diffusion model; The total loss function of the latent diffusion model is the weighted sum of the loss functions of each module: , in, , , is the weight of each loss item.

[0011] In step S33, the loss function of the autoencoder consists of two parts: S331, reconstruction loss: reconstruction loss is the core loss of Autoencoder, which is used to measure the difference between input data and reconstructed data; the mean square error is used to calculate the reconstruction loss, and its calculation formula is: , in: is the input data, is the reconstructed data, N is the number of data points; S332, Regularization loss: In order to prevent overfitting and improve the generalization ability of the latent space, the model adds regularization loss. The regularization loss method used is KL divergence Kullback-Leibler Divergence. KL divergence is used to constrain the potential distribution to be close to the prior distribution. The calculation formula of KL divergence is: , in: is the latent distribution of the encoder output, is the prior distribution; The total loss function of the autoencoder is the weighted sum of the above loss terms: in: , is the weight of each loss item, which is adjusted according to task requirements.

[0012] In step S34, a noise schedule is defined before the forward process begins. ,in: is the noise intensity at time step t, which usually satisfies ( ), at time step t, the latent representation From the previous step Adding noise gives: ,in, is standard Gaussian noise; by accumulating noise scheduling, from the initial potential representation Compute the latent representation at any time step t : , in: , , is standard Gaussian noise.

[0013] The formula for the reverse step is: At time step t, from recover : ; in: is the initial latent representation, is the noise added in the forward process, is the noise potential representation at time step t.

[0014] Step S4 specifically includes the following steps: S41. Dataset preparation: Prepare the promoter sequence data after one-hot encoding; S42, selecting an optimization algorithm: using an optimization algorithm to adjust parameters of the model; S43, set learning rate and batch size: select an appropriate learning rate to control the step size of each parameter update; S44. Model tuning and hyperparameter optimization: After the initial training is completed, the hyperparameters are further optimized through the grid search method.

[0015] Step S5 specifically includes the following steps: S51, using the trained latent diffusion model: in step S4, the fully trained latent diffusion model has the ability to generate the yeast core promoter sequence. At this time, the model is applied to the generation of the yeast core promoter sequence; S52. Determine the input of the latent diffusion model: In order to generate a new promoter sequence under sufficiently random conditions, the input in the generation process is a random data whose distribution conforms to the Gaussian distribution and whose size conforms to the latent space latent vector set by the model. The random vector z is input into the latent diffusion process and then passes through the decoder to finally obtain high-quality sequence data.

[0016] In step S6, known biological knowledge is used to determine the distribution and diversity of the generated sequence and the natural sequence by calculating the base ratio distribution of the generated sequence and the original natural sequence.

[0017] Compared with the prior art, the advantages of the present invention are: significantly reducing the computational cost, while the generation process well retains the high-quality generation capability of the diffusion model. This model method well overcomes the problems of high computational cost, slow generation speed, and poor quality of generated samples in traditional methods; at the same time, using the potential diffusion model to generate yeast core promoter sequences can reduce the cost of yeast promoter sequence biological experiments, accelerate the research process, and promote the optimization of industrial fermentation of yeast. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flow chart of the method of the present invention; Figure 2 is a schematic diagram of the architecture of the potential diffusion model in the present invention; Figure 3 It is a base probability comparison result diagram of the generated sequence provided in the present invention and the original yeast core promoter sequence. DETAILED DESCRIPTION

[0019] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] like Figure 1-3 As shown, a method for generating a yeast core promoter sequence based on a potential diffusion model comprises the following steps: S1. Collect and preprocess the public datasets of yeast core promoter sequences and their biological activity evaluation indicators; S2. Perform one-hot encoding on the preprocessed sequence data; S3, construct potential diffusion model; S4, using the training set to train the potential diffusion model, and using the optimization algorithm to adjust the model parameters; S5, using the trained latent diffusion model to randomly generate latent space data that conforms to the Gaussian distribution, and inputting the latent space data as input data into the diffusion process and decoder of the latent space of the model to generate a new promoter sequence; S6. Determine the isodistribution and diversity of the generated sequences with natural sequences.

[0021] Step S1 specifically includes the following steps: S11. Treat each set of data as an independent data unit and perform data preprocessing on the collected data. The sequence data after preprocessing must ensure that the sequence length is consistent; S12. Select data units with higher biological activity of promoter sequences to form a data set, and format the data set into CSV format to ensure consistency of data format for subsequent sequence encoding and model training.

[0022] In step S1, a public data set of yeast core promoter sequences and their biological activity evaluation indicators is downloaded from a public gene expression comprehensive database. Each public data set contains a column of yeast core promoter sequences and corresponding biological activity values. Each public data set is marked as a group of units, and the data set has a total of n groups of units.

[0023] The yeast core promoter sequence is a DNA sequence, and all sequences are preprocessed into standard CSV files. The length L of the promoter sequence is 118. In this example, the data set contains n=7536 groups of unit data.

[0024] In step S1, the preprocessed sequence data consists of four bases: A: adenine; T: thymine; C: cytosine; G: guanine; The one-hot encoding of the preprocessed sequence data is: The encoding of A is [1,0,0,0], The encoding of T is [0,1,0,0], The encoding of C is [0,0,1,0], The encoding of G is [0,0,0,1].

[0025] One-hot encoding of preprocessed sequence data is a method for converting DNA bases into binary vectors, which is commonly used in machine learning and deep learning tasks. One-hot encoding represents each base as a four-dimensional binary vector. The one-hot encoding of the present invention has the characteristics of uniqueness, sparsity, and no size relationship. It is suitable as a sequence encoding method for this experiment.

[0026] In step S3, the construction of the potential diffusion model includes three core parts: ① Encoder, which compresses high-dimensional sequence encoding data into a low-dimensional latent space; ② Latent Diffusion Process in latent space: forward diffusion and reverse generation of the diffusion model in latent data space; The dimensionality of the latent space is lower than the original data space, thus reducing computational cost; ③Decoder: maps the potential data representation to the original data space to generate the final high-quality sequence data.

[0027] Step S3 specifically includes the following steps: S31. Construct the encoder Encoder in the autoencoder Autoencoder; Autoencoders are used to map high-dimensional data to a low-dimensional latent space and reconstruct data from the latent space when generated. Encoder: compresses the input sequence encoding data x to the latent space z. The encoder consists of three two-dimensional convolutional layers, a Flatten layer, and a fully connected layer. The convolutional layer is used for feature extraction. The Flatten layer flattens the multi-dimensional feature map into a one-dimensional vector. The fully connected layer maps the flattened vector to a vector whose length is twice the latent dimension. S32, construct the decoder Decoder in the automatic encoder; Decoder: Reconstructs the original sequence data from the latent space z.

[0028] The decoder consists of a fully connected layer, an Unflatten layer, and three transposed convolutional layers; the fully connected layer maps the potential vector to an initial feature map shape; the Unflatten layer reshapes the one-dimensional vector into a multi-dimensional feature map; the transposed convolutional layer performs an upsampling operation to restore the feature map to single-channel data.

[0029] S33. Design the loss function of the autoencoder. S34. Construct a latent diffusion process in the latent space; in the latent diffusion model, the diffusion process is the core part of the model, which generates data by simulating the gradual addition and removal of noise in the latent space. The core idea of ​​the diffusion process is to transform the data (the latent representation of sequence encoding in this invention) from the original distribution to the Gaussian noise distribution by gradually adding noise, and then train a neural network to learn how to reverse this process, so as to recover the original data from the noise.

[0030] The diffusion process is divided into two stages: forward process and reverse process; The forward process gradually adds noise to transform data into noise and the reverse process gradually removes noise to recover data from noise; The reverse process is a learnable Markov chain that gradually recovers the original latent representation from the noise. The goal of the reverse process is to train a neural network , predict the noise added in the forward process , thereby gradually removing the noise; S35. Design the total loss function of the latent diffusion model; The total loss function of the latent diffusion model is the weighted sum of the loss functions of each module: , in, , , is the weight of each loss term.

[0031] In step S33, the loss function of the autoencoder consists of two parts: S331, reconstruction loss: reconstruction loss is the core loss of Autoencoder, which is used to measure the difference between input data and reconstructed data; the mean square error is used to calculate the reconstruction loss, and its calculation formula is: , in: is the input data, is the reconstructed data, N is the number of data points; S332, Regularization loss: In order to prevent overfitting and improve the generalization ability of the latent space, the model adds regularization loss. The regularization loss method used is KL divergence Kullback-Leibler Divergence. KL divergence is used to constrain the potential distribution to be close to the prior distribution. The calculation formula of KL divergence is: , in: is the latent distribution of the encoder output, is the prior distribution; The total loss function of the autoencoder is the weighted sum of the above loss terms: in , is the weight of each loss item, which is adjusted according to task requirements.

[0032] In step S34, a noise schedule is defined before the forward process begins. ,in: is the noise intensity at time step t, which usually satisfies ( ), at time step t, the latent representation From the previous step Adding noise gives: ,in, is standard Gaussian noise; by accumulating noise scheduling, from the initial potential representation Compute the latent representation at any time step t : Q , in: , , is standard Gaussian noise.

[0033] The formula for the reverse step is: At time step t, from recover : ; in: is the initial latent representation, is the noise added in the forward process, is the noise potential representation at time step t.

[0034] Noise Prediction Network It is the core component of the diffusion process, and its goal is to predict the noise added in the forward process. In this invention, the U-net network is used to predict the noise in the diffusion model. The multi-scale feature extraction of the U-net network can capture global and local information at the same time through the hierarchical structure of the encoder and decoder. And its jump connection can retain low-level detail information and improve the quality of the generated samples.

[0035] Step S4 specifically includes the following steps: S41. Dataset preparation: Prepare the promoter sequence data after one-hot encoding; S42, selecting an optimization algorithm: using an optimization algorithm to adjust parameters of the model; In order to effectively train the metric learning model, commonly used optimization algorithms include Adam optimizer and stochastic gradient descent (SGD). In the present invention, it is recommended to use Adam optimizer because it has good convergence and can automatically adjust the learning rate to adapt to the update speed of different parameters. During the training process, the optimization algorithm will gradually adjust the model weights according to the gradient information calculated by the total loss function L"Total" of the potential diffusion model to minimize the loss function value.

[0036] S43. Set the learning rate and batch size: Select an appropriate learning rate to control the step size of each parameter update; in this case, it is set to 1e-6. A learning rate that is too high will lead to an unstable training process, and a learning rate that is too small will lead to slow convergence. In addition, the batch size is also an important hyperparameter that affects the training process. In this case, it is set to 16. A suitable batch size can balance the training time and the convergence speed of the model.

[0037] S44. Model tuning and hyperparameter optimization: After the initial training is completed, the hyperparameters are further optimized through the grid search method.

[0038] Step S5 specifically includes the following steps: S51, using the trained latent diffusion model: in step S4, the fully trained latent diffusion model has the ability to generate the yeast core promoter sequence. At this time, the model is applied to the generation of the yeast core promoter sequence; S52. Determine the input of the latent diffusion model: In order to generate a new promoter sequence under sufficiently random conditions, the input in the generation process is a random data whose distribution conforms to the Gaussian distribution and whose size conforms to the latent space latent vector set by the model. The random vector z is input into the latent diffusion process and then passes through the decoder to finally obtain high-quality sequence data.

[0039] In step S6, using known biological knowledge, the distribution and diversity of the generated sequence and the natural sequence are determined by calculating the base ratio distribution of the generated sequence and the original natural sequence; Specifically include: the probability of DNA sequence bases. The probability of DNA sequence bases is of great significance in biology and bioinformatics, which reflects the occurrence law and distribution characteristics of bases A, T, C, and G in DNA sequences.

[0040] like Figure 3 As shown, the results show that the promoter sequence generated by the method in this example has the same distribution relationship as the original yeast core promoter sequence, and has diverse differences compared with the original sequence, and the functional region is also well preserved. The result diagram shows that the potential diffusion model can generate high-quality and diverse yeast core promoter sequences.

[0041] The present invention chooses to use this method to compare the association and difference between the generated sequence and the original natural sequence. Base probability can describe the basic composition characteristics of the DNA sequence. For example, GC content, that is, the ratio of G and C bases, is an important feature of the DNA sequence, which is usually related to the coding region, thermal stability and evolutionary selection of the gene. Base preference: Certain bases or base combinations appear more frequently in specific regions, such as promoter sequences and splicing sites, reflecting the characteristics of the functional region.

[0042] In summary, the principle of this embodiment is to train a deep learning-based latent diffusion model with an existing yeast core promoter sequence dataset, so that the model learns and has the ability to generate new yeast core promoter sequences, and generates high-quality, diversified yeast core promoter sequences. The latent diffusion model is a generative model that combines the diffusion model and the latent space. Its core idea is to migrate the diffusion process to the latent space and diffuse in the latent space, which significantly reduces the computational cost. At the same time, the generation process well retains the high-quality generation ability of the diffusion model.

[0043] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for generating a yeast core promoter sequence based on a potential diffusion model, characterized in that: The method comprises the following steps: S1. Collect and preprocess the public datasets of yeast core promoter sequences and their biological activity evaluation indicators; S2. Perform one-hot encoding on the preprocessed sequence data; S3, construct potential diffusion model; S4, using the training set to train the potential diffusion model, and using the optimization algorithm to adjust the model parameters; S5, using the trained latent diffusion model to randomly generate latent space data that conforms to the Gaussian distribution, and inputting the latent space data as input data into the diffusion process and decoder of the latent space of the model to generate a new promoter sequence; S6. Determine the isodistribution and diversity of the generated sequences with natural sequences.

2. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11. Treat each set of data as an independent data unit and perform data preprocessing on the collected data. The sequence data after preprocessing must ensure that the sequence length is consistent; S12. Select data units with higher biological activity of promoter sequences to form a data set, and format the data set into CSV format to ensure consistency of data format for subsequent sequence encoding and model training.

3. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 2, characterized in that: In step S1, the public data sets of the yeast core promoter sequence and the evaluation index of its biological activity are downloaded from the public gene expression comprehensive database, each group of public data sets contains a column of yeast core promoter sequences and corresponding biological activity values, each group of public data sets is marked as a group of units, and the data sets have a total of n groups of units; In step S1, the preprocessed sequence data consists of four bases: A: adenine; T: thymine; C: cytosine; G: guanine; The one-hot encoding of the preprocessed sequence data is: The encoding of A is [1,0,0,0], The encoding of T is [0,1,0,0], The encoding of C is [0,0,1,0], The encoding of G is [0,0,0,1].

4. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 1, characterized in that: In step S3, the potential diffusion model comprises three core parts: ① Encoder, which compresses high-dimensional sequence encoding data into a low-dimensional latent space; ② Latent Diffusion Process in latent space: forward diffusion and reverse generation of the diffusion model in latent data space; ③Decoder: maps the potential data representation to the original data space to generate the final high-quality sequence data.

5. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 4, characterized in that: Step S3 specifically includes the following steps: S31. Construct the encoder Encoder in the autoencoder Autoencoder; S32, construct the decoder Decoder in the automatic encoder; S33. Design the loss function of the autoencoder. S34, Constructing the Latent Diffusion Process of the Latent Space; The diffusion process is divided into two stages: forward process and reverse process; S35. Design the total loss function of the latent diffusion model; The total loss function of the latent diffusion model is the weighted sum of the loss functions of each module: , in, , , is the weight of each loss term.

6. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 5, characterized in that: In step S33, the loss function of the autoencoder consists of two parts: S331, reconstruction loss: reconstruction loss is the core loss of Autoencoder, which is used to measure the difference between input data and reconstructed data; the mean square error is used to calculate the reconstruction loss, and its calculation formula is: , in: is the input data, is the reconstructed data, N is the number of data points; S332, Regularization loss: In order to prevent overfitting and improve the generalization ability of the latent space, the model adds regularization loss. The regularization loss method used is KL divergence Kullback-Leibler Divergence. KL divergence is used to constrain the potential distribution to be close to the prior distribution. The calculation formula of KL divergence is: , in: is the latent distribution of the encoder output, is the prior distribution; The total loss function of the autoencoder is the weighted sum of the above loss terms: , in: , is the weight of each loss item, which is adjusted according to task requirements.

7. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 4, characterized in that: In step S34, a noise schedule is defined before the forward process begins. ,in: is the noise intensity at time step t, which usually satisfies ( ), at time step t, the latent representation From the previous step Adding noise gives: ,in, is standard Gaussian noise; by accumulating noise scheduling, from the initial potential representation Compute the latent representation at any time step t : , in: , , is standard Gaussian noise; The formula for the reverse step is: At time step t, from recover : ; in: is the initial latent representation, is the noise added in the forward process, is the noise potential representation at time step t.

8. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 3, characterized in that: Step S4 specifically includes the following steps: S41. Dataset preparation: Prepare the promoter sequence data after one-hot encoding; S42, selecting an optimization algorithm: using an optimization algorithm to adjust parameters of the model; S43, set learning rate and batch size: select an appropriate learning rate to control the step size of each parameter update; S44. Model tuning and hyperparameter optimization: After the initial training is completed, the hyperparameters are further optimized through the grid search method.

9. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 8, characterized in that: Step S5 specifically includes the following steps: S51, using the trained latent diffusion model: in step S4, the fully trained latent diffusion model has the ability to generate the yeast core promoter sequence. At this time, the model is applied to the generation of the yeast core promoter sequence; S52. Determine the input of the latent diffusion model: In order to generate a new promoter sequence under sufficiently random conditions, the input in the generation process is a random data whose distribution conforms to the Gaussian distribution and whose size conforms to the latent space latent vector set by the model. The random vector z is input into the latent diffusion process and then passes through the decoder to finally obtain high-quality sequence data.

10. The method for generating a yeast core promoter sequence based on a potential diffusion model according to claim 9, characterized in that: In step S6, known biological knowledge is used to determine the distribution and diversity of the generated sequence and the natural sequence by calculating the base ratio distribution of the generated sequence and the original natural sequence.

Citation Information

Patent Citations

  • Method for generating non-natural promoter based on diffusion model

    CN116978462A

  • Promoter activity optimization method based on de-noising diffusion model

    CN117789829A

  • Escherichia coli DNA promoter generation method based on diffusion model

    CN117953972A

  • Multi-modal medical image conversion method and system based on potential diffusion model

    CN118172237A

  • Dense fish school occlusion image restoration method based on diffusion model

    CN119784620A