Method for designing product synthesis optimization strain by enzyme constraint model fusion evolutionary algorithm
By combining genetic algorithms and enzyme constraint models, the metabolic engineering of microbial strains was optimized, solving the problem of multi-target combination prediction, achieving efficient strain design and improving the yield of target products, and demonstrating the potential of heuristic algorithms in strain optimization design.
Patent Information
- Application Number
- CN202411276176.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-09-12
AI Technical Summary
Existing technologies lack effective systematic design tools and methods to optimize the metabolic engineering of microbial strains, especially in the prediction and analysis of multi-target combinations. Furthermore, the application of enzyme constraint models is limited, resulting in low strain design efficiency and difficulty in accurately identifying metabolic engineering targets under time and resource constraints.
By combining genetic algorithms and enzyme constraint models, and through simulating the metabolic phenotype of wild-type strains, gene dimensionality reduction analysis and annotation, and genetic algorithm search, a method for optimizing strains for product synthesis is designed. This includes solving for maximizing the specific growth rate of the strain, dimensionality reduction of enzyme synthesis-related genes, genetic algorithm combined target search, and optimization of gene editing strategies.
It improves the accuracy and efficiency of strain design, reduces the search space, balances the algorithm's global search capability and convergence speed, significantly improves the yield of target products, provides a non-intuitive gene editing strategy, and enhances the ability to optimize metabolic engineering.
Smart Images

Figure CN119418752B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of combining biology and computers, and in particular to a method for designing optimal strains for product synthesis by fusing enzyme constraint models with evolutionary algorithms. Background Technology
[0002] The rapid development of computer and biotechnology has greatly promoted the application of microorganisms in many fields, including the production of biofuels, antibiotics, food and its additives, biopesticides, and various processing enzymes. Bioproduction pathways based on microbial fermentation technology are gradually replacing traditional chemical synthesis processes, marking a shift towards green and sustainable production methods. However, to win in the fierce competition with traditional petroleum-based production methods, the key lies in developing high-performance microbial cell factories. For industrial production, developing cost-effective industrial strains is crucial. Under the evolutionary pressure of natural selection, the metabolic phenotype of microorganisms in their natural habitats often maximizes cell growth rather than the synthesis of a particular product. This requires metabolic modification of strains to obtain the desired cell phenotype. Therefore, metabolic engineering of industrial strains—that is, optimizing the yield and efficiency of target products through gene editing—has become an important research direction in the field of bioengineering. The main challenges facing this process include the complexity of the microbial metabolic network itself and the incomplete understanding of the mechanisms regulating cellular metabolism.
[0003] Currently, most successful cases of metabolic engineering rely on intuitive, qualitative design principles, lacking rational, systematic design and analysis tools. Furthermore, most strain design algorithms are used to predict single targets, lacking prediction and analysis of multi-target combinations. The concept of constraint-based modeling has driven the development of reliable quantitative metabolic models, providing a new perspective and analytical platform for metabolic engineering. Among them, the Genome-Scale Metabolic Model (GSMM), as a comprehensive bioinformatics tool, can comprehensively describe all known metabolic pathways within cells and predict how cells adjust their metabolic processes under different environments. However, GSMM struggles to accurately simulate metabolic transitions occurring at high growth rates, such as bacterial overflow metabolism, the Crabtree effect in Saccharomyces cerevisiae, and the Warburg effect in human cancer cells. By integrating enzyme concentration and enzyme kinetic parameters into stoichiometric models to construct enzyme-constrained models, the predictive power of GSMM can be enhanced. The GECKO (GEM with Enzymatic Constraints using Kinetic and Omics data) approach was used to integrate proteomics data. This approach adds an upper bound to each metabolic flux by treating enzymes as part of the reaction—enzyme abundance and enzyme turnover (k). cat The product of ) . While significantly reducing the flux variability of traditional GSMM, it also reveals the use and distribution of enzymes across different reactions and metabolic pathways.
[0004] The sheer number of metabolites and reactions involved in GSMM and enzyme-constrained models leads to a large dimensionality in the solution space they define. A commonly used method for determining metabolic phenotypes is flux balance analysis (FBA). Furthermore, to simulate the metabolic flux distribution of mutant strains resulting from gene intervention, the Minimization of Metabolic Adjustment (MOMA) method has been proposed. The former primarily relies on linear programming techniques, with common methods including the simplex method. FBA is based on two fundamental assumptions: the steady-state assumption and the mass conservation assumption. The steady-state assumption states that the strain's growth has reached a stable state; the mass conservation assumption states that in a steady-state metabolic network, the total production and consumption of each metabolite should be stoichiometrically zero. The latter provides an approximate solution for handling suboptimal states of mutant strains through mathematical methods: by employing the distance minimization principle in the flux space, the problem is transformed into a quadratic programming problem. Flux variability analysis (FVA) is mainly used to determine the maximum and minimum flux values of each metabolic reaction under specific growth conditions and environments. This algorithm is used to understand the robustness of metabolic networks and their adaptability to environmental changes. The Forced Objective Flux (FSEOF) algorithm adds additional constraints to the pathway leading to the synthesis of the target product during flux balance analysis, forcibly increasing the flux of that pathway step by step, and then scanning the flux of all reactions. ecFSEOF is an improvement on this, adapted to an enzyme constraint model. This algorithm is implemented by solving a series of flux balance analysis problems with fixed biomass synthesis rates, scoring all reactions and simultaneously providing suggestions for overexpression, knockdown, and deletion of relevant genes.
[0005] Due to limitations in biometrics and the complexity of microbial metabolic regulation, it is currently impossible to obtain complete metabolic flux phenotypic data of strains under specific growth conditions. However, optimal strain design requires reference yields of target products, and the MOMA method also requires a set of reference flux vectors for solution. Therefore, it is essential to use models to predict the metabolic phenotype of wild-type strains with high accuracy.
[0006] Currently, various strain design algorithms and integrated model building and optimization algorithms have been developed into software toolkits, and there are also cases of integrating advanced algorithms such as intelligent algorithms, reinforcement learning, and machine learning into strain design. However, the development of algorithms for enzyme constraint models is still relatively limited, and accurately identifying metabolic engineering targets under time and resource constraints remains a major challenge. Summary of the Invention
[0007] The purpose of this invention is to combine genetic algorithms and enzyme-constrained models to obtain combinatorial gene targets and editing directions for modified microorganisms. This is achieved through steps such as simulating the metabolic phenotype of wild-type strains, dimensionality reduction analysis and annotation of genes in the model, and genetic algorithm search, thereby solving the problems mentioned in the background art.
[0008] To achieve the above-mentioned objectives, this invention provides a method for designing optimal strains for product synthesis using an enzyme constraint model fused with an evolutionary algorithm, comprising the following steps:
[0009] Step S1: Solve the enzyme constraint model with the objective function of maximizing the strain specific growth rate to obtain the simulated metabolic flux of the wild-type strain.
[0010] Step S2 involves solving a series of flux balance analysis problems with fixed biomass synthesis rates to perform dimensionality reduction analysis and labeling of genes directly related to enzyme synthesis within the model.
[0011] Step S3: Predict the yield of single-target editing, modify the model based on the results of dimensionality reduction and gene annotation, solve the strain phenotype of single-target mutation and calculate the yield of target product;
[0012] Step S4 involves using a genetic algorithm to search for combined target points after dimensionality reduction, designing and adjusting relevant parameters, and obtaining the optimal gene editing combination strategy.
[0013] Step S5: Repeat the experiment to verify the stability of the algorithm, and perform statistical analysis and biological interpretation on the obtained combinatorial strategies.
[0014] Furthermore, before performing the solution analysis of the enzyme constraint model and the combined optimization of the genetic algorithm, a series of necessary adjustments and verifications were made to the model to ensure that it could accurately simulate the metabolic state of the strain under different environmental conditions. This included releasing the absorption pathways of oxygen and the main carbon and nitrogen sources, and setting the upper bound of all reactions with an infinite upper limit to b. u (b u The model precision is set to Ω (Ω∈[100,2000]) to support subsequent flux scanning, flux variability analysis, and metabolic adjustment minimization calculations. Considering the very small absolute flux of protein synthesis during strain growth, the model precision is set to Ω (Ω∈[100,2000]). -8 10 -10 This increases our ability to understand protein synthesis.
[0015] Furthermore, the flux balance analysis method solves the enzyme constraint model to simulate the metabolic flux of wild-type strains, including the following three methods:
[0016] Method 1: Given a minimum yield limit for the target product, a single-step flux balance analysis is performed with the objective function of maximizing the specific growth rate of the strain.
[0017] Method 2, based on Method 1, restricts the specific growth rate to be no less than its maximum growth rate by λ.
[0018] (λ∈[0.5,0.9]), the FBA is solved with the objective function of minimizing glucose consumption, and the glucose consumption rate is limited to no more than its minimum consumption amount α (α∈[1.1,1.5]) times, and the FBA is solved with the objective function of minimizing the total enzyme consumption;
[0019] Method 3: Without limiting the yield of the target product, perform FBA calculations three times as in Method 2, and then limit the total enzyme consumption to no more than β (β∈[1.1,1.5]) times its minimum value, and use maximizing product yield as the objective function to perform FBA calculations.
[0020] Furthermore, the flux balance analysis used is expressed by the following formula:
[0021] max c T ·v
[0022]
[0023] Here, c is a constant vector consisting of 0s and 1s, and v is the flux vector representing all reactions. The product of the transpose of c and v defines the optimization direction of the model, i.e., maximizing the flux of a certain reaction. S is the stoichiometry matrix of the model. and These represent the minimum and maximum constraint boundaries of the reaction, respectively.
[0024] Further, step 2 requires screening all genes related to enzyme synthesis. The enzyme constraint model within the GECKO framework adds new rows representing enzymes and new columns indicating their usage to the GSMM framework, directly treating enzymes as reaction components to expand the model. Each reaction related to enzyme synthesis has a corresponding gene; these genes are then subjected to dimensionality reduction analysis and annotation.
[0025] First, the biomass synthesis flux is set to its theoretical maximum. The algorithm then explores 16 different biomass synthesis rate conditions that are gradually reduced from this maximum (from 100% of its maximum predicted value to 25%) to investigate the changes in metabolic responses under different biomass synthesis rate conditions. Next, the flux of all reactions is obtained by solving for the FBA with the yield of the target product as the objective function. These changes are quantified through a two-step scoring method; reactions showing significant changes are considered as targets for engineering adjustments. The total flux score (V) is then calculated. scoreThis represents the average ratio of the flux of the reaction to the flux under all forced target flux settings, compared to the flux under the maximum biomass synthesis rate setting. The specific formula is:
[0026]
[0027] in The flux of reaction i is under the condition of maximum biomass synthesis rate, while the molecule in the formula is the sum of the fluxes of the reaction under all simulation conditions, and n is the number of simulations.
[0028] Furthermore, any V greater than A (A∈[1000,5000]) score Values will be truncated to 1000 to avoid infinitely large values. Undetermined values (0 divided by 0) are assigned 1. Based on these flux scores, gene scores are then calculated. As the average reaction flux score for each gene:
[0029]
[0030] In the formula is the score of gene target i, molecule is the sum of the reaction scores of all reactions catalyzed by the gene product, and m is the total number of reactions catalyzed by the gene product.
[0031] according to Genes with a value higher than 1 are suggested as overexpression targets. Between k a to k b The target is between k and k. a The target is to be knocked out. Where k is... a ∈[0.02,0.08], k b ∈[0.2,0.8].
[0032] Furthermore, a representation for gene editing is designed within the model, wherein:
[0033] In enzyme constraint models, gene knockout can be intuitively implemented by setting both the upper and lower bounds of the reaction flux for the corresponding enzyme's synthetic pathway to 0, indicating that the enzyme's synthetic pathway is shut down. For gene overexpression and knockdown, the direction and degree of flux adjustment need to be manually set. Based on the predicted wild-type flux, a reference flux is provided for each enzyme. Overexpression is designed to increase the lower bound of the relevant enzyme's synthetic flux to N times the reference value, while knockdown adjusts the upper bound to M times the reference value. If the reference flux value for the target site to be overexpressed is 0, its lower bound is set to v. low (v low ∈[10 -9 4×10 -9 ]).
[0034] The metabolic adjustment minimization method for solving the metabolic flux of mutant strains defines the distance as:
[0035]
[0036] Here, vector w represents the metabolic flux of the wild-type strain and serves as a reference vector. Vector x represents the metabolic flux distribution of the mutant strain.
[0037] Furthermore, yield predictions were performed for single-target editing, and the edited enzyme constraint model was solved using a metabolic adjustment minimization method to obtain the metabolic phenotype of the mutant strain and calculate the yield of the target product.
[0038] Furthermore, in step S4, a genetic algorithm is used to search for combined target points after dimensionality reduction. During the initialization phase, the following population is generated:
[0039] pop = {X1, X2, ..., X} NP},
[0040] X i =[x i1 ,x i2 ,...,x iL ],
[0041] NP represents the population size, and L represents the dimension of an individual, i.e. the number of genes contained in an individual, where i∈[1,NP].
[0042] It includes two encoding formats, namely:
[0043] In binary encoding, each dimension of the individual packet can only be either "0" or "1". "0" indicates that the gene was not selected by this gene editing strategy, while "1" indicates that it was selected. The editing method for the gene is then obtained based on the annotation results. The individual initialization method is as follows:
[0044]
[0045] oi is a random number between 0 and 1. The value of O in the judgment condition controls the number of times 1 appears in the sequence of an individual. Taking a smaller value of O∈[0.05,0.2] ensures that the gene editing combination strategy represented by the individual only needs to intervene in a small number of target sites, thus ensuring the feasibility of actual experiments.
[0046] Symbolic encoding-based genetic algorithms rely solely on gene dimensionality reduction results to simultaneously find combined target points and their editing directions within a larger search space. In the symbolic encoding case, the individual's encoding sequence contains four values: 0, 1, 2, and 3, representing that the gene was not selected, needs overexpression, knockdown, and deletion operations, respectively. The individual initialization method is as follows:
[0047]
[0048] Where 0 < O1 ≤ 0.04 < O2 ≤ 0.08 < O3 ≤ 1.2. Each gene has an equal probability of being initialized in one of the three editing directions, and a greater probability of not requiring editing.
[0049] In the selection and population update, a roulette wheel selection strategy was adopted. After each individual was evaluated, they were sorted according to their fitness values, and the probability of each individual being selected was calculated according to the following formula.
[0050]
[0051] The individual evaluation function f is solved using a model. This formula calculates the cumulative probability of all individuals being selected, then generates a random number between 0 and 1, and selects the corresponding individuals based on the cumulative probability range the random number falls into. The selected individuals will then be considered as candidates for the crossover operation.
[0052] After the algorithm completes one iteration, the original population and the individuals newly generated by crossover mutation are mixed and sorted, and a portion of the individuals at the top are selected according to the set population size to enter the next search cycle.
[0053] This algorithm uses single-point crossover, with a crossover probability V. c (V c The algorithm determines whether two parent individuals should cross over within the range [0.1, 0.9]. The two parent individuals selected for crossover exchange the coding sequence after a certain gene locus, where the gene locus is randomly chosen.
[0054] In the case of binary encoding, if an individual is to undergo mutation, a gene locus in that individual is randomly selected for mutation: from "0" to "1" or from "1" to "0".
[0055] For symbolic encoding, first determine whether the individual needs to be mutated, then determine the direction of mutation bit by bit: generate a random number for each position of the individual, and mutate according to the interval where the random number is located. If the value exceeds the mutation probability V, the mutation is considered complete. r (V r If the value is in the range [0.01, 0.5], then the point will not change.
[0056]
[0057] Where 0 < M1 ≤ 0.02 < M2 ≤ 0.04 < M3 ≤ 0.06. Flux variability analysis was performed on the enzyme constraint model. The flux variability value for each reaction was calculated according to the formula:
[0058] flux variability i =max flux i -min flux i ,
[0059] Where, max flux i and min flux i It is obtained by changing the solution objective through FBA, and the specific growth rate is limited to a variation range of 1% of the maximum specific growth rate during the solution process.
[0060] Compared with existing technologies, this system and method have the following advantages:
[0061] (1) The metabolic phenotype of wild-type strains with the lowest target product yield was predicted by enzyme constraint model and used as a reference for subsequent design, which solved the problem of lack of experimental data.
[0062] (2) Dimensionality reduction analysis and annotation were performed on genes directly related to enzyme synthesis in the enzyme-constrained metabolic network model, which effectively reduced the search space size and ensured that research resources and time were concentrated on the most promising candidate targets.
[0063] (3) The specific implementation of gene editing in the model was designed, and a genetic algorithm was used to search for gene editing combination strategies, balancing the algorithm's global search capability and convergence speed.
[0064] (4) The target product yield predicted by the combined target editing model is significantly higher than the yield predicted by single target editing.
[0065] (5) The genetic algorithm in symbolic encoding can simultaneously search for and combine target points and their editing directions, and has a higher improvement effect on strains in the model calculation simulation results.
[0066] (6) Genetic algorithms are used to discover non-intuitive gene editing strategies, providing strong algorithmic support for solving complex metabolic engineering optimization problems, and demonstrating the great potential of applying heuristic algorithms to strain optimization design. Attached Figure Description
[0067] Figure 1 A schematic diagram of a method for designing the optimal strain for product synthesis using an enzyme constraint model fused with an evolutionary algorithm.
[0068] Figure 2 Flowchart of the method for designing the optimal strain for product synthesis using an enzyme constraint model fused with an evolutionary algorithm.
[0069] Figure 3 This is a graph showing the variability analysis of the Saccharomyces cerevisiae enzyme constraint model.
[0070] Figure 4 This is a scatter plot of the target product yield after single-target editing in this invention.
[0071] Figure 5 The graph shows the search results of the genetic algorithm in binary and symbolic encoding forms.
[0072] Figure 6 This is a convergence curve of ten repeated experiments in this invention.
[0073] Figure 7 Stacked bar chart to statistically analyze gene editing methods in repeated experiments. Detailed Implementation
[0074] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0075] Genetic algorithms (GAs) simulate the mechanisms of natural selection and heredity in biological evolution, enabling rapid searches in high-dimensional spaces and proving highly effective for solving combinatorial optimization problems. GAs use crossover, mutation, and selection operators to simulate biological genetic processes, iteratively searching within a population to find the optimal solution. This invention aims to leverage the powerful search capabilities of GAs and the accurate simulation of microorganisms using enzyme-constrained models to solve the combinatorial optimization search problem for large-scale gene editing strategies. This shortens the time required for strain optimization design, improves the accuracy of model target prediction, and reduces the number of targets involved in gene editing strategies, facilitating experimental verification and ensuring the feasibility of experimental verification and the economic viability of microbial cell factories. By combining heuristic algorithms and advanced enzyme-constrained models, this invention overcomes the limitations of traditional methods and mechanistic knowledge, enabling the large-scale gene editing strategy search problem to be solved within an acceptable timeframe, opening new avenues for improving the efficiency and stability of industrial microbial production.
[0076] In this embodiment, the latest Saccharomyces cerevisiae enzyme-constrained metabolic network model ecYeast 8.3.4 was used, and the impact of applying additional constraints to the enzyme-constrained model was analyzed using FVA. This model contains 1148 genes and 8144 reactions. Based on the gene-protein-reaction correspondence in the model, 968 genes correspond to reactions directly related to enzyme synthesis. Metabolic phenotype prediction schemes for three wild-type strains were compared using 2-phenylethanol as the target product. Gene dimensionality reduction was performed using the ecFSEOF algorithm, and scores for each reaction were calculated to label relevant genes. Before starting the search for combined gene target editing strategies, MOMA was used to predict the results of single target editing. Then, genetic algorithms were tested based on different encoding formats. Finally, the results and iteration curves were analyzed in detail to obtain valuable data on algorithm optimization and model adjustment, as well as non-intuitive gene editing strategies. Figure 1 This is a schematic diagram of the present invention.
[0077] like Figure 2 The diagram shown is a flowchart of the method of the present invention, and the specific steps are as follows:
[0078] Step S1: Solve the enzyme constraint model with the objective function of maximizing the strain specific growth rate to obtain the simulated metabolic flux of the wild-type strain.
[0079] Comparing the three simulation methods described above, the results show that the model's solution accuracy has a more significant impact on the number of zero values in the prediction results, while the differences between single-step and multi-step single-objective solutions are not substantial. Furthermore, even repeated experiments using the same scheme yielded slightly different results, which will affect the final analysis of the experimental results. Although all schemes effectively predicted the metabolic phenotype of the strains, significant differences were observed in metabolic flux distribution, energy efficiency, and target product yield among the schemes. These differences not only reflect the influence of model settings but also reveal the potential impact of different growth conditions and biological limitations on the metabolic regulation of the strains.
[0080] Results of flux variability analysis as follows Figure 3 The red and blue curves, representing the addition of additional constraints, are both above and to the right of the green curve, representing the absence of additional constraints. This indicates that by imposing appropriate biological constraints on the model, the variability of flux can be effectively reduced, the feasible solution space of the model can be narrowed, and thus the accuracy and reliability of predictions can be improved.
[0081] Step S2 involves solving a series of flux balance analysis problems with fixed biomass synthesis rates to perform dimensionality reduction analysis and labeling of genes directly related to enzyme synthesis within the model.
[0082] An improved forced target throughput scanning algorithm was used to effectively reduce the dimensionality and label the genes in the Saccharomyces cerevisiae enzyme constraint model. The number of genes was reduced from 968 to 52, and each gene was labeled with a suggested editing direction: 11 genes need to be overexpressed, 35 genes need to be knocked down, and 6 genes need to be knocked out.
[0083] Step S3: Predict the yield of single-target editing.
[0084] Yield prediction for single targets was performed based on ecFSEOF results combined with MOMA. The results are as follows: Figure 4 As shown in the figure, the yield of the target product is defined as the ratio of the strain's production of 2-phenylethanol per unit time to its glucose consumption. The red curve in the figure represents the 2-phenylethanol yield of the wild-type strain. It can be observed that the results of most single-target treatments are very close to the reference values. However, the model cannot find the optimal solution for some targets after individual editing, and only one target, when edited according to the indicated direction, improved the yield of 2-phenylethanol. This is because the synthesis of specific chemicals in biology involves multiple reaction pathways, and changes to a single gene may not directly lead to an increase in product yield, indicating that searching for combined editing strategies is necessary.
[0085] Step S4 involves using a genetic algorithm to search for combined target points after dimensionality reduction, adjusting relevant parameters, and obtaining the optimal gene editing combination strategy.
[0086] In the initial experiment, the population size was set to 50, the maximum number of iterations to 500, the crossover probability to 0.8, and the mutation probability to 0.05, with other parameters remaining unchanged. When using the roulette wheel selection method to choose parents, the number of candidate individuals was set to 4. The binary encoding form is directly based on the gene dimensionality reduction and annotation results of the eCFSEOF algorithm, fixing the regulatory direction of key genes as the basis of genetic coding. The results are as follows: Figure 5 The upper part is shown. The results show that the fitness value of the best individual in each generation generally increases, eventually converging to a relatively large value. In the initial population, the fitness value of the first-generation best individual was 0.1374. After 500 iterations, the fitness value of the globally optimal individual reached 0.1412, an improvement of 14.7% compared to the reference fitness value of 0.1231 for the wild-type strain. This result also significantly exceeded the prediction results for a single target. The blue curve shows the mean change in the fitness value of all individuals in each generation. It generally shows an upward trend, but with small fluctuations. This is because there are individuals in the population that prevent the model from finding the optimal solution (fitness value set to 0.01). As the genetic algorithm runs, these poor individuals are gradually eliminated; however, crossover and mutation may produce new, unsolvable individuals, lowering the mean fitness value.
[0087] The symbolic encoding form relies solely on the gene dimensionality reduction results of eCFSEOF. While selecting combinations of genes to be edited, it searches for the editing methods of each gene. This increases the computational complexity of the search but may also reveal new regulatory relationships, especially the interactions between genes. The population size is set to 100, the maximum number of iterations is 2000, and a convergence detection condition is introduced (the search terminates early if the optimal fitness value changes by no more than 0.01 over 1000 consecutive generations). The algorithm results are shown in... Figure 5 The second half. The results show that the fitness value of the best individual in the initial population was 0.1375, and the fitness value of the best individual obtained by the algorithm was 0.1630, representing a 32.4% improvement over the reference yield.
[0088] Comparing the results of the genetic algorithms under the two encoding forms reveals that both design approaches can find gene-editing strategies that result in higher 2-phenylethanol synthesis yields in Saccharomyces cerevisiae. The binary encoding genetic algorithm, leveraging dimensionality reduction and gene annotation results from ecFSEOF, converges more quickly to obtain an optimal gene-editing strategy. In contrast, the symbolic encoding genetic algorithm, based solely on its dimensionality reduction results, can overcome the convergence limit of the former and obtain a better gene-editing strategy, but this requires a larger population size and more iterations, resulting in higher time costs.
[0089] Step S5: Repeat the experiment to verify the stability of the algorithm, and perform statistical analysis and biological interpretation on the obtained combination strategies.
[0090] To explore the impact of different parameters on the genetic algorithm, this experiment adjusted the number of parent individuals selected using the roulette wheel selection method. Increasing the number of candidate individuals to 20 increases the crossover probability of the algorithm. Using binary encoding and expanding the population size to 100, with a maximum allowed number of iterations of 2000, the results are as follows. Figure 6 The upper part is shown. The results show that the algorithm's convergence curve is steeper, and the algorithm converges quickly. The fitness value of the final optimal individual is 0.1411, which is close to the previous experimental results. Due to the increased number of new individuals generated by crossover and mutation, the number of infeasible solutions generated in each iteration also increases, leading to greater volatility in the blue curve.
[0091] Genetic algorithms are heuristic search algorithms, inherently possessing a degree of randomness. This is reflected in the initialization phase and the operation phases of genetic operators such as crossover and mutation. To investigate the stability of the algorithm, this study conducted ten repeated experiments with the same set of parameters. Figure 6The lower half shows the results of these ten algorithm runs, using a genetic algorithm in symbolic encoding. In most cases, the algorithm converges to 0.15 to 0.17. The red curve is derived from the average fitness value of the best individual in each generation of the ten repeated experiments, converging on average to around 0.16.
[0092] In each experiment, the optimal individual and the best fitness value of each generation were retained during the algorithm iteration process. Based on the results of ten repeated experiments, statistical analysis was performed on the ten optimal gene editing strategies, screening for those where each gene target was encoded as non-zero more than 6 times in the ten combined editing strategies. This indicates that the gene was selected as the target for gene intervention in multiple experiments. The results are as follows: Figure 7 As shown, most genes exhibit three different editing methods under different combined editing strategies. This is due to the diversity of biological systems, where different gene intervention methods may lead to the same optimization results.
[0093] In summary, the solution accuracy of the enzyme constraint model significantly impacts the prediction results of the reference strain, and applying additional constraints to the enzyme constraint model effectively reduces the variability of throughput. The gene dimensionality reduction and annotation performed by the ecFSEOF algorithm not only simplifies the search space of the genetic algorithm but also improves search efficiency. Both encoding schemes of the genetic algorithm can identify gene editing strategies that improve the yield of 2-phenylethanol, and key genes in *Saccharomyces cerevisiae* such as Pdx1p are also effectively identified. The results of combined target searches are significantly better than those of single-target treatments, providing more practical guidance for wet experimental design. These findings have important theoretical and practical significance for understanding and improving microbial metabolic engineering.
[0094] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for designing optimal strains for product synthesis using an enzyme constraint model combined with an evolutionary algorithm, characterized in that: Includes the following steps: Step S1: Solve the enzyme constraint model with the objective function of maximizing the strain specific growth rate to obtain the simulated metabolic flux of the wild-type strain. Step S2 involves solving a series of flux balance analysis problems with a fixed biomass synthesis rate, and performing dimensionality reduction analysis and labeling on genes directly related to enzyme synthesis within the model. This gene dimensionality reduction and labeling involves analyzing metabolic response changes in the FBA solution results under 16 different conditions with progressively decreasing biomass synthesis rates for all genes related to enzyme synthesis. The changes in reaction flux are quantified through a two-step scoring method, and corresponding genes are labeled according to different scoring thresholds. Labeling methods include overexpression, knockdown, and deletion. Dimensionality reduction of gene targets involves screening out genes most relevant to the synthesized target product, reducing background noise, and allowing the algorithm to focus on genes most influential in improving the yield of the target product. Gene target labeling additionally provides the intervention methods for the selected genes. Step S3: Predict the yield of single-target editing; Step S4 involves using a genetic algorithm to search for combined target points after dimensionality reduction, adjusting relevant parameters, and obtaining the optimal gene editing combination strategy. Step S5: Repeat the experiment to verify the stability of the algorithm, and perform statistical analysis and biological interpretation on the obtained combination strategies; Step S2 also includes designing a representation for gene editing in the model, wherein: Gene knockout involves setting both the upper and lower bounds of the reaction flux of the corresponding enzyme's synthetic pathway to 0, indicating that the enzyme's synthetic pathway is shut down. Gene overexpression and knockdown require manual adjustment; using the predicted wild-type throughput as a reference, overexpression raises the lower bound of the throughput of the relevant enzyme synthesis reaction to the reference value. The knockdown operation adjusts the upper limit of flux to the reference value. If the reference flux value corresponding to the target to be overexpressed is 0, then its flux lower bound is set to 0. , ; The two-step scoring quantification operation that reflects changes in flux is: Total Flux Score The average ratio of the flux of this reaction to the flux under all forced target flux settings is given by the formula: , in The reaction was carried out under the condition of maximum biomass synthesis rate. The flux is given by , where the molecule in the equation is the sum of the fluxes of the reaction under all simulation conditions. The number of simulations; In addition, any greater than , of The values will be truncated to 1000 to avoid infinitely large values; undetermined values will be assigned 1. Based on these flux scores, the gene scores will then be calculated. As the average reaction flux score for each gene: , In the formula Genetic targets The score is the sum of the reaction scores of all reactions catalyzed by the gene product. It is the total number of reactions catalyzed by the gene product; according to Genes with a value higher than 1 are suggested as overexpression targets. Between arrive The target is between [a certain value] and below [a certain value]. The target to be eliminated is ; among them, , .
2. The method for designing optimal strains for product synthesis using an enzyme-constrained model fusion evolutionary algorithm according to claim 1, characterized in that, The enzyme constraint model was adjusted and validated based on GSMM to ensure it accurately simulates the metabolic state of the strain under different environmental conditions, including the absorption pathways of oxygen, carbon, and nitrogen sources, with all upper limits set to... , Set the model's precision to , This increases our ability to understand protein synthesis.
3. The method for designing optimal strains for product synthesis using an enzyme-constrained model fusion evolutionary algorithm according to claim 1, characterized in that, The flux balance analysis method solves the enzyme constraint model to simulate the metabolic flux of wild-type strains, including the following three methods: Method 1: Given a minimum yield limit for the target product, a single-step flux balance analysis is performed with the objective function of maximizing the specific growth rate of the strain. Method 2, based on Method 1, restricts the specific growth rate to not be lower than its maximum growth rate. , The Free Flow Absorption (FBA) algorithm is used to solve for the FBA with the objective function of minimizing glucose consumption, limiting the rate of glucose consumption to not exceed its minimum consumption amount. times, The FBA is solved with the objective function of minimizing the total enzyme consumption. Method 3: Without limiting the yield of the target product, perform FBA calculations three times as per Method 2, then limit the total enzyme consumption to not exceed its minimum value. times, The FBA is solved with the objective function of maximizing product yield.
4. The method for designing optimal strains for product synthesis using an enzyme constraint model fused with an evolutionary algorithm according to claim 3, characterized in that, The flux balance analysis used is expressed by the following formula: , in, It is a constant vector consisting of 0s and 1s. It is the flux vector representing all reactions. transpose and Multiplication defines the optimization direction of a model, namely, maximizing the flux of a certain reaction. The stoichiometry matrix of the model, and These represent the minimum and maximum constraint boundaries of the reaction, respectively.
5. The method for designing optimal strains for product synthesis using an enzyme constraint model fused with an evolutionary algorithm according to claim 1, characterized in that, In step S3, based on the results of gene dimensionality reduction and annotation, the yield prediction of single-target editing is performed one by one. The metabolic adjustment minimization method is used to solve the edited enzyme constraint model to obtain the metabolic phenotype of the mutant strain and calculate the yield of the target product.
6. The method for designing optimal strains for product synthesis using an enzyme constraint model fusion evolutionary algorithm according to claim 1, characterized in that, In step S4, a genetic algorithm is used to search for combined target points after dimensionality reduction, including two encoding forms: Binary-encoded genetic algorithms rely on gene dimensionality reduction and annotation results to search for combined target points to be edited; Symbolic encoding genetic algorithms rely solely on the results of gene dimensionality reduction to simultaneously find combined targets and the editing direction of targets within a larger search space.