A metabolic network key factor mining method based on flux balance analysis

By employing a metabolic network approach based on flux balance analysis, and utilizing FBA and genetic algorithms to calculate flux changes, this approach addresses the challenge of uncovering key factors in metabolic networks in existing technologies, thereby improving the production efficiency and yield of target products.

CN115762663BActive Publication Date: 2026-01-02WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211432147.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2026-01-02
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently discover and locate key factors in metabolic networks, resulting in insufficient production of target metabolites in cell factories. Furthermore, biological experiments and in vitro simulation methods lack comprehensiveness and guidance.

Method used

A key factor mining method for metabolic networks based on flux balance analysis was adopted. By constructing a metabolic network model, flux changes were calculated using FBA and genetic algorithms to locate the key reactions and enzymes affecting the target compound.

Benefits of technology

It enables rapid discovery of key factors in metabolic networks without complex parameter settings, improving the production efficiency and yield of target products and simplifying the metabolic network optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762663B_ABST
    Figure CN115762663B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on flux balance analysis's metabolic network key factor mining method.The method does not need complex field knowledge, by input metabolic network and target compound, can directly mine the multiple reactions that the biggest influence of target compound is affected, for quickly assisting biologist optimized metabolic network.The application first models existing metabolic network, and determines the target product that needs to be produced, and according to FBA, the reaction that target product can be produced in metabolic network is analyzed, and then using the maximum flux variation calculation method based on genetic algorithm to calculate the influence of each reaction on target metabolite flux, and in turn determine the top-k metabolic reactions that the biggest influence on flux as key factor.The application provides candidate key factor for metabolic pathway design and optimization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of bioengineering, and particularly relates to a method for mining key factors by using flux balance analysis and metabolic network, which can be used to determine the key factors in the metabolic network affecting the metabolic product by analyzing the dynamic adjustment results of the target reaction flux on the metabolic product. BACKGROUND

[0002] In industrial production, many life-essential compounds such as fuel and plastic, polymer and pharmaceutical materials can be produced by microbial metabolic engineering. Metabolic engineering provides a green and sustainable metabolic pathway for manufacturing natural products by manipulating the genes or enzymes of microorganisms for further pathway modification and large-scale production. However, in many cases, due to the insufficient enzyme activity of the reactions in the pathway and the insufficient utilization of the intermediate metabolites, the yield of the target metabolic product of the cell factory is insufficient. Such reactions in the metabolic network that have a great influence on the yield of the metabolic product are defined as key factors for promoting the target metabolic product.

[0003] Due to the corresponding relationship between biochemical reactions and enzymes in the metabolic network, some works are based on the key enzymes to locate the key reactions. For example, isopentenyl diphosphate (IPP)-dimethylallyl diphosphate (DMAPP) isomerase (IDI) is a rate-limiting enzyme of isoprene biosynthesis, which has the effect of adjusting the intracellular concentration of IPP and DMAPP, and it is found that the ordered co-expression of the corresponding gene can significantly improve the production of isoprene in E. coli. The overexpression of these rate-limiting enzymes can effectively increase the metabolic flux of the corresponding key reactions, thereby increasing the flux of the MEP pathway, and ultimately improving the supply of precursors and the production of final products of the target product.

[0004] At present, the mining and positioning of key factors are mostly based on biological experiments, in vitro simulation of pathways, etc. For example, by traversing all genes in the metabolic pathway, sequentially knocking out a single target gene, comparing the yield of the target compound before and after knocking out the target gene, and judging whether the biochemical reaction corresponding to the target gene is a key factor affecting metabolism. However, due to the experimental cost and technical limitations, it is difficult to expand the evaluation of the target pathway. In vitro experiments are to start the metabolic pathway by adding a mixture of enzymes catalyzing the reactions in the pathway to a solution containing substrates and auxiliary factors to simulate the metabolic process in vivo. This method moves the entire pathway in vivo to extract and analyze the product in vitro, thereby locating the key reactions corresponding to the target enzyme. Due to the complexity of the metabolic system, this method lacks the interaction information of enzymes in vivo, and cannot completely simulate the metabolic process.

[0005] Previous work cannot effectively extract multiple metabolic key factors due to the limitations of biological experiments and in vitro simulation approaches. Moreover, for most enzymes in the metabolic pathway, it is not feasible to directly detect the high-throughput products. In addition, it is extremely difficult to modify multiple key enzymes individually, and it may generate new key factors. Some work is based on literature to collect a database of rate-limiting enzymes of different species, which lacks annotation of key enzymes of other unmeasured strains or organisms. Some methods are based on the study of key factors or rate-limiting steps of a single chassis bacterium, which lacks guidance value for other strains. SUMMARY

[0006] In view of the above technical problems, the purpose of the present application is to provide a metabolic network key factor mining method based on flux balance analysis. This method does not require complex domain knowledge, and by inputting the metabolic network model and the target compound, the k reactions that have the greatest impact on the target compound can be directly mined to quickly assist biologists in optimizing the metabolic network. The key factor described in this method refers to a specific biochemical reaction in the metabolic network. Then, according to the correspondence between the reactions and enzymes in the metabolic network, the corresponding catalytic enzyme of the biochemical reaction can be located, so as to analyze the target pathway from proteomics and metabolomics.

[0007] The technical scheme provided by the present application is as follows:

[0008] A metabolic network key factor mining method based on flux balance analysis, comprising the following steps:

[0009] Step 1: Obtain the reaction set R and the metabolite set M from the known metabolic network N, and give the target metabolite m T , wherein m T ∈M;

[0010] Step 2: Obtain the reaction set R T from the reaction set R T , wherein

[0011] Step 3: Use flux balance analysis (FBA) to calculate the flux maximum V of the reaction set R T under the initial condition;

[0012] Step 4: Construct a flux change calculation method A to calculate the target yield change value AV of the given reaction R under a given scale, wherein scale is a variable used to control the flux disturbance of the reaction;

[0013] Step five: Construct a genetic algorithm-based maximum flux change calculation method B using the method A in step four, for calculating the maximum target yield change value AV that each reaction can achieve max , and the disturbance amount scale at this time;

[0014] Step six: Key factor mining using method B in step five: input the metabolic network N and the target metabolite m T Method B outputs the maximum target yield change value corresponding to each reaction in the reaction set R, and then selects the top k reactions with the largest yield change value as the key factors as the result output, k being a set parameter.

[0015] Specifically,

[0016] The known metabolic network N in step one is collected from public resources or provided by researchers, and the storage format of the metabolic network model is SBML. The metabolic network N contains all reaction sets R and metabolite sets M. Each reaction R i contains a metabolite set M corresponding to the reaction i , wherein,

[0017] Specifically,

[0018] In step three, the FBA calculation method is as follows:

[0019] 3.1 Initialize the metabolic network and extract all reactions in the metabolic network;

[0020] 3.2 Construct the stoichiometric matrix S, where each row represents a metabolite and each column represents a reaction;

[0021] 3.3 Construct the linear equation set under steady-state conditions S·v = 0, where v represents the flux constraint of each reaction, and for any reaction i, its range is [α i ,β i ];

[0022] 3.4 Define the maximum optimization problem: Z = c T v, where c is the weight vector and Z is the optimization target, i.e. the maximum flux corresponding to the reaction set R T ; c represents the contribution of the reaction to the target, which is a 0-1 vector (the elements in this vector are only 0 and 1, for example [0, 1, 0, …, 1]), which is 1 only at the reaction position of interest, and 0 elsewhere, i.e. only at the position corresponding to c of the reaction set R T in step two is 1;

[0023] 3.5 Solve the constrained optimization problem using linear programming method to get the maximum target value V.

[0024] In particular,

[0025] The flux variation calculation method A in step four is as follows for a given reaction R and a perturbation scale:

[0026] 4.1 Calculate the flux constraints [a, b] of each reaction R using Flux Variability Analysis (FVA);

[0027] 4.2 Modify the flux constraints of reaction R, let a = a x (1 + scale), b = b x (1 - scale), where a < b;

[0028] 4.3 Calculate the maximum objective value V' of the modified metabolic network using FBA;

[0029] 4.4 Calculate AV = |V - V'| as the output of method A.

[0030] In particular,

[0031] The steps of the genetic algorithm-based maximum flux variation calculation method B in step five are as follows:

[0032] 5.1 Initialize the population popo, set the population size as pop_size, and set the fitness function Func as the flux variation calculation method A in step four to calculate the fitness value of each individual in the initial population;

[0033] 5.2 Iterate the population: set the iteration number g max , and set the current iteration number g = 1; if g < g max , then: (1) generate a new population pop' g from the current population pop g , (2) calculate the fitness value obj' g of the new population pop' g , (3) select pop_size individuals with larger fitness values from the mixed population pop g ∪ pop' g to join the new population pop g+1 , (4) update g = g + 1;

[0034] 5.3 Output the optimal individual in the population.

[0035] More specifically,

[0036] In 5.1, pop is composed of individuals imd, and each individual is composed of a decision variable, i.e., scale, and a fitness value, i.e., the change in target yield AV. That is, pop = {ind1, ind2, …, indi ..., ind pop_size}, ind i = {scale i , ΔV i}, where ΔV i = Func(scale i ).

[0037] Compared with the prior art, the present application has the beneficial effects that:

[0038] 1. The present application is the first key factor mining method based on metabolic network in the field of metabolic engineering, which can well promote the development of the field of metabolic engineering.

[0039] 2. The present application does not need complex parameter setting, only needs to give the metabolic model to be optimized and the target compound to be produced, and the key factor in the metabolic network can be obtained.

[0040] 3. The method in the present application does not need to have any understanding of metabolite concentration, does not need to understand the enzyme kinetics of the system, and only uses stoichiometric coefficients to mine the key factor of the specific target compound. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 is a flow chart of the metabolic network key factor mining method based on flux balance analysis of the present application;

[0042] Figure 2 is a flow chart of the flux balance analysis calculation method in the present application;

[0043] Figure 3 is a flow chart of the flux change calculation method in the present application;

[0044] Figure 4 is a flow chart of the maximum flux change calculation method based on genetic algorithm in the present application. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0046] The technical solutions of the present application will be further described below in combination with the drawings and embodiments.

[0047] Embodiment 1

[0048] A metabolic network key factor mining method based on flux balance analysis, comprising:

[0049] Step one: obtain reaction set R and metabolite set M from known metabolic network N, given target metabolite m T , wherein m T ∈M;

[0050] Step two: obtain reaction set R T from reaction set R, where all reaction products are m T , wherein

[0051] Step three: calculate flux maximum V of reaction set R under initial conditions using flux balance analysis FBA; T

[0052] Step four: construct flux change calculation method A to calculate target yield change value ΔV of given reaction R at given scale, where scale is a variable used to control the flux disturbance of the reaction;

[0053] Step five: use method A in step four to construct maximum flux change calculation method B based on genetic algorithm, which is used to calculate the maximum target yield change value ΔV max and the disturbance scale of each reaction;

[0054] Step six: use method B in step five to mine key factors: input metabolic network N and target metabolite m T , method B outputs the maximum target yield change value corresponding to each reaction in reaction set R, and then selects the top k reactions with the largest yield change value as the key factors as the result output, k is a set parameter.

[0055] Specifically, in step one, the known metabolic network N is collected from public resources or provided by researchers, and the storage format of the metabolic network model is SBML. The metabolic network N contains all reaction set R and metabolite set M. Each reaction R i ∈r contains a set of metabolites corresponding to the reaction M i , wherein,

[0056] Specifically, in step three, the calculation method of FBA is as follows:

[0057] 3.1 Initialize the metabolic network and extract all reactions in the metabolic network;

[0058] 3.2 Construct stoichiometric matrix S, each row represents a metabolite, and each column represents a reaction;

[0059] ​3.3 Construct the linear equation system S·v=0 under steady-state conditions, where v represents the flux constraint for each reaction, and for any reaction R i Its range is [α i ,β i ]; c represents the contribution of the reaction to the target, which is a 0-1 vector, with 1 only at the position of the reaction of interest and 0 at the other positions, that is, only in the reaction set R in S1.2. T The corresponding value at point c is 1;

[0060] 3.4 Define the maximization optimization problem: Z = c T v, where c is the weight vector and Z is the optimization objective, i.e., the reaction set R. T The corresponding maximum flux;

[0061] 3.5 Solve the constrained optimization problem using linear programming to obtain the maximum objective value V.

[0062] Specifically, in step four, the flux change calculation method A, given the response R and the perturbation scale, follows the following process:

[0063] 4.1 Flux constraint [,β] for each reaction R is calculated using flux change analysis (FVA);

[0064] 4.2 Modify the flux constraint of reaction R, let α=α×(1+scale) , β=β×(1-scale), where α≤β;

[0065] 4.3 Calculate the maximum target value V′ of the metabolic network after modifying the constraints of reaction R using FBA;

[0066] 4.4 Calculate ΔV = |VV′| and use it as the output of method A.

[0067] Specifically, in step five, the steps of the maximum flux change calculation method B based on the genetic algorithm are as follows:

[0068] 5.1 Initialize the population pop0, set the population size to pop_size, and set the fitness function Func to flux change calculation method A in step four, which is used to calculate the fitness value of each individual in the initial population;

[0069] 5.2 Iterate the population: Set the number of iterations g max And set the current iteration number g = 1; if g ≤ g max Then: (1) pop from the current population g A new population pop′ is generated in the middle. g (2) Calculate the new population pop′ g fitness value obj′ g (3) Pop from mixed populationg ∪pop′ g The pop_size individuals with greater fitness value are selected into the new population pop g+1 , (4) update g = g + 1;

[0070] 5.3 Output the optimal individual in the population.

[0071] More specifically, in 5.1 of the step five, the pop is composed of individuals ind, and each individual is composed of a decision variable, i.e. scale, and a fitness value, i.e. the target product variation value ΔV. That is, pop = {ind1, ind2, …, ind i , …, ind pop_size}, ind i = {scale i , ΔV i}, where ΔV i = Func(scale i ).

[0072] More specifically, in 5.2 of the step five, the generation method of the new population pop′ g includes roulette method, competition method, etc.

[0073] This embodiment takes Escherichia coli model iJO1366 as an example, in which the target product is set as acetyl coenzyme A (accoa_c), and the model is from the public resource database BIGG (King ZA, Lu JS, A, Miller PC, Federowicz S, Lerman JA, Ebrahim A, Palsson BO, and Lewis NE. BiGG Models: A platform for integrating, standardizing, and sharing genome-scale models (2016) Nucleic Acids Research 44 (D1): D515-D522. doi: 10.1093 / nar / gkv1049). The model contains 1805 metabolites and 2583 reactions.

[0074] This embodiment adopts the above-mentioned FBA-based metabolic network key factor mining method to calculate the top 10 key factors when the target product is accoa_c, as shown in Table 1.

[0075] Table 1 Key factors

[0076]

[0077] It should be understood that the above description is merely a detailed specification of the preferred embodiments of the application and is not intended to limit the scope of the application in any way. Various modifications of the application in accordance with the principles of the application will be apparent to those with skill in the art upon reference to the description. It is therefore contemplated that the application can admit of still other modifications and embodiments, all of which are intended to be within the scope of the following claims.

[0078] The above description is merely the preferred specific embodiments of the application and is not intended to limit the scope of the application in any way. Any modification, equivalent replacement and improvement made by any person skilled in the art within the technical scope of the application should be included in the scope of the protection of the application.

Claims

1. A metabolic network key factor mining method based on flux balance analysis, characterized in that, The steps are as follows: Step one: Obtain a reaction set from a known metabolic network and a metabolite set , given a target metabolite wherein, ;​ Step two: Collect all reaction products from the reaction collection as the reaction collection where ; Step three: Calculate flux maxima of the reaction set using flux balance analysis under initial conditions ;​ The calculation method of the flux balance analysis is as follows: 3.1 Initialize the metabolic network and extract all reactions in the metabolic network; 3.2 Constructing the stoichiometric matrix each row represents one metabolite and each column represents one reaction; 3.3 Constructing the linear system under steady-state conditions where, denotes the flux constraint for each reaction, for any reaction ranging from ; 3.4 Defining the maximization optimization problem: where, is the weight vector, is the optimization objective, i.e. the set of reactions corresponding to the maximum flux; denotes the contribution of a reaction to the objective, is a 0-1 vector, which is 1 only at the position of the reaction of interest and 0 elsewhere, i.e. only in the set of reactions in step two corresponding to the maximum flux; is 1 at that position. 3.5 Solving the constrained optimization problem using linear programming to obtain the maximum objective value ; Step four: Construct flux change calculation method A to calculate the target change in output for a given reaction At a given target output change value where is the variable used to control the flux perturbation of the reaction; Method A for flux change calculation for a given reaction and perturbation The specific procedure is as follows: 4.1 Calculate flux constraints for each reaction using flux variation analysis ;​ 4.2 Modification of the reaction of flux constraints, allowing , where ; 4.3 Calculating modified reactions using FBA the maximum objective value of a constrained metabolic network ; 4.4 Calculation , as output from Method A; Step five: Construct a genetic algorithm based maximum flux change calculation method B using the method A in step four for calculating the maximum target production change value that can be achieved by each reaction and the amount of disturbance at this time ; Step six: Key factor mining using method B in step five: input the metabolic network and target metabolite , method B outputs the maximum change in target production value for each reaction in the reaction set ; and then filters out the top reactions with the largest change in production value, i.e. the key factors, as the result output, which is a set parameter.

2. The method of claim 1, wherein: In the step one, the known metabolic network N is collected from public resources or provided by researchers, wherein the metabolic network model is in SBML format, and the metabolic network contains all reaction sets and metabolite sets ; each reaction contains a set of metabolites corresponding to the reaction , wherein .

3. The method of claim 1, wherein: The steps of the calculation method B of the maximum flux variation based on the genetic algorithm in step five are as follows: 5.1 Initialize population Set the population size to Set the fitness function Method A for flux change calculation for step four, used to calculate the fitness value of each individual in the initial population; 5.2 Iteration of the population: set the number of iterations , and set the current iteration number ; if , then: (1) generate a new population from the current population , (2) calculate the fitness values of the new population , (3) select individuals with higher fitness values from the mixed population to join the new population , (4) update ; 5.3 Output the optimal individual in the population.

4. The method of claim 3, wherein: Said step five's 5.1, consists of individual consists of individual and fitness value, i.e. target yield change value ; i.e. , wherein .