Method for predicting metabolic flux of living body, and storage medium

By constructing metabolic network models and prediction network models, and using technical means such as recurrent neural networks, the problems of low efficiency and low accuracy of biological metabolic flux calculation in the existing technology are solved, efficient and accurate metabolic flux prediction are achieved, and the dynamic mechanism of metabolic regulation is revealed.

CN120299503APending Publication Date: 2025-07-11GUIZHOU NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510261480.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems such as low accuracy, long time consumption, low efficiency and high computing resource requirements in the calculation of biological metabolic flux, making it difficult to obtain biological metabolic flux efficiently and accurately.

Method used

By constructing a differential equation model of a metabolic network, combining data simulation and prediction network model, metabolic flux prediction is used to predict traditional recurrent neural networks (SRN), long and short-term memory networks (LSTM) and gated recurrent units (GRU). Optimization algorithms such as least squares method fit experimental data and inversely deduce the metabolic flux distribution.

Benefits of technology

It improves the accuracy and efficiency of metabolic flux prediction, can better capture the time-dependent behavior of metabolic networks, reveal the dynamic mechanism of metabolic regulation, and provide biological research with greater reference value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299503A_ABST
    Figure CN120299503A_ABST
Patent Text Reader

Abstract

The invention discloses an organism metabolic flux prediction method and a storage medium, and the method comprises the steps: responding to instruction information, and determining a metabolic network of an organism to be predicted; performing data simulation generation on the metabolic network to obtain simulation data which is used for representing flux distribution of the metabolic network; and inputting the simulation data into a prediction network model, and obtaining the predicted metabolic flux of the metabolic network output by the prediction network model. According to the technical scheme, the metabolic flux of organisms is efficiently and accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biological metabolism, and specifically relates to a method for predicting the metabolic flux of an organism and a storage medium. Background Art

[0002] With the development of technology, biological research is becoming more and more intelligent, and the research on the metabolic flux of organisms is a key research direction in the field of biological metabolism. The analysis of the metabolic flux of organisms can reveal the activity level of specific metabolic pathways in organisms, reveal the sites and mechanisms of metabolic regulation, and deepen our understanding of how metabolites participate in cell decision-making.

[0003] In the past, the calculation method of biological metabolic flux usually used the method of isotope labeling to directly observe organisms, and then calculated the metabolic flux of a certain metabolic network in biological activities. However, this method has low accuracy, long time consumption, low efficiency, and little reference significance. To address this problem, those skilled in the art introduced computer technology and calculated through mathematical models and fluxes: by constructing a differential equation model based on the metabolic network to describe the spatio-temporal changes of isotope labeling. By optimizing algorithms (such as the least squares method) to fit experimental data and inversely deduce the metabolic flux distribution.

[0004] However, due to the high complexity of the experiment, non-steady-state experiments require precise control of the labeling time points, which is difficult to operate, and the detection of isotope labeling data (such as mass spectrometry) is costly, time-consuming, and unstable. Secondly, there are bottlenecks in the calculation. The solution of dynamic models involves complex differential equations and optimization problems, requires a large amount of computing resources, and the existing tools have limited support for biological analysis.

[0005] Therefore, how to efficiently and accurately obtain the metabolic flux of organisms is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0006] The purpose of this application is to efficiently and accurately obtain the metabolic flux of organisms.

[0007] Other features and advantages of this application will become apparent through the following detailed description, or will be learned in part through the practice of this application.

[0008] According to one aspect of the embodiments of this application, a method for predicting the metabolic flux of an organism is provided, including:

[0009] Respond to the instruction information to determine the metabolic network of the organism to be predicted;

[0010] Perform data simulation generation on the metabolic network to obtain simulation data, where the simulation data is used to represent the flux distribution of the metabolic network;

[0011] Input the simulation data into the prediction network model to obtain the predicted metabolic fluxes of the metabolic network output by the prediction network model.

[0012] According to one aspect of the embodiments of the present application, before generating simulation data by simulating the metabolic network, the method further includes:

[0013] Determine the isotope-labeled object and isotope-labeled substrate in response to the instruction information.

[0014] According to one aspect of the embodiments of the present application, the metabolic network includes at least one metabolic reaction; generating simulation data by simulating the metabolic network includes:

[0015] Select a set proportion of metabolic reactions in all metabolic networks as target metabolic reactions;

[0016] When the duration of the target metabolic reaction reaches the set duration and the flux value of the target metabolic reaction is within the set flux interval, obtain the flux value of the target metabolic reaction as the sampling flux value of the target metabolic reaction at the set duration, so that the sampling flux value can be used as the non-equilibrium flux distribution of the target metabolic reaction;

[0017] Simulate the isotope distribution in the metabolic network according to each initial sampling flux value to obtain the simulation data.

[0018] According to one aspect of the embodiments of the present application, simulating the isotope distribution in the metabolic network according to each initial sampling flux value to obtain the simulation data includes:

[0019] Obtain a set number of sampling flux values, and simulate the isotope distribution in the metabolic network according to the sampling flux values to obtain the sampling data corresponding to the sampling flux values, so as to obtain the simulation data. The sampling flux values include metabolic substrate data and isotope isomer distribution data when the target metabolic reaction reaches each set duration.

[0020] According to one aspect of the embodiments of the present application, inputting the simulation data into the prediction network model includes:

[0021] Remove noise and transform the simulation data according to the metabolic substrate data and isotope isomer distribution data in the simulation data to obtain standardized simulation data;

[0022] Input the standardized simulation data into the network model.

[0023] According to one aspect of the embodiments of the present application, the method for training the network model includes:

[0024] Select a first metabolic network and a second metabolic network from each of the metabolic networks according to the number of metabolic reactions in each metabolic network, where the number of metabolic reactions in the first metabolic network is less than that in the second metabolic network;

[0025] Obtain first simulation data in the first metabolic network and input the first simulation data into an initial prediction network model, where the first simulation data has corresponding first sample data;

[0026] Calculate a first prediction error of the initial prediction network model according to the prediction result of the initial prediction network model and the first sample data, and adjust the parameters of the initial prediction network model according to the first prediction error until the first prediction error is less than a first error threshold;

[0027] Obtain second simulation data in the second metabolic network and input the second simulation data into the initial prediction network model, where the second simulation data has corresponding second sample data;

[0028] Calculate a second prediction error of the initial prediction network model according to the prediction result of the initial prediction network model and the second sample data, and adjust the parameters of the initial prediction network model according to the second prediction error until the second prediction error is less than a second error threshold, where the second error threshold is less than the first error threshold;

[0029] Take the initial prediction network model as the prediction network model.

[0030] According to one aspect of the embodiments of the present application, the input flux values of the first metabolic network and the second metabolic network are respectively within corresponding input setting intervals; the global flux splits of the first metabolic network and the second metabolic network are within corresponding global setting intervals.

[0031] According to one aspect of the embodiments of the present application, there is provided a computer device, including a memory, a processor, and a readable program stored on the memory, where the processor executes the readable program to implement the method as described in any one of the above.

[0032] According to one aspect of the embodiments of the present application, there is provided a readable storage medium, on which a readable program / instruction is stored, and when the readable program / instruction is executed by a processor, the method as described in any one of the above is implemented.

[0033] According to one aspect of the embodiments of the present application, there is provided a program product, including a readable program / instruction, and when the readable program / instruction is executed by a processor, the method as described in any one of the above is implemented.

[0034] In this application, by responding to the user's operation to form instruction information, the metabolic network that the user wants to analyze is selected, and then combined with data simulation technology, the flux distribution of the metabolic network is obtained to form simulation data. Finally, the simulation data is input into the prediction network model to obtain the predicted metabolic flux of the metabolic network output by the prediction network model. In the embodiments of this application, a prediction network model is introduced in the field of biological metabolism to predict the metabolic flux of the metabolic network, replacing the traditional mathematical calculation method, which greatly improves the accuracy and efficiency of predicting the metabolic flux.

[0035] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings

[0037] The accompanying drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0038] Figure 1 Shows a flowchart of a method for predicting the metabolic flux of an organism according to an embodiment of this application.

[0039] Figure 2 Shows a flowchart of generating data simulation for the metabolic network according to an embodiment of this application to obtain simulation data.

[0040] Figure 3 Shows a flowchart of inputting simulation data into a prediction network model according to an embodiment of this application.

[0041] Figure 4 Shows a schematic diagram of the original feature composition of a metabolic network according to an embodiment of this application.

[0042] Figure 5 Shows a flowchart of training a prediction network model according to an embodiment of this application.

[0043] Figure 6 Shows a schematic diagram of the hyperparameter settings of an initial prediction neural network according to an embodiment of this application.

[0044] Figure 7 Shows a comparison chart of the errors of various hyperparameters according to an embodiment of this application.

[0045] Figure 8 Shows a control schematic diagram of the activation function settings of the initial prediction neural network according to an embodiment of the present application.

[0046] Figure 9 Shows a schematic diagram of hyperparameters of different neural networks as the initial prediction network model according to an embodiment of the present application.

[0047] Figure 10 Shows the training results under different network models and different hyperparameters according to an embodiment of the present application.

[0048] Figure 11 Shows a schematic diagram of the hyperparameter settings of the model when different activation functions are selected according to an embodiment of the present application.

[0049] Figure 12 Shows the hyperparameter settings of the model when different optimization algorithms are selected according to an embodiment of the present application.

[0050] Figure 13 Shows a comparison chart of the error rate and training time of different activation functions according to an embodiment of the present application.

[0051] Figure 14 Shows a comparison chart of the error rate and training time of different optimization algorithms according to an embodiment of the present application.

[0052] Figure 15 Shows a comparison chart of the model performance using two loss functions, MAE and MSE, according to an embodiment of the present application.

[0053] Figure 16 Shows a schematic diagram of the evaluation of the prediction network model used in the present application according to an embodiment of the present application.

[0054] Figure 17 Shows a block diagram of the computer system for implementing the prediction method of biological metabolic flux according to an embodiment of the present application. Detailed implementation manners

[0055] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0056] In addition, the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0057] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0058] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0059] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0060] Please refer to Figure 1 , Figure 1 , which shows a flowchart of a method for predicting the metabolic flux of an organism according to an embodiment of the present application. The embodiments of the present application provide steps of a method for predicting the metabolic flux of an organism, including:

[0061] Step S110, in response to the instruction information, determine the metabolic network of the organism to be predicted;

[0062] Step S120, perform data simulation generation on the metabolic network to obtain simulation data, where the simulation data is used to represent the flux distribution of the metabolic network;

[0063] Step S130, input the simulation data into the prediction network model to obtain the metabolic flux of the metabolic network predicted by the prediction network model.

[0064] The above three steps will be described in detail below.

[0065] It should be clear that each organism contains multiple metabolic networks, each metabolic network contains multiple metabolic pathways, and each metabolic pathway includes multiple metabolic reactions. Exemplarily, the central metabolic network of photosynthesis in plant leaves includes pathways such as the C3 cycle, photorespiration, starch synthesis, sucrose synthesis, and sucrose synthesis. These metabolic pathways include 55 metabolites and 57 metabolic reactions, including 23 reversible metabolic reactions, 1 input metabolic reaction, and 12 output metabolic reactions.

[0066] It should be further clear that metabolic flux refers to the rate at which substances pass through a specific metabolic reaction path in a metabolic network, usually expressed as the number of moles of a substance per unit time (such as mmol / gDW / h). It reflects the dynamic process of metabolic reactions and is a core indicator of the function of the metabolic network. That is, the metabolic network has a metabolic flux, specifically referring to the overall rate of biological metabolic reactions in the metabolic network. The metabolic pathway also has a metabolic flux, specifically referring to the overall rate of biological metabolic reactions in the metabolic pathway. And each metabolic reaction also has a metabolic flux, specifically referring to the rate at which the metabolic reaction proceeds. The metabolic flux of the metabolic network is not simply the sum of the metabolic fluxes of each metabolic pathway and each metabolic reaction. It needs to consider the types of each metabolic reaction (such as forward and reverse), which is a complex calculation process.

[0067] It should also be clear that in the following text, the term "metabolic flux" specifically refers to the metabolic flux of the metabolic network. The metabolic flux of a metabolic pathway or a metabolic reaction is referred to as "flux".

[0068] In step S110, corresponding instruction information is generated in response to the user's operation behavior. The instruction information is used to select the object for predicting the metabolic flux, that is, the metabolic network that the user wants to predict. It should be clear that the user's operation behavior is a behavior with specific meaning emitted by the user by controlling their own body. Such as manually operating a button, or outputting information, or specific body language.

[0069] After determining the metabolic network, continue to generate corresponding new instruction information in response to the user's operation behavior, or in the content indicated by the instruction information in step S110, set metabolic parameters for the metabolic network, such as metabolic substrates, isotope labeling objects, input flux values, and global flux value limits. Make the metabolic network simulate and run under the set parameters, because only in this way, the metabolic flux of the obtained metabolic network is more meaningful for reference.

[0070] In some embodiments, after determining the metabolic network of the organism for which prediction is needed, continue to determine the metabolic parameters corresponding to the metabolic network, where the metabolic parameters are used to determine the operating environment of the metabolic network and the operations to be performed on the metabolic network. Exemplarily, determine the isotope-labeled objects, isotope-labeled substrates, isotope types, and isotope concentrations that need to be clarified for metabolic flux prediction of the metabolic network.

[0071] In some embodiments, carbon-13 is used as the isotope, an organic or inorganic substance containing carbon-13 is used as the isotope-labeled object, and carbon dioxide is used as the isotope-labeled substrate.

[0072] In some embodiments, the input flux value of the metabolic network is restricted to be between 32.5 and 42.5.

[0073] In step S120, according to the metabolic parameters set for the metabolic network, simulate the operation of the metabolic network to obtain the fluxes of the metabolic reactions in the metabolic network, and then obtain the flux distribution of the metabolic network. The flux distribution of the metabolic network refers to the fluxes of the metabolic reactions in the metabolic pathways in the metabolic network. Use the flux distribution of the metabolic network as the simulation data.

[0074] In some embodiments, the simulation data only includes the fluxes of some metabolic reactions in some metabolic pathways. To create a non-equilibrium flux distribution so that the metabolic network can meet the conditions of non-steady-state 13C metabolic flux analysis, and further enable the finally predicted metabolic fluxes of the metabolic network to capture the time-dependent behavior of the metabolic network and reveal the dynamic mechanism of metabolic regulation. Provide greater reference value for biological research.

[0075] Please refer to Figure 2 , Figure 2 which shows a flowchart of generating simulation data for the metabolic network according to an embodiment of the present application. The metabolic network includes at least one metabolic reaction. The embodiment of the present application provides step S120 for generating simulation data for the metabolic network, including:

[0076] Step S121, select a set proportion of metabolic reactions in all metabolic networks as target metabolic reactions;

[0077] Step S122, when the duration of the target metabolic reaction reaches the set duration and the flux value of the target metabolic reaction is within the set flux range, obtain the flux value of the target metabolic reaction as the sampling flux value of the target metabolic reaction at the set duration, so that the sampling flux value can be used as the non-equilibrium flux distribution of the target metabolic reaction;

[0078] Step S123, according to each sampling flux value, simulate the isotope distribution in the metabolic network to obtain the simulation data.

[0079] The above three steps will be described in detail below.

[0080] In step S121, each metabolic network includes a plurality of metabolic reactions, and a set proportion of the metabolic reactions are randomly selected from the plurality of metabolic reactions as target metabolic reactions. Exemplarily, 20% of the metabolic reactions in the metabolic network are randomly selected as the target metabolic reactions of the metabolic network.

[0081] In step S122, when the duration of the target metabolic reaction reaches a set duration and the flux value of the target metabolic reaction is within a set flux interval, the flux value of the target metabolic reaction is obtained, because only the data collected when the flux value of the target metabolic reaction is within the set flux interval is meaningful.

[0082] The obtained flux value of the target metabolic reaction is used as the sampling flux value of the target metabolic reaction at the set duration. That is, when the metabolic reaction reaches the set duration, the obtained flux value is used as the sampling flux value, and each sampling flux value has a corresponding time dimension. It should be clear that the set duration is a set. For example, when the reaction duration of the metabolic reaction is 3 seconds, 6 seconds... N seconds, N + 3 seconds, they are all set durations. Among them, the metabolic duration of the metabolic network is the metabolic duration of the target metabolic reaction.

[0083] In some embodiments, the set durations are respectively: the time nodes when the metabolic network conducts for 1s, 100s, 500s, 1000s, 2000s, 4000s, 6000s.

[0084] In this way, the sampling flux value can be used as the non-equilibrium flux distribution of the corresponding target metabolic reaction. The non-equilibrium flux distribution can enable the finally obtained simulation data to perform non-steady-state carbon-13 metabolic flux analysis. It can capture the time-dependent behavior of the metabolic network and reveal the dynamic mechanism of metabolic regulation.

[0085] In step S123, according to each initial sampling flux value, the isotope distribution in the metabolic network is simulated to obtain simulation data. The isotope distribution refers to the distribution of organic or inorganic substances with isotope labels in each metabolic reaction or metabolic pathway in the metabolic network. The distribution includes the "amount" of the isotope-labeled organic and inorganic substances distributed in each metabolic reaction or metabolic pathway, as well as the transformation rate between various isotope-labeled substances, which can reflect the reaction rate or the intensity of the reaction of each metabolic reaction.

[0086] It should be clear that the sampling flux values of the target metabolic reaction can be obtained multiple times, and a set number of sampling flux values are obtained. The isotope distribution in the metabolic network is simulated through the obtained sampling flux values, and the sampling data corresponding to each sampling flux value is obtained. The simulation data includes all the sampling data. It should be clear that the sampling data includes at least the metabolic substrate data and the isotopologue distribution data when the target metabolic reaction reaches each set duration. A metabolic substrate is the starting substance or reactant participating in a metabolic reaction. They are key components in metabolic pathways and are converted into other molecules through a series of enzyme-catalyzed chemical reactions, thereby providing energy for cells, constructing biomolecules, or completing other physiological functions. An isotopologue (or isotopic isomer) is a molecule with the same chemical composition but different isotopic compositions. These molecules are exactly the same in terms of the types and numbers of atoms, but some atoms are replaced by their isotopes, resulting in differences in molecular mass. In the present application, isotopologues are often used to label molecules to trace reaction paths or metabolic processes.

[0087] In some embodiments, for the sampling flux values obtained for the metabolic network, the sampling data corresponding to the set number of sampling flux values is taken as the simulation data, and each sampling data includes metabolic substrate data and isotopologue distribution data.

[0088] In some embodiments, a set number of sampling flux values are obtained, and the isotope distribution in the metabolic network is simulated according to the sampling flux values to obtain the sampling data corresponding to the sampling flux values, so as to obtain the simulation data. The sampling flux values include the metabolic substrate data and the isotopologue distribution data when the target metabolic reaction reaches each set duration.

[0089] In some embodiments, for the central metabolic network of plant leaf photosynthesis, 10,000 sets of sampling data are collected as the simulation data. Each set of sampling data includes metabolic substrate data and isotopologue distribution data. The metabolic substrate data includes two isotopologues of carbon dioxide, and the distribution data includes 7 compounds such as ACE, AKG, CIT, GLN, GLU, MAL, and SUC (the sampling time points are 1 s, 100 s, 500 s, 1000 s, 2000 s, 4000 s, and 6000 s).

[0090] In step S130, the obtained simulated data is input into a pre-trained prediction network model, and the prediction network model outputs the predicted metabolic flux based on the simulated data. It should be noted that the prediction network model can be any one of a traditional recurrent neural network (SRN), a long short-term memory network (Long Short-Term Memory, LSTM), and a gated recurrent unit (Gated Recurrent Unit, GRU).

[0091] In an embodiment of the present application, the traditional recurrent neural network (SRN) is taken as the optimal embodiment of the present application. In some embodiments, the traditional recurrent neural network (SRN) uses the Adam algorithm to predict the metabolic flux of the metabolic network. In some embodiments, the number of hidden layers of the traditional recurrent neural network (SRN) is 2 layers. In some embodiments, the activation function of the traditional recurrent neural network (SRN) is ReLU (Chinese).

[0092] Please refer to Figure 3 , Figure 3 which shows a flowchart of inputting simulated data into the prediction network model according to an embodiment of the present application. The embodiment of the present application provides step S130 of inputting simulated data into the prediction network model, including:

[0093] Step S131: According to the metabolic substrate data and isotope isomer distribution data in the simulated data, noise removal and transformation are performed on the simulated data to obtain standardized simulated data;

[0094] Step S132: Input the standardized simulated data into the network model.

[0095] The above two steps are described in detail below.

[0096] In step S131, the performance of the prediction network model highly depends on the quality of simulated data preprocessing. Before inputting the obtained data into the prediction network model, it is necessary to perform cleaning, transformation, anomaly processing, etc. on the simulated data to obtain standardized simulated data for inputting into the prediction network model. For the simulated data, its specific structure is as Figure 4 shown Figure 4The figure shows a schematic diagram of the original feature composition of a metabolic network according to an embodiment of the present application. Among them, the metabolite MID information of the metabolic network is the distribution characteristic of the isotope isomer distribution data. The Metabolic Network Reaction Pool refers to the set of all possible metabolic reactions in an organism. These reactions constitute the metabolic network, supporting cell growth, maintenance, and response to environmental changes. Its dimensions can include the following aspects: 1 Reactants and products, each reaction pool contains specific reactants and products, and these substances are converted into each other during metabolism. 2 Enzyme catalysis, reactions are usually catalyzed by specific enzymes, and the activity and expression level of the enzymes affect the reaction rate. 3 Reaction rate, the reaction rate is determined by the substrate concentration, enzyme activity, and regulatory factors, and is usually described by kinetic equations. 4 Thermodynamic feasibility, whether a reaction is feasible depends on the change in Gibbs free energy (ΔG), and a negative value indicates that the reaction can proceed spontaneously. 5 Regulatory mechanisms, reactions are regulated by mechanisms such as feedback inhibition and feedforward activation to ensure metabolic balance. 6 Metabolic flux, flux balance analysis (FBA) is used to quantify the flow of substances in the metabolic network. 7 Gene association, reactions are associated with the genes encoding the enzymes, and gene expression affects the production and activity of the enzymes. 8 Metabolite concentration, metabolite concentration affects the reaction rate and direction and is the key to dynamic balance. 9 Network topology, the position and connection method of the reaction pool in the metabolic network affect the overall function. 10 Environmental response, the metabolic network can adjust the reaction rate and path according to environmental changes (such as nutrient supply).

[0097] By Figure 4 It is not difficult to see that the characteristic value range of the reaction pool information is quite different from that of the characteristics composed of other data. Therefore, the range is scaled to [0,1] through Min-Max normalization.

[0098] The formula is:

[0099]

[0100] Among them, for the data we want to process, such as x1, x2, x3, x4 to x n A total of n groups of sampling data, and n groups of new data y1, y2, y3, y4 to yn are obtained by transforming them through min-max normalization. In the formula, both i and j belong to the interval [1, n]. i is the sequence number of the sampling flux in the simulated data. It should be noted that there is no restriction on the sorting method of the sampling data. x i is the reaction pool information of the i-th sampling data. Among them, represents the maximum value of the reaction pool information corresponding to each sampling data in the simulated data, represents the minimum value of the reaction pool information corresponding to each sampling data in the simulated data.

[0101] In addition, since the proportion of carbon-13 labeled information in the abundance distribution of some mass isotopologues is extremely low (generally less than 10-5) and it is basically impossible to measure in experiments, it is necessary to screen the characteristics of the samples. For some sample data with missing values, processing is also required: if the missing proportion < 5%, directly delete the samples with missing values, otherwise use linear interpolation or fill the missing data with the mean value of adjacent time points.

[0102] After the above process, noise removal and transformation of the simulation data are completed to obtain the standardized simulation data. In step S132, the obtained standardized simulation data is input into the prediction network model as described above, and the metabolic flux of the metabolic network output by the prediction network model is obtained.

[0103] Please refer to Figure 5 , Figure 5 which shows a flowchart of training a prediction network model according to an embodiment of the present application. The embodiment of the present application provides steps for training a prediction network model, including:

[0104] Step S201, according to the number of metabolic reactions in each metabolic network, select a first metabolic network and a second metabolic network in each metabolic network, where the number of metabolic reactions in the first metabolic network is less than that in the second metabolic network;

[0105] Step S202, obtain the first simulation data in the first metabolic network and input the first simulation data into the initial prediction network model, and the first simulation data has corresponding first sample data;

[0106] Step S203, according to the prediction result of the initial prediction network model and the first sample data, calculate the first prediction error of the initial prediction network model, and adjust the parameters of the initial prediction network model according to the first prediction error until the first prediction error is less than the first error threshold;

[0107] Step S204, obtain the second simulation data in the second metabolic network and input the second simulation data into the initial prediction network model, and the second simulation data has corresponding second sample data;

[0108] Step S205, according to the prediction result of the initial prediction network model and the second sample data, calculate the second prediction error of the initial prediction network model, and adjust the parameters of the initial prediction network model according to the second prediction error until the second prediction error is less than the second error threshold, and the second error threshold is less than the first error threshold;

[0109] Step S206, use the initial prediction network model as the prediction network model.

[0110] The above 6 steps are described in detail below.

[0111] In step S201, according to the number of metabolic reactions in each metabolic network, the first metabolic network and the second metabolic network are selected from each metabolic network, and the number of metabolic reactions in the first metabolic network is less than that in the second metabolic network. In some embodiments, the metabolic networks used for training the initial prediction network model are divided into two types: the first metabolic network and the second metabolic network according to their sizes. Among them, the small-scale first metabolic network A is a part of the tricarboxylic acid cycle. Since there are few precedents for using only metabolic fluxomics data to predict metabolic fluxes in neural networks currently, the feasibility of the initial prediction network model is verified through the first metabolic network, so as to judge whether the initial prediction network model can capture the dynamic non-linear relationship of the metabolic network. The second metabolic network selects the central metabolic network of plant leaf photosynthesis with a larger scale to judge whether the initial prediction network model can improve the prediction efficiency and shorten the prediction time.

[0112] In step S202, the first simulation data in the first metabolic network is obtained, and the first simulation data is input into the initial prediction network model. The first simulation data has corresponding first sample data. Exemplarily, first, the target metabolic reaction is determined in the first metabolic network, then the sampling data is determined, and finally, a set number of sampling data values are collected as the first simulation data. Each sampling data includes metabolic substrate data (CIT not labeled with carbon 13, [1,2,3- 13 C3]CIT, [U- 13 C6]CIT), isotope isomer distribution data (including mass isotope isomer distribution data at time points of 1s, 100s, 500s, 1000s, 2000s, 4000s, 6000s), and prediction labels (the prediction labels refer to the correct metabolic flux data corresponding to the first metabolic network).

[0113] Then, data denoising and cleaning are performed on the above simulation data by the technical means described in steps S131 - S132, and the normalized first simulation data is input into the prediction network model.

[0114] In step S203, according to the prediction result of the initial prediction network model and the prediction label in the first sample data, the prediction result and the prediction label are compared to calculate the first prediction error of the initial prediction network model, and the parameters of the initial prediction network model are adjusted according to the first prediction error until the first prediction error is less than the first error threshold.

[0115] In some embodiments, a Simple Recurrent Neural Network (SRN) model is selected for training. The dataset is divided into a training set and a test set at a ratio of p (p is a positive number):1. During the experiment, the model is trained multiple times in the form of K-fold cross-validation (in this case, 3-fold cross-validation is used). The prediction results are measured by the Mean Absolute Error (MAE). The absolute difference between the prediction result and the prediction label is calculated, and the average of all absolute differences is obtained. The formula:

[0116]

[0117] MAE (Mean Absolute Error) represents the average of the absolute errors between the prediction label and the prediction result in the first sample data, that is, the first prediction error described in the steps. Among them, y(k) is the prediction label, is the prediction result, M is the number of sampling data in the first simulation data, and k represents the k-th sampling data in the first simulation data.

[0118] The parameters of the initial prediction network model are adjusted according to the first prediction error until the first prediction error is less than the first error threshold. The preliminary training of the initial prediction network model is completed.

[0119] In step S204, the second simulation data in the second metabolic network is obtained and input into the initial prediction network model. The second simulation data has corresponding second sample data.

[0120] Exemplarily, first, the target metabolic reaction is determined in the first metabolic network, then the sampling data is determined, and finally, a set number of sampling data is collected as the second simulation data. Each sampling data includes metabolite substrate data (CIT, [1,2,3-13C3]CIT, [U-13C6]CIT not labeled or labeled with carbon-13), isotope isomer distribution data (including mass isotope isomer distribution data at time points of 1s, 100s, 500s, 1000s, 2000s, 4000s, 6000s), and a prediction label (the correct metabolic flux data corresponding to the first metabolic network).

[0121] In step S205, according to the prediction result of the initial prediction network model and the prediction label in the second sample data, the prediction result and the prediction label are compared to calculate the second prediction error of the initial prediction network model, and the parameters of the initial prediction network model are adjusted according to the second prediction error until the second prediction error is less than the second error threshold.

[0122] In some embodiments, a Simple Recurrent Neural Network (SRN) is selected for training. The dataset is divided into a training set and a test set at a ratio of p (p is a positive number):1. During the experiment, the model is trained multiple times in the form of K-fold cross-validation (in this case, 3-fold cross-validation). The prediction results are measured by the Mean Absolute Error (MAE). The absolute difference between the prediction result and the prediction label is calculated, and the average of all absolute differences is taken. The formula:

[0123]

[0124] MAE (Mean Absolute Error) represents the average of the absolute errors between the prediction result and the prediction label. Among them, y(Q) is the prediction label (i.e., the second sample data), is the prediction result, S is the number of sampled data in the second simulation data, and Q represents the Qth sampled data in the second simulation data.

[0125] According to the second prediction error, the parameters of the initially trained initial prediction network model are further adjusted until the second prediction error is reduced to be less than the second error threshold. The final training of the initial prediction network model is completed. Among them, the second error threshold is less than the first error threshold.

[0126] In step S206, the initial prediction network model is used as the prediction network model.

[0127] In the embodiments of the present application, the first metabolic network and the second metabolic network are gradually used as sample data to train the prediction network model, so that the prediction network model can perform accurate metabolic flux prediction for metabolic networks of different sizes.

[0128] In some embodiments, to ensure the training effect on the metabolic network, the input flux values of the first metabolic network and the second metabolic network are respectively within the corresponding input setting intervals; the global flux distributions of the first metabolic network and the second metabolic network are within the corresponding global setting intervals.

[0129] In some embodiments, carbon dioxide is used as the metabolic substrate for the second metabolic network, and its input flux value is randomly sampled between 32.5 and 42.5. The global flux limit is set to [-300, 300].

[0130] Please refer to Figure 6 , Figure 6 which shows a schematic diagram of the hyperparameter settings of the initial prediction neural network according to an embodiment of the present application. To screen out the most suitable hyperparameters for the initial prediction neural network. Please refer to Figure 7 ,Figure 7 A comparison chart of the errors of various hyperparameters according to an embodiment of the present application is shown. The prediction network model uses a Simple Recurrent Neural Network (SRN). The error comparison is selected for the initial prediction network model. It should be clear that when using the first metabolic network and the second metabolic network, the hyperparameters of the initial prediction network model are determined in the following manner. The "error rate" is used to represent the error size of the initial prediction network model. The larger the error rate, the less accurate the initial prediction network model. That is, it is most appropriate that the number of hidden layers of the initial prediction network model (the number of hidden layers is one of the hyperparameters of the initial prediction network model) is 3 layers.

[0131] Please refer to Figure 8 , Figure 8 A comparison diagram of the activation function settings of the initial prediction neural network according to an embodiment of the present application is shown. To screen out the most suitable activation function for the initial prediction neural network. As Figure 8 can be seen, through experiments, it is shown that when the ReLU function, Leaky ReLU function, Tanh function, and Sigmoid function are used as the activation functions of the initial prediction network model, the error rate and consumption time of the initial prediction network model are within an acceptable range, and the experiment is feasible. Different activation functions can be set according to the different needs of users. For example, if the user needs high precision, the ReLU function is used as the activation function; if the user needs fast prediction, the Leaky ReLU function is used as the activation function.

[0132] In some embodiments, in order to explore the neural network type, number of neural network hidden layers, optimization algorithm, activation function, and loss function required for training the prediction network model, the following verification is carried out in the laboratory to determine the optimal embodiment and prove the timeliness and accuracy of the prediction network model.

[0133] First, determine the neural network type and the number of neural network hidden layers required for the initial prediction network model:

[0134] To compare the performance between different recurrent neural network models, three recurrent neural network models, namely, the traditional recurrent neural network (SRN), Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU), are respectively constructed for prediction. To determine the neural network type required for the prediction network model. The preset hyperparameters of the model are as Figure 9 shown.

[0135] For the prediction network model studied in the embodiments of the present application, according to Figure 9Debug the network model with the hyperparameters to be selected. The performance of the models trained with different hyperparameters is as Figure 9 , Figure 10 shown. As the scale of the dataset increases, the training time of the model increases significantly. At this time, not only the prediction performance of the model needs to be considered, but also whether the time cost is within an acceptable range. The neural network structure and experimental results used in the experiment are as Figure 10 shown. Among them, Figure 9 shows a schematic diagram of hyperparameters of different neural networks as the initial prediction network model according to an embodiment of the present application. Figure 10 shows the training results of different network models and different hyperparameters according to an embodiment of the present application.

[0136] Through analysis, it can be seen that the SRN model has better comprehensive performance in this research project, and the time cost of the model is within an acceptable range. By changing other hyperparameters, the optimal model of SRN in this experiment is obtained, and the optimal number of hidden layers is 3 layers.

[0137] Secondly, after determining the type of neural network used in the prediction network model, determine the influence of the activation function, optimization algorithm, and loss function required for the prediction network model on the model performance.

[0138] The prediction results and training time of different models are as Figure 11 , 12 shown, and the hyperparameter settings of the model are shown in Tables 11 and 12. Figure 11 shows a schematic diagram of hyperparameter settings of the model when different activation functions are selected according to an embodiment of the present application. Figure 12 shows the hyperparameter settings of the model when different optimization algorithms are selected according to an embodiment of the present application.

[0139] Please continue to refer to Figure 13 , 14 , Figure 15 , Figure 13 shows a comparison chart of the error rate and training time of different activation functions according to an embodiment of the present application. Figure 13 In the histogram in , the ordinate is the loss rate, and the ordinate of the broken line represents the time required for training. Figure 13 It is composed of a bar chart and a broken line chart. The title of the abscissa is the type of activation function. The abscissa is divided into four groups in total. Each group represents an activation function. There are four activation functions in total, namely Tanh activation function, Sigmoid activation function, ReLu activation function. The ordinate in the figure is divided into two types on the left and right. The left ordinate represents the error rate (unit: %) of the model when using this activation function, which is represented by the column in the figure. The right ordinate represents the training time (unit: h) of the model, which is represented by the broken line in the figure.

[0140] Figure 14 It shows a comparison chart of the error rate and training time of different optimization algorithms according to an embodiment of the present application. Figure 14 In the histogram, the vertical axis represents the error rate, and the vertical axis of the broken line represents the time required for training. Figure 14 It is composed of a bar chart and a line chart. The title of the horizontal axis is the type of optimization algorithm. The horizontal axis is divided into three groups, each group represents an optimization algorithm, and there are three optimization algorithms in total, namely the Gradient Descent optimization algorithm (also called the gradient descent optimization algorithm), the AdaDelta optimization algorithm, and the Adam optimization algorithm.

[0141] Figure 15 It shows a comparison chart of the model performance using two loss functions, MAE and MSE, according to an embodiment of the present application.

[0142] Among them, according to the figure, in some embodiments, the neural network type used in the initial prediction network model is a simple recurrent neural network model (Simple Recurrent Neural Network, SRN). The activation function used is ReLU, and the optimization algorithm used is Adam.

[0143] According to the user's selection, the loss function used during training is MAE or MSE. If high accuracy is required, MSE is used as the loss function. If fast training speed is required, MAE is used as the loss function.

[0144] Similarly, for the prediction network model obtained through training, the neural network type used is a simple recurrent neural network model (Simple Recurrent Neural Network, SRN). The activation function used is ReLU, the optimization algorithm used is Adam, and the number of hidden layers is 2. This can make the prediction of the prediction network model the most efficient and accurate.

[0145] Please refer to Figure 16 , Figure 16 It shows a schematic diagram of the evaluation of the prediction network model used in the present application according to an embodiment of the present application. From Figure 16 it can be seen that compared with the existing technical means, the embodiment of the present application has significant technical effects, greatly improving the prediction efficiency and prediction accuracy of metabolic flux.

[0146] Figure 17 It shows a block diagram of the computer system for implementing the method for predicting biological metabolic flux according to an embodiment of the present application.

[0147] It should be noted that Figure 17 the computer system 800 shown is only an example and should not bring any idle to the functions and usage scope of the embodiments of the present application.

[0148] As Figure 17 shown, the computer system 800 includes a central processing unit 801 (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory 802 (ROM) or a program loaded from a storage section 808 into a random access memory 803 (RAM). In the random access memory 803, various programs and data required for system operation are also stored. The central processing unit 801, the read-only memory 802, and the random access memory 803 are connected to each other via a bus 804. An input / output interface 805 (Input / Output interface, i.e., I / O interface) is also connected to the bus 804.

[0149] The following components are connected to the input / output interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as required. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as required so that a computer program read from it can be installed into the storage section 808 as required.

[0150] Specifically, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit 801, various functions defined in the system of the present application are executed.

[0151] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0153] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0154] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0155] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0156] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only defined by the appended claims.

Claims

1. A method for predicting the metabolic flux of an organism, characterized in that Including: Determine the metabolic network of the organism to be predicted in response to the instruction information; Generate data simulation for the metabolic network to obtain simulation data, where the simulation data is used to represent the flux distribution of the metabolic network; Input the simulation data into the prediction network model to obtain the predicted metabolic flux of the metabolic network output by the prediction network model.

2. The method according to claim 1, characterized in that, Before generating data simulation for the metabolic network to obtain simulation data, the method further includes: Determine the isotope labeling object and isotope labeling substrate in response to the instruction information.

3. The method according to claim 2, wherein The metabolic network includes at least one metabolic reaction; generating data simulation for the metabolic network to obtain simulation data includes: Select a set proportion of metabolic reactions in all metabolic networks as target metabolic reactions; When the duration of the target metabolic reaction reaches the set duration and the flux value of the target metabolic reaction is within the set flux range, obtain the flux value of the target metabolic reaction as the sampling flux value of the target metabolic reaction at the set duration, so that the sampling flux value can be used as the non-equilibrium flux distribution of the target metabolic reaction; Simulate the isotope distribution in the metabolic network according to each initial sampling flux value to obtain the simulation data.

4. The method according to claim 3, characterized in that, Simulating the isotope distribution in the metabolic network according to each initial sampling flux value to obtain the simulation data includes: Obtain a set number of sampling flux values, and simulate the isotope distribution in the metabolic network according to the sampling flux values to obtain the sampling data corresponding to the sampling flux values, so as to obtain the simulation data. The sampling flux values include metabolic substrate data and isotope isomer distribution data when the target metabolic reaction reaches each set duration.

5. The method according to claim 4, characterized in that, Inputting the simulation data into the prediction network model includes: According to the metabolic substrate data and isotope isomer distribution data in the simulation data, remove noise and transform the simulation data to obtain standardized simulation data; Input the standardized simulation data into the network model.

6. The method according to claim 1, wherein The method for training the network model includes: According to the number of metabolic reactions in each metabolic network, select a first metabolic network and a second metabolic network in each metabolic network, where the number of metabolic reactions in the first metabolic network is less than that in the second metabolic network; Obtain the first simulation data in the first metabolic network and input the first simulation data into the initial prediction network model. The first simulation data has corresponding first sample data; According to the prediction result of the initial prediction network model and the first sample data, calculate the first prediction error of the initial prediction network model, and adjust the parameters of the initial prediction network model according to the first prediction error until the first prediction error is less than the first error threshold; Obtain the second simulation data in the second metabolic network and input the second simulation data into the initial prediction network model. The second simulation data has corresponding second sample data; Calculate the second prediction error of the initial prediction network model according to the prediction result of the initial prediction network model and the second sample data, and adjust the parameters of the initial prediction network model according to the second prediction error until the second prediction error is less than the second error threshold, where the second error threshold is less than the first error threshold. Use the initial prediction network model as the prediction network model.

7. The method according to claim 5, characterized in that The input flux values of the first metabolic network and the second metabolic network are respectively within the corresponding input setting intervals; the global flux splits of the first metabolic network and the second metabolic network are within the corresponding global setting intervals.

8. A computer device, comprising a memory, a processor, and a readable program stored on the memory, characterized in that, The processor executes the readable program to implement the method according to any one of claims 1-7.

9. A readable storage medium, characterized in that, Stored thereon are readable programs / instructions which, when executed by a processor, implement the method according to any one of claims 1 to 7.

10. A program product, comprising readable programs / instructions, characterized in that, When the readable program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.