System and method for comparative interactive genome scale flux balance analysis

An interactive human-computer interface using FDSeM and RiG addresses the lack of benchmarking methods for organism comparisons, enabling optimized cell factory design and high-yield production by selecting organisms with minimal impurities.

WO2025225233A1PCT designated stage Publication Date: 2025-10-30YOKOGAWA ELECTRIC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/011026
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2025-03-21
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current methods fail to provide a benchmark for comparing organisms based on their metabolic capacity, and there is no quantitative method to evaluate organisms based on their metabolic capacity, which limits the ability to optimize their growth parameters, thus preventing the scaling up of production in large batches.

Method used

The development of an interactive human-computer interface that allows for the formulation of appropriate objective functions for cell factory design, utilizing flux-derived demand-supply exchange (FDSeM) and reaction interaction graphs (RiG) to select organisms with high yields and minimal impurities, and optimize growth through genetic engineering.

Benefits of technology

Enables quantitative comparisons and optimizations of organisms for high-yield production of target compounds by providing interpretability in simulation outputs, allowing for the selection of organisms with improved productivity and purity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025011026_30102025_PF_FP_ABST
    Figure JP2025011026_30102025_PF_FP_ABST
Patent Text Reader

Abstract

One example of a method according to the present invention includes a step for loading a model associated with a first organism. The method further includes: a step for defining a molar relationship including the stoichiometry between one or more intermediate metabolites and a target substance; a step for defining a first ratio-based coefficient that establishes a relationship between a production rate of the target substance and a consumption rate of consumed substances on the basis of the molar relationship and one or more reaction pathways from among a plurality of reaction pathways; a step for performing a first biclustering operation in order to group two or more of the plurality of reaction pathways associated with the target substance; and a step for optimizing the first ratio-based coefficient by imposing one or more constraints on the production of the target substance on the basis of the first biclustering operation.
Need to check novelty before this filing date? Find Prior Art

Description

Systems and methods for comparative interactive genome-scale flux balance analysis

[0001] The present disclosure relates generally to simulation systems, and more particularly to simulation systems that enable comparative interactive genome-scale flux balance analysis (FBA).

[0002] The biofoundry industry is being developed on organisms that produce high-yield, high-value products. To make the bioindustry successful, the existing potential of organisms to produce bioactive compounds, industrial precursor compounds, and green alternatives needs to be analyzed. The potential and range of applications of organisms can be further expanded through genetic engineering and optimization of growth parameters, which allows for the scaling up of production in large batches.

[0003] There are two major challenges in this field. Currently, there is no benchmarking method to compare organisms based on their metabolic capacity (to the best of our knowledge, no quantitative method exists). Furthermore, experts in various fields require different "cell design objective functions" for target-compound-based optimization. These must be tailored to the application, such as obtaining the highest yield without affecting organism viability, balancing product yield with growth rate, or introducing minimal engineering into the organism. The cell design objective function is crucial for the success of scaling up the growth of the organism and optimizing the yield of the target compound. However, the cell design objective function for strain improvement is complex and cannot be predicted a priori.

[0004] Genome-scale metabolic models (GSMMs) are unique to each strain and organism. GSMMs are complex structures that capture information about all metabolites present in an organism, reactions occurring within the organism, and the genes associated with those reactions. A properly curated GSMM and experimentally derived parameters can be used as inputs to run simulations using various algorithms to optimize the yield of a target compound. This is an iterative process that is performed until a suitable objective is derived. Therefore, SMEs must run multiple iterative simulations. If the simulation output lacks interpretability and the iterations lack automation, the decision-making process becomes a major obstacle. Therefore, manual search for an appropriate cell design objective function does not guarantee the quality of the results. Furthermore, the lack of interpretability in the simulation output makes comparisons between pools of organisms difficult due to the complexity of the GSMM and the simulation results.

[0005] These and other needs are addressed by the various embodiments and configurations of the present disclosure, which can provide several advantages depending on the particular configuration.

[0006] Cell factory design broadly refers to selecting an organism, engineering the organism, and enabling the growth of cells under optimal conditions to function as a "factory" for target compound production. Embodiments of the present disclosure propose an interactive human-computer interface that addresses the formulation of an appropriate objective function for cell factory design. The embodiments disclosed herein enable SMEs to understand the mathematical output from a GSMM simulation system through representations.

[0007] Embodiments of the present disclosure also describe the construction of a flux-derived demand-supply exchange of metabolites (FDSeM) and a reaction interaction graph (RiG), which enables SMEs to a) select appropriate organisms that produce greater yields of target compounds, b) select organisms with minimal impurities for downstream processes, and c) improve the yield of target compounds within organisms while optimizing growth. The FDSeM and experimentally derived parameters are used for simulations (flux balance / flux fluctuation analysis). The FDSeM can include an interactive representation that captures the weighted flux distribution from source metabolites to target metabolites. Additionally, the FDSeM can be used to analyze the productivity of metabolites within organisms. A fingerprint of an organism's metabolite production can be constructed by interacting with the FDSeM and imposing constraints on the ordering of metabolites based on attributes (e.g., the pathways they participate in / their chemical properties) followed by biclustering.

[0008] RiGs can be constructed from sequential reaction topologies with weighted fluxes. RiGs can be supported by gene reaction databases, thus interactively altering gene expression parameters (e.g., overexpression, deletion, etc.) allows SMEs to tailor the production of target compounds.

[0009] FDSeMs and RiGs can be constructed for organisms for which curated GSMMs exist. Aligning FDSeMs / RiGs across organisms involves nontrivial optimizations where reaction equivalence is established and reactions are sorted by pathway / attribute to enable interpretability.

[0010] Embodiments of the present disclosure also provide a mechanism to enable quantitative comparisons by involving matrix biclustering to allow comparisons between organisms.

[0011] Aspects of the present disclosure include using a database to associate reactions in RiG with attributes such as gene-protein rules, pathways, characteristics of the reaction (e.g., growth, survival, toxicity, etc.) Aspects of the present disclosure also include the ability to use the database to interactively impose constraints and order reactions.

[0012] If a major contaminant with a similar chemical profile to the compound of interest is present, the SME can use simulations using genetic modification and RiG to identify advantageous, minimal gene deletions that ensure a better yield of the compound of interest without impurities. One approach to minimize contaminants and ensure ease of extraction of the compound of interest may involve using an in-house database to obtain reactions associated with the contaminant, iteratively deleting genes associated with the reaction (e.g., using a gene-protein association database), and calculating RiG for each deletion. Iterative deletions may be performed to minimize contaminant yield, optimize the yield of the compound of interest, optimize a biomass objective function, and / or minimize toxicity.

[0013] In some embodiments, the techniques described herein relate to a method including: loading a model associated with a first organism, the model representing a plurality of reaction pathways including one or more intermediate metabolites for producing a target substance from a consumed substance in the first organism; defining a molar relationship including a stoichiometry between the one or more intermediate metabolites and the target substance; defining a first ratio-based coefficient that establishes a relationship between a production rate of the target substance and a consumption rate of the consumed substance based on the molar relationship and one or more reaction pathways from the plurality of reaction pathways; performing a first biclustering operation to group two or more of the plurality of reaction pathways associated with the target substance; and optimizing the first ratio-based coefficient by imposing one or more constraints on the production of the target substance based on the first biclustering operation.

[0014] In some aspects, the techniques described herein relate to a system including a processor and a memory device coupled to the processor, wherein the memory device includes data stored on the memory device that, when processed by the processor, enables the processor to: load a model associated with a first organism, the model representing a plurality of reaction pathways including one or more intermediate metabolites for producing a target substance from a consumed substance in the first organism; define a molar relationship including a stoichiometry between the one or more intermediate metabolites and the target substance; define a first ratio-based coefficient that establishes a relationship between the production rate of the target substance and the consumption rate of the consumed substance based on the molar relationship and one or more reaction pathways from the plurality of reaction pathways; perform a first biclustering operation to group two or more of the plurality of reaction pathways associated with the target substance; and optimize the first ratio-based coefficient by imposing one or more constraints on the production of the target substance based on the first biclustering operation.

[0015] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium having processor-executable instructions stored thereon, which, when executed, enable a processor to: load a model associated with a first organism, the model representing a plurality of reaction pathways including one or more intermediate metabolites for producing a target substance from a consumed substance in the first organism; define a molar relationship including a stoichiometry between the one or more intermediate metabolites and the target substance; define a first ratio-based coefficient that establishes a relationship between the production rate of the target substance and the consumption rate of the consumed substance based on the molar relationship and one or more reaction pathways from the plurality of reaction pathways; perform a first biclustering operation to group two or more of the plurality of reaction pathways associated with the target substance; and optimize the first ratio-based coefficient by imposing one or more constraints on the production of the target substance based on the first biclustering operation.

[0016] Aspects of the present disclosure also contemplate one or more means for carrying out any one or more of the above aspects or aspects of the embodiments described herein.

[0017] The foregoing is a simplified summary of the invention to provide an understanding of some aspects of the invention. This summary is not an extensive or exhaustive overview of the invention and its various embodiments. This summary is not intended to identify key or critical elements of the invention or to delineate the scope of the invention, but rather to present selected concepts of the invention in a simplified form as a prelude to the more detailed description that is presented below. It will be understood that other embodiments of the invention can utilize, alone or in combination, one or more of the features noted above or described in detail below. Also, while the present disclosure has been presented in terms of exemplary embodiments, it should be understood that individual aspects of the disclosure can be claimed separately.

[0018] The present disclosure will be described in conjunction with the accompanying figures.

[0019] FIG. 1 illustrates a computing system according to an embodiment of the present disclosure. FIG. 2 illustrates a simulation pipeline according to an embodiment of the present disclosure. FIG. 3 illustrates a process for making simulation data interpretable for cell factory design according to an embodiment of the present disclosure. FIG. 4 illustrates an example of a simulation process according to an embodiment of the present disclosure. FIG. 5 illustrates a simulation process using metabolite flux-derived supply and demand exchange (FDSeM) according to an embodiment of the present disclosure. FIG. 6 illustrates a simulation process using reaction interface graphs (RiG) according to an embodiment of the present disclosure. FIG. 7 illustrates a process for quantifying the production of metabolites in an organism according to an embodiment of the present disclosure. FIG. 8 illustrates another example of a simulation process according to an embodiment of the present disclosure. FIG. 9 illustrates another example of a simulation process according to an embodiment of the present disclosure.

[0020] The following description provides only embodiments and is not intended to limit the scope, applicability, or configuration of the claims. Rather, the following description provides those skilled in the art with an effective description for implementing the embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.

[0021] Where a sub-reference identifier is present in the Figures, any reference in the description including a numeric reference number without an alphabetic sub-reference identifier, when used in the plural, is a reference to any two or more elements having the same reference number. Where such a reference is made in the singular and without identification of a sub-reference identifier, it is a reference to one of the similarly numbered elements but without limitation to the particular one of the referenced elements. Any explicit use herein to the contrary, or providing further limitation or identification, shall take precedence.

[0022] Before describing various embodiments, it is useful to understand the possible meanings of various terms used herein.

[0023] As used herein, an "organism" may refer to a cluster of cells. An organism may contain one or more attributes defined by a genome.

[0024] As used herein, a "cell" may refer to a cluster of pathways. A cell may contain one or more attributes defined by its genome.

[0025] As used herein, a "pathway" may refer to a cluster of reactions. A pathway may include one or more attributes that are defined by its role within a cell (e.g., survival, carbon, metabolism, etc.).

[0026] As used herein, a "reaction subsystem" may refer to a niche cluster of reactions.

[0027] As used herein, a "reaction" may refer to a cluster of metabolites (e.g., inputs and / or outputs). Oxidation, proton transfer, and metabolite transport are some exemplary classes of reactions. An enzyme that facilitates a reaction may be associated with the reaction.

[0028] As used herein, "enzyme" or "protein" may refer to the product of a gene. Attributes of an enzyme or protein may include annotations regarding cellular location, structure, whether it is a transporter, an enzyme, or part of a complex, etc.

[0029] As used herein, a "gene" may refer to a functional segment of a genome. Attributes of a gene may include a gene class.

[0030] As used herein, "nutrients" may refer to elements supplied to cells. Attributes of "nutrients" may include growth media in which components are quantified.

[0031] As used herein, "metabolites" may refer to nutrients and outputs or products of internal reactions. Attributes of metabolites may include internal molecules of a cell, as well as chemical attributes such as polarity, hydrophobicity, and / or toxicity.

[0032] As used herein, a "target compound" or "target substance" may refer to an element to be extracted from a reaction. A target compound or target substance may correspond to a waste product (e.g., a by-product) for a cell, but may be considered valuable to industry.

[0033] As used herein, the term "flux" may refer to a rate that is altered or optimized.

[0034] 1 illustrates an exemplary computing system 100 according to an embodiment of the present disclosure. As discussed in further detail herein, computing system 100 may be utilized to perform one or more simulations that facilitate decisions in cell factory design and / or optimization.

[0035] In one embodiment, computing system 100 may include a computing device 102 with various components and connections to other components and / or computing devices for performing certain embodiments described herein. The components of computing device 102 may be variously embodied and may include a processor 104 and memory 106. As used herein, the term "processor" may refer to any suitable type of processing and / or microprocessing device. By way of example, and without limitation, processor 104 may include an integrated circuit (IC) chip, a microprocessor, a central processing unit (CPU), and a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a semiconductor device, combinations thereof, etc.

[0036] Processor 104 may include programmable logic functions as determined at least in part from accessing machine-readable instructions held in non-transitory data storage, which may be embodied as circuitry, on-chip read-only memory, computer memory 106, data storage 108, etc., that cause processor 104 to perform steps or processes in accordance with the instructions. Processor 104 may be embodied as a single electronic microprocessor or a multi-processor device (e.g., multi-core) having electronic circuitry therein that may further include control units, input / output units, arithmetic logic units, registers, primary memory, and / or other components that access information (e.g., data, instructions, etc.) as received via bus 114, execute instructions, and output data, again via bus 114, etc.

[0037] In some embodiments, processor 104 may comprise a shared processing device that can be utilized by other processes and / or process owners, such as within a processing array in a system (e.g., a blade, a multiprocessor board, etc.) or a distributed processing system (e.g., a “cloud” farm, etc.). It should be understood that processor 104 is a non-transient computing device (e.g., an electronic machine with circuitry and connections for communicating with other components and devices). Processor 104 may operate a virtual processor, such as to process machine instructions that are not native to the processor (e.g., translating a VAX operating system and VAX machine instruction code set into Intel® 9xx chipset code to enable a VAX-specific application to run on a virtual VAX processor). However, as one skilled in the art will appreciate, such a virtual processor is hardware, more specifically, the electrical circuitry underlying the processor (e.g., processor 104) and applications executed by other hardware. Processor 104 may alternatively or additionally be executed by a virtual processor, such as when an application (i.e., a Pod) is orchestrated by Kubernetes. A virtual processor allows an application to be presented with what appears to be a static and / or dedicated processor that executes the application's instructions, while the underlying non-virtual processor is executing the instructions and may be dynamic and / or divided among several processors.

[0038] In addition to the processor 104 components, the computing device 102 may utilize computer memory 106 and / or data storage 108 for storage of accessible data, such as instructions, values, etc. Examples of data and / or instructions that may be stored in the memory 106 include, but are not limited to, one or more GSM models 116 and one or more simulation instruction sets 118. As described in further detail herein, the processor 104 may be configured to access the simulation instruction set 118, which loads the GSM model 116 as part of performing a simulation or multiple simulations to aid in analyzing target substance production from an organism given a set of constraints.

[0039] Computing device 102 may further include a communication interface 110 that facilitates communication with other devices in system 100. As shown, communication interface 110 may provide connectivity between computing device 102 and communication networks 120, 124. Communication networks 120, 124 may provide machine-to-machine connectivity and communication, thereby allowing a user to remotely access components of computing device 102. Communication interface 110 may be embodied as a network port, card, cable, or other configured hardware device.

[0040] Additionally or alternatively, human input / output interface 112 connects to one or more interface components to receive information (e.g., instructions, data, values, etc.) from humans and / or electronic devices and / or present information (e.g., instructions, data, values, etc.) to humans and / or electronic devices. Examples of input / output devices 130 that may be connected to the input / output interface include, but are not limited to, a keyboard, a mouse, a trackball, a printer, a display, sensors, switches, relays, speakers, microphones, still and / or video cameras, etc. In another embodiment, communication interface 110 may comprise or be configured by human input / output interface 112. Communication interface 110 may be configured to communicate directly with networked components or may be configured to utilize one or more networks, such as network 120 and / or network 140.

[0041] Network 120 may include a wired network (e.g., Ethernet), a wireless (e.g., WiFi, Bluetooth, cellular, etc.) network, or a combination thereof, and may enable computing device 102 to communicate with networked components 122. In other embodiments, network 120 may be embodied, in whole or in part, as a telephone network (e.g., a public switched telephone network (PSTN), a private branch exchange (PBX), a cellular network, etc.).

[0042] Additionally or alternatively, one or more other networks may be utilized. For example, network 124 may represent a second network that may facilitate communication with components utilized by computing device 102. For example, network 124 may be an internal network of a business or other organization whereby the components are more trusted (or at least more trusted) than networked components 122 that may be connected to network 120, including public networks (e.g., the Internet) that may be less trusted.

[0043] Components attached to network 124 may include computer memory 126, data storage 128, input / output devices 130, and / or other components that may be accessible to processor 104. For example, computer memory 126 and / or data storage 128 may supplement or replace computer memory 106 and / or data storage 108, either entirely or for a particular task or purpose. As another example, computer memory 126 and / or data storage 128 may be an external data repository (e.g., a server farm, array, “cloud,” etc.) that may allow computing device 102 and / or other devices to access its data. Similarly, input / output devices 130 may be accessed by processor 104 directly via human input / output interface 112 and / or communication interface 110, via network 124, via network 120 alone (not shown), or via networks 120 and 124. Each of computer memory 106, data storage 108, computer memory 126, and data storage 128 may comprise non-transitory data storage comprising a data storage device.

[0044] It should be understood that computer-readable data may be transmitted, received, stored, processed, and presented by various components. It should also be understood that illustrated components may control other components, whether shown herein or otherwise. For example, one input / output device 130 may be a router, switch, port, or other communication component, such that a particular output of processor 104 enables (or disables) input / output device 130 that may be associated with network 120 and / or network 124 to permit (or prohibit) communication between two or more nodes on network 120 and / or network 124. Those skilled in the art will understand that other communication equipment may be utilized in addition to or instead of those described herein without departing from the scope of the embodiments.

[0045] 2 , additional details of a simulation pipeline 200 will be described, according to at least some embodiments of the present disclosure. The simulation pipeline 200 may be implemented, in whole or in part, by the processor 104 executing a simulation instruction set 118 stored in the memory 106. In some embodiments, the simulation instruction set 118 may be included as part of a simulation software package 204 that receives one or more inputs 212 and generates one or more outputs 216 based on processing of the one or more inputs 212. In some embodiments, a graphical user interface (GUI) for the simulation software package 204 may be presented via a display device, such as the input / output devices 130. The presentation of the GUI via the display device may enable a user 208 to define one or more inputs 212 for the simulation, define one or more constraints to impose on the simulation, and / or view one or more outputs 216 generated by the simulation software package 204.

[0046] According to at least some embodiments, the simulation software package 204 provides the capability for human-computer interaction. By interacting with the simulation software package 204, a user 208 can execute the simulation pipeline 200 to address the formulation of complex "objective functions" for biological optimization and cell factory design. The simulation software package 204 is shown to include: a) an exploration module for interactively loading genome-scale metabolic models and inputting initial constraints; b) a simulation module for running simulations with a desired starting goal (production of a target compound); c) a data visualization module for making simulation data interpretable and / or visible to the user 208 based on the objective function; and d) an iteration module that can specify further simulations for experimental design.

[0047] The simulation software package 204 can be configured to receive several different inputs to support the simulation processes disclosed herein. Examples of inputs 212 that can be provided to the simulation software package 204 include, but are not limited to, GSM models 116, gene expression data, protein abundance data, enzyme activity values, 13-C labeling data, and gene regulatory network data.

[0048] Based on processing of the inputs 212, the simulation software package 204 may generate one or more outputs 216. The outputs 216 may be provided to another computing device for further processing and / or may be presented to a user 208 via a display device as described herein. Examples of outputs 216 that may be provided by the simulation software package 204 include, but are not limited to, pathway selection for targeted genetic mutation, organism selection to optimize a product of interest, and / or the search for an ideal host organism for expressing a non-natural product.

[0049] As seen in FIG. 3 , a process 300 for making simulation data interpretable for cell factory design is shown. The process 300 may include running a GSMM simulation 304 (e.g., using a simulation software package 204). The GSMM simulation 304 may generate one or more outputs 216, including an FDSeM 308 presentation and / or an RiG 312 presentation. A user 208 may be able to update 316 the GSMM simulation 304 through several simulation iterations 320. Each iteration 320 may generate a different set of outputs, which ultimately leads the user 208 to a solution 324 for the cell factory design.

[0050] It should be understood that the FDSeM 308 or the RiG 312 may generate output 216 that is ultimately adopted by the user 208 as the solution 324. In some embodiments, output from multiple simulations, including output from the FDSeM 308 and / or the RiG 312, may generate the solution 324.

[0051] FIG. 4 shows additional details of a simulation process utilizing one or more GSMMs 116. According to at least some embodiments, the GSMMs 116 may include data 404 describing consumed substances, data 408 describing reactions, data 412 describing intermediate metabolites, and data 416 describing produced substances. The GSMMs 116 may also include data 420 describing gene-reaction relationships. In some embodiments, data from the GSMMs 116 may be used to simulate the production of target substances and normalize them for comparison. Using the simulation process described herein with the GSMMs 116 may also be useful for deriving ratiometric flux relationships between produced and consumed substances, studying the impact of modified gene expression on reactions / fluxes resulting in produced substances, and biclustering organisms with constraints to maximize metabolite production.

[0052] In some embodiments, the simulation process may utilize GSMM 116 to define molar relationships (424), define ratiometric relationships between metabolites (e.g., produced and / or consumed) (428), perform biclustering on the ratiometric relationships (432), and optimize production of a target substance given a set of constraints (436). In some embodiments, the simulation process may utilize GSMM 116 to modify gene expression (440), determine flux changes due to modified gene expression (444), and optimize production of a target substance given modified gene expression (448). Additional details of both simulation processes are provided below.

[0053] Simulations utilizing FDSeM 308 can represent weighted fluxes due to nutrient uptake in the medium. In some embodiments, a bipartite reaction network 504 can be generated from GSMM 116. Additionally, weighted fluxes can be represented in a source-target metabolite matrix 508. Internal fluxes are calculated for each reaction within the organism, and FDSeM 308 generates an output 512 that allows assessment of the net depletion and net production of metabolites. Net production of metabolites can be analyzed by summing columns of FDSeM 308.

[0054] In some embodiments, an organism can be considered an industrial complex. A single cell of an organism can be considered a factory. A single cell of an organism consumes raw materials (nutrients) so that energy is produced to support growth, development, and life processes. Reactions within a cell involve several entities: input metabolites, enzymes that act as catalysts for the reaction, and output metabolites that are created at the end of the reaction. Reactions are also related to pathways that enable life processes. Pathways are like divisions of a factory, and reactions are grouped together to form processes under a pathway.

[0055] Because cell factories have order and hierarchy, a series of steps can be used to understand the workings of the factory. In some embodiments, GSMM 116 captures a partial blueprint of the factory. GSMM 116 can store data describing the nutrients consumed by the cells, how those nutrients are processed, how energy is produced, all the processes that allow the cell to survive, other participants within the cell such as various metabolites (with different chemical properties), protein enzymes and transporters that use the metabolites / nutrients, and / or reactions involving metabolites and enzymes.

[0056] Organizing information about metabolites and enzymes in the analysis provided in FIG. 5 helps the user 208 understand the cell factory as a whole. In particular, metabolites can be associated with specific pathways, as shown in network 504. Metabolites also participate in reactions based on their participation in standard formats (e.g., stoichiometry matrix 508). Metabolites can be related to each other based on whether they function as raw materials to become another metabolite (e.g., serve as input to a reaction). A non-limiting example is A+B->C+D, which means that A, B can be directionally related to C and B. This understanding allows for the construction of a matrix 512 in which input elements are arranged on one side and output elements on another side.

[0057] Another case is A+B<->C+D, which means that A, B are related to C and B. The rate at which C, D are driven to form may differ from the back-reactions involved in forming A, B. This nuance can also be captured in the stoichiometry matrix 508. The matrix 512 relating cellular metabolites also captures information about the steady-state rates (flux) of metabolite consumption and formation. In some embodiments, each cell of the matrix 512 is filled with the weighted contribution of the input metabolite flux to the output metabolite flux.

[0058] The output matrix 512 of the FDSeM 308 can be further analyzed because the input metabolites can be arranged in rows and the output metabolites can be arranged in columns. Specifically, the row values ​​can be analyzed to estimate total consumption, and the column values ​​can be analyzed to estimate total production. Additionally, the input feedstock concentrations and steady-state (production and consumption) rates can be summed to describe the net production of each metabolite. This allows the user 208 to analyze the production status.

[0059] At this stage, FDSeM308 captures everything from how the raw materials are utilized after being loaded into the cell factory (e.g., uptake into cells) to how the raw materials are converted into precursors (e.g., intermediate metabolites) to produce products (e.g., the target metabolite of interest).

[0060] Figure 6 provides another mechanism for analyzing different layers of a factory blueprint. In particular, it may be desirable to analyze how processes are connected to one another. As an example, it may be useful to distinguish between fast, efficient processes and slow processes. It may also be useful to determine how processes are connected versus which processes are isolated. As seen in Figure 6, reactions (as well as metabolites) belong to a particular pathway. Many reactions in turn make up a process, and many processes make up a pathway. Such information is stored in GSMM 116 and / or some other data structure stored in data storage 108. Pathway information and reaction information may be organized into a reaction network 604. In the reaction network 604, reactions are immediately connected to one another if they share a common participating metabolite (input or output). In the reaction network 604, a matrix of reactions 608 can be generated by ordering precursor reactions in rows and successive reactions in columns. Steady-state rates can be converted to ratiometric coefficients and used to fill the cells of the matrix of reactions 608. The matrix of reactions 608 may also be referred to as RiG312 or its output. In RiG312 (or the matrix of reactions 608), all connected reactions have values ​​in their cells, but unconnected reactions may provide null cells. Thus, the matrix of reactions 608 or RiG312 may be considered a sparse matrix.

[0061] Reactions are processes, and enzymes are the machines that facilitate the process. The enzymes themselves are produced from genes (higher levels of control). Genes are the controllers for each machine. In some embodiments, the user 208 can turn off the master controller (gene) and shut down a specific machine (enzyme), which will interrupt that specific process (reaction). Of course, this will affect neighboring processes. The user 208 can also increase the production rate by a specific machine by amplifying the controller (gene), which will lead to overproduction of a minority raw material and change the flux of production. RiG312 interactively allows the user 208 to turn off (delete) or amplify (overexpress) controllers (genes). This allows the user 208 to simulate and recalculate the process effects resulting from any of the changes described above. In some embodiments, a database can be used to associate genes with enzymes in the backend of RiG312.

[0062] As can be seen, if the product of a source reaction becomes the input of a target reaction, a link exists between the source and target reactions. Thus, the network topology described herein captures the reaction flow of an organism. Furthermore, the weighted fluxes input into RiG312 provide estimates of high- and low-flow pathways that may lead to the product of interest. RiG312 is also associated with a gene-protein rule database in which genes that control enzymes for reactions are stored.

[0063] Figures 5 and 6 highlight how FDSeM308 and RiG312 can be used for simulations where a user 208 wishes to optimize the productivity of a single cell. In some embodiments, FDSeM308 and RiG312 can be used to compare the productivity of two different organisms with different genomes and reactions, but a common product of interest. Additional details of this approach are described with reference to Figure 7, which shows a unique layout of processes (reactions, Ri) and raw materials that are converted into products (metabolites, mi).

[0064] In organism A, there is a single pathway within reaction network 704a to produce compound of interest m5. Similarly, in organism B, there is a single pathway within reaction network 704b to produce the same compound of interest m5. However, in organism B, in addition to making m5, there is also a process that uses up raw materials for some other compound. There are layout differences in each organism's blueprint. Sometimes, there may be more than one route to create a product, such as m5. Thus, embodiments of the present disclosure contemplate the ability to have a flexible platform for comparing two organisms, their constituent processes, and productivity.

[0065] A determination can be made by comparing the output 708a, 708b of each organism and the layout in the FDSeM 712a, 712b of each organism. Here, the shunting of raw materials can be represented without directly representing the process. This makes the analysis more straightforward, less complicated, and easier for the user 208 to understand.

[0066] To make the comparison interpretable, the following process can be performed: a) Select one of the organisms being compared as a reference; b) Use a database to annotate raw materials (mi) based on chemical attributes such as formula, hydrophobicity, etc.; c) Specify attributes (process-related or chemical properties) by which raw materials are sorted. This enriches the neighborhood of the matrix so that relevant mi are identified and separated from unrelated mi. d) Calculate the FDSeM using the recipe for reference organism A; e) Subject the FDSeM to a biclustering algorithm. This is optimization 1, in which mi attributes and mi usage are used to sort raw materials. Raw materials that are frequently used in the process are clustered together, and raw materials that are moderately used are clustered together. This is the fingerprint for organism A. f) Next, the template for organism A is used to subject organism B to alignment (optimization 2). g) The raw materials and processes common to organism A and organism B are aligned using the attribute database. h) Optimization is performed to find the best match for aligning FDSeM 712a of organism A with FDSeM 712b of organism B. i) There are numerous processes and raw materials. Without a database, drawing correspondences between components of organism A and those of organism B is difficult to calculate and interpret. Similarly, alignment is complex and must therefore be achieved using optimization. j) The net productivity of all compounds produced in organisms A and B can be compared. This profile is a direct fingerprint and an interpretable comparison between organism A and organism B. This can immediately enable the selection of organisms.

[0067] When no clear choice exists or in complex cases, further decisions may need to be made. Some decisions can be validated after the optimization and alignment of the FDSeMs. In some embodiments, it may be possible to search for all compounds with a similar metabolic profile to the target compound and focus the analysis on them. This approach may allow for evaluation of whether organism A or organism B has a better production ratio between the target compound and impurities. In some embodiments, it may also be possible to impose constraints on the amount of raw material made available to the process. Thus, it becomes possible to simulate and recalculate the FDSeMs 712a, 712b for both organisms. After iterations, the user 208 can choose whether to invest in organism A or organism B.

[0068] According to at least some embodiments, the compound of interest may be naturally produced by many organisms. However, the selection of the ideal or preferred organism for growth expansion may depend on many factors, which complicates the analysis. The original yield of the compound should be taken into consideration. Furthermore, there should be no potential contaminants that are produced in equal proportions to the compound of interest during downstream extraction (purity factor).

[0069] Therefore, embodiments of the present disclosure propose a solution for organism selection that includes: (1) an alignment between two organisms that can impose constraints on the order of metabolites, and (2) quantitative analysis of the production of each compound in the organisms of interest.

[0070] The process of quantifying the production of each metabolite in an organism is now described with reference to FIG. 8 . The inputs for the simulation platform 816 may include an appropriately curated GSMM 804 for the organism, a description of the medium uptake components, and constraints 812 in the form of normalized enzyme activities. In some embodiments, the simulation platform 816 may construct a stoichiometry matrix from the GSMM and experimental data to run simulations using flux balance analysis (FBA) and flux fluctuation analysis (FVA). The output fluxes 820 are used to derive overall reaction weights and construct the FDSeM 824. Rows indicate source (precursor) metabolites, and columns indicate target metabolites. Interpretability with the FDSeM 824 is achieved by ordering the metabolites in the FDSeM based on their attributes. The ordering is non-trivial and requires a metabolite attribute database and an optimization framework. An embodiment of the present disclosure proposes the use of a database of metabolite attributes, i.e., metabolite identifiers (KEGG / MetCyc), SMILES formulas, pathway associations, and chemical properties (mass, charge, polarity, solubility), which can be used to impose constraints while aligning the FDSeMs of organisms.

[0071] In some embodiments, metabolites can be grouped according to constraints, non-limiting examples of which include clustering metabolites belonging to similar pathways, clustering metabolites with similar chemical properties by optimization, and grouping metabolites with similar production profiles by matrix biclustering.

[0072] The process of comparing yields of compounds of interest can be performed by: a) creating an FDSeM for the organisms; b) using the FDSeM of one organism as a reference and using a database to interactively impose an order of metabolites based on pathway occurrence; c) utilizing a global optimization to maximize the proximity of similar metabolites based on the imposed constraints (Optimization 1); d) performing biclustering on the FDSeM of the reference organism to generate a "metabolite fingerprint"; e) performing reference biclustering to align the metabolites of the second organism globally (pathway level) and locally (metabolite ID level) (Optimization 2); and f) summing the column values ​​of weighted fluxes, which allows deriving the scaled production of compounds within the organism.

[0073] A metabolite database can be used to impose constraints on grouping metabolites based on chemical similarity prior to optimization 1. After alignment with organism 2, the production profiles of all metabolites with chemical similarity to the compound of interest can be quantitatively evaluated. The decision in organism selection can be based on (1) the maximum native yield of the compound of interest and / or (2) the minimum amount of contaminating metabolites present within each organism.

[0074] It can also be useful to compare two competing organisms for compound extraction. Once the natural yield of the compound of interest has been quantified within each organism, efforts can be directed toward increasing yield. One way to achieve this is through enzyme deletion or overexpression, which can allow more product to be directed toward making the compound of interest. Deletion / overexpression strategies can be interactively constrained and simulated. Problems arise when a change may benefit yield in silico but be detrimental to the growth and survival of the organism. This presents a complex choice that cannot be determined without a simulation and analysis platform, as described herein.

[0075] According to at least some embodiments, interpretability is possible by using RiG928 to determine appropriate genetic modifications to optimize yield. To help determine appropriate genetic modifications to optimize yield, the following non-limiting, exemplary process is provided: a) Build a post-simulation (FBA / FVA) of RiG828 for one or more organisms to be compared. b) The user 208 evaluates RiG828 and performs gene knockout or overexpression (via a gene-protein rules database) in the simulation software package 204. c) The representation interface then globally groups the survival responses of organism A with toxicity-related responses by retrieving response attributes from the database. d) The RiGs of both organisms are aligned by (i) biclustering the RiGs of the reference organism using weighted flux, and / or (ii) aligning the RiGs of organism B to A using optimization. e) Aligned and updated RiGs are created for both organisms (the alignment method can be by optimization and biclustering). f) A reaction attribute database can be used to identify responses related to survival, growth (summarized as the biomass objective function), and toxicity for both organisms. g) If a genetic modification affects survival and growth responses or promotes toxicity responses, the genetic modification is eliminated and a new iteration is initiated. h) The results of the genetic modifications are compared using the FDSeM productivity graph. Those that enhance the compound of interest while promoting the flux of the biomass objective function are stored for decision-making and ranked from most effective to least effective. i) The user 208 can access genetic modifications of organisms that enhanced the compound of interest without affecting survival.

[0076] If a major contaminant with a similar chemical profile to the compound of interest is present, user 208 can use simulations using genetic modifications and RiG to find advantageous, minimal gene deletions that ensure a better yield of the compound of interest without impurities. According to at least some embodiments, the process of minimizing contaminants and ensuring ease of extraction of the compound of interest can include: a) obtaining reactions associated with the contaminant using a database; b) iteratively deleting genes associated with the reactions (e.g., using a gene-protein association database); c) calculating RiG for each deletion and performing iterative deletions to (i) minimize the yield of the contaminant, (ii) optimize the yield of the compound of interest, (iii) optimize a biomass objective function, and / or (iv) minimize toxicity.

[0077] 9, additional details of a process 900 for selecting an organism will be described in accordance with at least some embodiments of the present disclosure. Process 900 begins by loading a GSMM into the simulation software package 204 (step 904). In some embodiments, the GSMM may include one or more reaction pathways that include one or more intermediate metabolites required to produce a target substance. The GSMM may also include associations between first and second reactions present in the reaction pathway and multiple genes, where the multiple genes are used to produce the target substance.

[0078] The process 900 continues with the simulation software package 204 initially defining a cell design objective function (step 908). The cell design objective function may be defined, at least in part, based on user 208 input provided to the simulation software package 204. The initial cell design objective function may provide an indication of the amount of target compound produced.

[0079] The process 900 may further continue by running a simulation 912 using the GSMM and the objective function (step 912). In some embodiments, the simulation software package 204 may be configured to make the simulation data (e.g., inputs and / or outputs) interpretable to the user 208 based on the objective function (step 916). Embodiments of the present disclosure contemplate multiple mechanisms for making the simulation data interpretable through visualization. One mechanism is FDSeM, and another is RiG. As discussed herein, weighted fluxes due to nutrient uptake in the medium are represented in a source-target metabolite matrix. Internal fluxes can be calculated for each reaction within an organism, and FDSeM allows for evaluation of the net depletion and net production of metabolites. The net production of metabolites can be analyzed by summing columns of the FDSeM. Meanwhile, RiG can be constructed from weighted fluxes from source-target reactions present in the GSMM. A link exists between source and target reactions when the product of a source reaction becomes the input of the target reaction. This network topology thus captures the reaction flow of the organism. Furthermore, the weighted fluxes input to RiG provide estimates of high- and low-flow pathways that may lead to the product of interest. RiG is also linked to a gene-protein rule database, which stores the genes that control the enzymes for the reactions.

[0080] Process 900 may further include determining whether the simulation performed in step 912 satisfies the cell design objective function defined in step 908 (step 920). If the query is answered affirmatively, the organism may be selected (step 924). On the other hand, if the query of step 920 is answered negatively, the input constraints may be modified through the simulation software package 204 (step 928). The cell design objective function may also be updated based on the modified input constraints (step 932). Process 900 may then return to step 908 or 912, where an iteration of the simulation is performed.

[0081] 10 , additional details of another process 1000 for simulating and selecting an organism are described in accordance with at least some embodiments of the present disclosure. The steps of process 1000 may be combined in any order with the steps of any other process shown and described herein. For example, one or more steps of process 1000 may be combined with one or more steps of process 800 and / or process 900 without departing from the scope of the present disclosure.

[0082] Process 1000 begins by obtaining a GSMM for a first organism (step 1004). In some embodiments, the GSMM may be loaded from memory 106, from data storage 108, from memory 126, and / or from data storage 128. The GSMM may include multiple reaction pathways involving one or more intermediate metabolites required to produce a target substance.

[0083] The process 1000 continues by loading the GSMM into the simulation instruction set 118 (step 1008). A molar relationship can then be defined for the first organism using the simulation software package 204 (step 1012). In some embodiments, the molar relationship includes the stoichiometry between the consumed metabolite and the target substance.

[0084] Process 1000 may further include defining a first ratio-based coefficient (step 1016). The first ratio-based coefficient may be defined using the molar relationships and reaction pathways of the GSSM. In some embodiments, the first ratio-based coefficient may establish a relationship between the production rate of the target substance and the consumption rate of the consumer substance.

[0085] Process 1000 may further include performing a first biclustering operation to group two or more of the reaction pathways associated with the target substance (step 1020) and optimizing the first ratio-based coefficient by imposing one or more constraints on the production of the target substance (step 1024). In some embodiments, the first biclustering operation is performed with a constraint that minimizes one or more undesired substances. In some embodiments, the first biclustering operation is performed with an additional constraint to reduce the production of one or more substances having substantially similar chemical characteristics to the one or more undesired substances. The additional constraint may apply to at least one of the solubility and polarity of the one or more undesired substances. In some embodiments, the first biclustering operation is performed with an additional constraint to maintain a ratio of the target substance to the one or more undesired substances at a value greater than or equal to one.

[0086] With respect to the one or more constraints (or additional constraints), the one or more constraints imposed on the production of the target material include at least one of (i) defining a maximum amount of the target material to produce, and (ii) defining a minimum amount of the target material to produce. Alternatively or additionally, the one or more constraints imposed on the production of the target material may include defining an amount of a consumable material available for use in the production of the target material.

[0087] Process 1000 may further include the optional step of performing a second biclustering operation to align second ratio-based coefficients from a second organism with the first ratio-based coefficients (step 1028). It should be understood that the first organism and the second organism may both be capable of producing the target substance.

[0088] The process 1000 may also include displaying one or more simulation outputs (step 1032). In some embodiments, the simulation outputs may correspond to outputs generated by the simulation software package 204 that are rendered on a display device. The outputs may include some data and / or graphical output generated based on processing of the GSSM under defined constraints, as described above.

[0089] 11 , another process 1100 for simulating and selecting an organism will be described in accordance with at least some embodiments of the present disclosure. The steps of process 1100 may be combined in any order with the steps of any other process shown and described herein. For example, one or more steps of process 1100 may be combined with one or more steps of process 800, process 900, and / or process 1000 without departing from the scope of the present disclosure.

[0090] Process 1100 begins by retrieving a GSMM from an appropriate data storage location (e.g., memory and / or database) (step 1104). The GSMM may be loaded from a database of models curated and indexed by organism type. The GSMM may be associated with a first organism and may express multiple reaction pathways, including one or more intermediate metabolites, for producing a target substance from a consumed substance within the first organism. The GSMM may also provide an association between the first and second reactions present in each reaction pathway and one or more genes used to produce the target substance.

[0091] The GSMM is then loaded into the simulation instruction set 118 (step 1108). Process 1100 may also include obtaining a cellular design objective function associated with the first organism and the target substance produced thereby (step 1112). Specifically, without limitation, the cellular design objective function may include one or more constraints on the production of the target substance that can be used to optimize the production of the target substance. In some embodiments, the cellular design objective function includes at least one of a minimization function (e.g., minimizing unwanted by-products) and a maximization function (e.g., maximizing target substance production). In some embodiments, the one or more constraints include constraints on at least one of pathway relationships, chemical properties, metabolite attributes, and / or formulas. The one or more constraints may also include defining the amount of a consumable substance available for use in the production of the target substance.

[0092] The process 1100 may also include determining a first flux flow between the first reaction and the second reaction (step 1116). Both the first reaction and the second reaction may be performed to produce a target substance using a consumer substance.

[0093] Process 1100 may include the optional step of grouping the first and second reactions (step 1120). In some embodiments, the first and second reactions may be grouped using a plurality of survival reactions, where the plurality of survival reactions occur as part of maintaining the first organism. In some embodiments, the first and second reactions may be grouped using a toxicity measure associated with one or more intermediate metabolites.

[0094] Process 1100 may further include modifying the GSMM (step 1124). Modifying the GSMM may include modifying at least one of the associations between the first reaction and the second reaction to result in an altered association, and modifying one or more genes used to produce the target substance. Such an approach may result in an altered GSMM.

[0095] Process 1100 may further include determining a second flux flow between the first reaction and the second reaction using the modified GSMM (step 1128). An optional step in process 1100 may further include summing the first flux flow and the second flux flow as part of obtaining the net production of the metabolite (step 1132). The flux flow can be converted to production by normalizing the concentration in millimoles. The production can be calibrated to the input substrate supplied per unit time.

[0096] The process 1100 may also include optimizing one or both of the target material and the consumer material using the second flux flow (step 1136). As discussed herein, optimization may involve efforts to maximize and / or minimize various outputs of a reaction. Such optimization may be defined with the help of user 208 input provided to the simulation software package 204. Accordingly, the process 1100 may further include displaying one or more outputs of the simulation on a display device (step 1140). The outputs displayed as part of the simulation may allow the user 208 to decide to define one or more inputs and / or one or more reactions to use as part of producing the target material using one or more inputs.

[0097] In the preceding description, the method was described in a particular order for purposes of explanation. It should be understood that in alternative embodiments, the method may be performed in an order different from that described without departing from the scope of the embodiments. It should also be understood that the method described above may be performed as an algorithm executed by hardware components (e.g., circuits) specifically constructed to execute one or more algorithms, or portions thereof, described herein. In another embodiment, the hardware components may include a general-purpose microprocessor (e.g., CPU, GPU) that is initially switched to a dedicated microprocessor. The dedicated microprocessor is then loaded with encoded signals, thereby causing the dedicated microprocessor to maintain machine-readable instructions that enable the microprocessor to read and execute a set of machine-readable instructions derived from the algorithms and / or other instructions described herein. The machine-readable instructions utilized to execute the algorithms, or portions thereof, are not unlimited and utilize a finite instruction set known to the microprocessor. In one or more embodiments, the machine-readable instructions may be encoded within the microprocessor as signals or values ​​in signal-generating components through the selective use of voltages in memory circuits, configurations of switching circuits, and / or specific logic gate circuits. Additionally or alternatively, the machine-readable instructions may be accessible to a microprocessor and may be encoded within a medium or device as a magnetic field, a voltage value, a charge value, a reflective / non-reflective portion, and / or a physical indicator.

[0098] The present invention, in various embodiments, configurations, and aspects, includes components, methods, processes, systems, and / or apparatus substantially as illustrated and described herein, including various embodiments, subcombinations, and subsets thereof. Those skilled in the art will understand how to make and use the present invention after understanding this disclosure. The present invention, in various embodiments, configurations, and aspects, includes providing devices and processes in the absence of items not shown and / or described herein or in its various embodiments, configurations, or aspects, including the absence of items that may have been used in previous devices or processes, for example, to improve performance, achieve ease, and / or reduce cost of implementation.

[0099] The foregoing discussion of the present invention has been presented for purposes of illustration and description. It is not intended to limit the invention to the form(s) disclosed herein. For example, in the foregoing Detailed Description, various features of the invention are grouped together in one or more embodiments, configurations, or aspects for the purpose of streamlining the disclosure. Features of the embodiments, configurations, or aspects of the invention may be combined in alternative embodiments, configurations, or aspects other than those discussed above. This method of disclosure is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment, configuration, or aspect. Thus, the following claims are hereby incorporated into this Detailed Description, with each claim standing on its own as a separate preferred embodiment of the present invention.

[0100] Furthermore, while the description of the invention includes a description of one or more embodiments, configurations, or aspects, and certain variations and modifications, other variations, combinations, and modifications are within the scope of the invention and may be within the skill and knowledge of one of ordinary skill in the art after understanding the present disclosure. The invention is intended to acquire the right to include alternative embodiments, configurations, or aspects to the extent permitted, including alternative, interchangeable, and / or equivalent structures, functions, ranges, or steps to those claimed, whether or not such alternative, interchangeable, and / or equivalent structures, functions, ranges, or steps are disclosed herein, and without intending to dedicate patentable subject matter to the public.

[0101] 100 Computing system 102 Computing device 104 Processor 106 Memory, computer memory 108 Data storage 110 Communication interface 112 Human input / output interface 116 GSM model, GSMM 118 Simulation instruction set 120 Communication network 124 Communication network 126 Computer memory 128 Data storage 130 Input / output device 200 Simulation pipeline 204 Simulation software package 208 User 212 Input 216 Output 304 GSMM simulation 308 FDSeM 312 RiG 320 Iteration 324 Solution 404 Data describing consumed substances 408 Data describing reactions 412 Data describing intermediate metabolites 416 Data describing produced substances 420 Data describing gene-reaction relationships 504 Bipartite reaction network 508 Source-target metabolite matrix, stoichiometry matrix 512 Output, matrix, output matrix 604 Reaction network 608 Matrix of reactions 704a Reaction network 704b Reaction network 708a Output 708b Output 712a FDSeM 712b FDSeM 804 GSMM 812 Constraints on the form of normalized enzyme activity 816 Simulation platform 820 Output flux 824 FDSeM 828 RiG

Claims

1. A method comprising: loading a model associated with a first organism, the model representing a plurality of reaction pathways including one or more intermediate metabolites for producing a target substance from a consumer substance in the first organism; defining a molar relationship including a stoichiometry between the one or more intermediate metabolites and the target substance; defining a first ratio-based coefficient that establishes a relationship between a production rate of the target substance and a consumption rate of the consumer substance based on the molar relationship and one or more reaction pathways from the plurality of reaction pathways; performing a first biclustering operation to group two or more of the plurality of reaction pathways associated with the target substance; and optimizing the first ratio-based coefficient by imposing one or more constraints on the production of the target substance based on the first biclustering operation.

2. The method of claim 1, wherein the model comprises a genome-scale metabolic (GSM) model.

3. The method of claim 1, wherein the one or more constraints imposed on the production of the target material include at least one of (i) defining a maximum amount of the target material to produce, and (ii) defining a minimum amount of the target material to produce.

4. The method of claim 1, wherein the one or more constraints imposed on the production of the target material include defining an amount of the consumable material available for use in the production of the target material.

5. The method of claim 1, further comprising performing a second biclustering operation to align second ratio-based coefficients from a second organism with the first ratio-based coefficients, wherein both the first organism and the second organism are capable of producing the target substance.

6. The method of claim 1, wherein the first biclustering operation is performed with a constraint that minimizes one or more undesirable materials.

7. The method of claim 6, wherein the first biclustering operation is performed with additional constraints to reduce production of one or more substances having substantially similar chemical characteristics to the one or more undesirable substances.

8. The method of claim 7, wherein the additional constraint applies to at least one of the solubility and polarity of the one or more undesirable substances.

9. The method of claim 6, wherein the first biclustering operation is performed with an additional constraint to maintain a ratio of the target material to the one or more undesired materials to a value greater than or equal to 1.

10. A system comprising: a processor; and a memory device coupled to the processor, the memory device including data stored on the memory device that, when processed by the processor, enables the processor to: load a model associated with a first organism, the model representing a plurality of reaction pathways including one or more intermediate metabolites for producing a target substance from a consumed substance in the first organism; define a molar relationship including a stoichiometry between the one or more intermediate metabolites and the target substance; define a first ratio-based coefficient that establishes a relationship between a production rate of the target substance and a consumption rate of the consumed substance based on the molar relationship and one or more reaction pathways from the plurality of reaction pathways; perform a first biclustering operation to group two or more of the plurality of reaction pathways associated with the target substance; and optimize the first ratio-based coefficient by imposing one or more constraints on the production of the target substance based on the first biclustering operation.

11. The system of claim 10, wherein the model comprises a genome-scale metabolic (GSM) model.

12. The system of claim 10, wherein the one or more constraints imposed on the production of the target material include at least one of (i) defining a maximum amount of the target material to produce, and (ii) defining a minimum amount of the target material to produce.

13. The system of claim 10, wherein the one or more constraints imposed on the production of the target material include defining an amount of the consumable material available for use in the production of the target material.

14. The system of claim 10, wherein the data further enables the processor to perform a second biclustering operation to align second ratio-based coefficients from a second organism with the first ratio-based coefficients, wherein both the first organism and the second organism are capable of producing the target substance.

15. The system of claim 10, wherein the first biclustering operation is performed with a constraint that minimizes one or more undesirable materials.

16. The system of claim 15, wherein the first biclustering operation is performed with additional constraints to reduce production of one or more substances having substantially similar chemical characteristics to the one or more undesirable substances.

17. The system of claim 16, wherein the additional constraint applies to at least one of solubility and polarity of the one or more undesirable substances.

18. The system of claim 15, wherein the first biclustering operation is performed with an additional constraint to maintain a ratio of the target material to the one or more undesired materials to a value greater than or equal to one.

19. The system of claim 10, further comprising a user interface that enables display of one or more outputs received from the model and enables a user to define one or more constraints for a simulation.

20. A non-transitory computer-readable medium having processor-executable instructions stored thereon, the instructions, when executed, enabling a processor to: load a model associated with a first organism, the model representing a plurality of reaction pathways including one or more intermediate metabolites for producing a target substance from a consumer substance in the first organism; define a molar relationship including a stoichiometry between the one or more intermediate metabolites and the target substance; define a first ratio-based coefficient that establishes a relationship between a production rate of the target substance and a consumption rate of the consumer substance based on the molar relationship and one or more reaction pathways from the plurality of reaction pathways; perform a first biclustering operation to group two or more of the plurality of reaction pathways associated with the target substance; and optimize the first ratio-based coefficient by imposing one or more constraints on the production of the target substance based on the first biclustering operation.

Citation Information

Patent Citations

  • Biological improvement method using flux sum profiling of metabolites

    JP2009535055A

  • Information analysis system, method, and program

    JP2021016325A

  • How to verify the performance of a culture device

    JP2021534782A

  • Method for in silico Modeling of Gene Product Expression and Metabolism

    US20150127317A1