Method and device for determining combination of microbial strains
Through genome analysis and metabolic information models, combined with potential factor collaborative filtering algorithms, the combination of microbial strains is optimized, which solves the problems of time-consuming and labor-intensive combination and poor reproducibility in existing technologies, and achieves efficient optimization of strain combinations and functional improvement.
Patent Information
- Application Number
- CN202480017595.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2024-03-07
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies consume a lot of time and manpower in selecting and combining microbial strains, and subtle differences in experimental conditions affect the reproducibility of results, making it difficult to effectively consider the cooperative and competitive relationships between individuals in complex community environments.
Through genome analysis, metabolic information and growth index information, reinforcement learning methods and devices are used to determine the combination of microbial strains, including model estimation of genome analysis information, metabolic information and growth index information, combined with latent factor collaborative filtering algorithm and transformer model to optimize the strain combination.
The functionality of microbial strain combinations is improved, the fitness of individuals within the community and the determination of optimal actions are enhanced, and the growth, activity and stability of target microorganisms are improved.
Smart Images

Figure CN120752704A_ABST
Abstract
Description
Technical Field
[0001] Various embodiments of this specification relate to a method for optimizing microbial strain combinations and predicting the growth status of the strains. In particular, the method can be applied to determine specific microbial strain combinations by utilizing genomic analysis, metabolic information, and growth indicator information to optimize productivity, efficiency, or production of specific metabolites. Background Art
[0002] The selection and combination of microbial strains plays a fundamental role in various fields, including bioengineering, pharmaceuticals, environmental engineering, and food engineering. These strains are used for a variety of purposes, such as the production of specific compounds, the breakdown of harmful substances, and the expression of special biological effects. In particular, in industrial fermentation processes, the appropriate combination of microbial strains is essential to produce large quantities of specific substances, and in biological treatment processes, the selection of strains to effectively remove environmental pollutants is crucial.
[0003] Generally speaking, the selection and combination of microbial strains relies heavily on experimental methods. This involves conducting large-scale culture experiments to evaluate the growth rates of various strains, the types and quantities of products produced, and the interactions between the strains. However, these methods are not only time-consuming and labor-intensive, but also have limitations in finding the optimal combination among the numerous possible combinations. Furthermore, subtle differences in experimental conditions can significantly influence the results, leading to problems with reproducibility.
[0004] The above-mentioned background technology is technical information retained by the inventor in order to derive the present invention, or obtained in the process of deriving the present invention. It cannot necessarily be said that it is technology that has been disclosed to the public before the application of the present invention. Summary of the Invention
[0005] Technical issues
[0006] The various embodiments described in this specification are proposed to solve the above-mentioned problems and provide a reinforcement learning method and device that considers relationships such as cooperation and competition between individuals in a complex community environment.
[0007] Technical Solution
[0008] A method for determining a microbial strain combination according to an embodiment of the present specification for achieving the above-mentioned technical problem may include the steps of obtaining genome analysis information related to a target microorganism; obtaining first metabolic information related to each of a plurality of first microbial colonies including the target microorganism; estimating first growth index information related to each of the plurality of first microbial colonies by using the genome analysis information and the metabolic information as inputs of a first model; and determining a strain combination based on at least one of the metabolic information and the growth index information.
[0009] Each of the plurality of first microbial colonies may be composed of any composite strain that combines sequence data of one or more individual strains with genome data of the target microorganism.
[0010] The genome analysis information may include species composition data, predicted gene data, and metabolite data related to the target microorganism.
[0011] The species composition data includes phylogenetic information of the target microorganism, specifically, may include at least one of species classification information, strain classification information, phylogenetic tree information, and phylogenetic distance information of the target microorganism.
[0012] The predicted gene data includes information related to all genes that can be derived from the complete genome sequence of the target microorganism, which may include a list of all coding gene types and / or functional annotation data of the genes. In addition, the predicted gene data may include at least one of the following: metabolic-related gene information, antibiotic-related gene information, toxin gene information, optimal culture medium composition information, and biosynthetic gene group information of the target microorganism.
[0013] The metabolite data may include information on all genes identifiable from the whole genome sequence of the target microorganism and information on metabolites identified by estimating biochemical pathways and metabolic pathways based on information on proteins (enzymes) encoded by the genes.
[0014] The metabolic information may include the results of analyzing metabolic interactions between microorganisms in two or more microbial colonies. The metabolic interaction analysis results are analyzed by considering metabolic resource overlap, metabolic interaction potential, metabolic inconsistency, and / or predicted minimum nutrient diversity for growth between the microorganisms. Specifically, the metabolic interaction analysis results may include at least one of metabolic resource overlap (MRO) and metabolic interaction potential (MIP).
[0015] The MRO is a value obtained by calculating the similarity of metabolites required when each microorganism in a microbial community exists independently. It is an indicator of the degree of competition for given nutrients among microorganisms in the community and can be calculated using the following mathematical formula 1.
[0016] [Mathematical formula 1]
[0017]
[0018] (M i : The minimum amount of nutrients required for the growth of each microorganism i in a community composed of n species)
[0019] The MIP is a value indicating the maximum amount of metabolites that can be exchanged between microorganisms present in a microbial community. It is an indicator of metabolic dependency between microorganisms constituting the community and can be calculated using the following Mathematical Formula 2.
[0020] [Mathematical formula 2]
[0021]
[0022] (M: the minimum total number of metabolites required for the microorganisms that make up the community to exist independently, I: the minimum number of metabolites required for community growth when metabolite exchange is allowed within the community)
[0023] The MRO and MIP can be calculated using the SMETANA tool, and detailed descriptions of the algorithm of the SMETANA analysis tool are described in the document “Metabolic dependency drive species co-occurrence indiverse microbiotic groups (PNAS May 19, 2015 112 (20) 6449-6454)”.
[0024] Growth indicator information refers to information related to indicators that can confirm or evaluate the growth, growth, and / or proliferation level of a microbial colony (composite strain group) comprising a single microorganism or a microbial colony (composite strain group) comprising two or more microorganisms. In one embodiment, the growth indicator information can include one or more selected from the group consisting of optical density (OD) information of microbial growth measured by absorption spectrophotometry, time to maximum OD value (time to maxOD), growth rate, and growth rate during each phase of microbial growth (lag phase, exponential growth phase, stationary phase, and death phase). Any information is acceptable as long as it can confirm whether a microorganism is growing and the extent of its growth. Furthermore, the growth indicator information can include a comparison value [Δ(single / colony)] between the growth indicator information of any single microorganism and the growth indicator information of the complex strain group comprising the single microorganism.
[0025] The first metabolic information may include at least one of MRO data and MIP data.
[0026] The first model may be a regression model that uses the genomic analysis information and the first metabolic information as features.
[0027] The method may further include the steps of obtaining second metabolic information and second growth index information related to each of a plurality of second microbial colonies stored in a database, and determining the strain combination by using the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information as inputs of a second model.
[0028] The second model may include a latent factor collaborative filtering algorithm model, and the step of determining the strain combination may include a step of using the target microorganism and the single strain as users, a step of using third metabolic information and third growth indicator information of data concatenating the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information as items, and a step of determining a microorganism strain combination for at least one of the third metabolic information and the third growth indicator information.
[0029] The step of determining the strain combination may further include the step of re-determining the determined strain combination by determining respective weights of the estimated third metabolic information and the third growth indicator information.
[0030] The second model may include a transformer model, and the step of determining the strain combination may include the step of determining the strain combination based on the second model, wherein the second model learns the relationship between the target microorganism and one or more single strains based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information.
[0031] A method for generating a model for estimating growth index information according to one embodiment of the present specification for achieving the above-mentioned technical problem may include the steps of obtaining genomic analysis information related to a plurality of microorganisms, obtaining metabolic information and growth index information related to each of a plurality of microbial colonies including each of the plurality of microorganisms, generating a data set, the data set using the genomic analysis information and the metabolic information as features and using the growth index information as a label, and generating a model based on the data set, the model estimating growth index information related to each of the plurality of microbial colonies included.
[0032] A computer device according to an embodiment of the present specification for achieving the above technical problem may include a processor, which may obtain genome analysis information related to a target microorganism, obtain first metabolic information related to each of a plurality of first microbial colonies including the target microorganism, estimate first growth index information related to each of the plurality of first microbial colonies by using the genome analysis information and the metabolic information as inputs of a first model, and determine a strain combination based on the growth index information.
[0033] Each of the plurality of first microbial colonies may be composed of any composite strain that combines sequence data of one or more individual strains with genome data of the target microorganism.
[0034] The genome analysis information may include species composition data, predicted gene data, and metabolite data related to the target microorganism.
[0035] The first metabolic information may include at least one of MRO data and MIP data.
[0036] The first model may be a regression model that uses the genomic analysis information and the first metabolic information as features.
[0037] The processor may obtain second metabolic information and second growth index information related to each of a plurality of second microbial colonies stored in a database, and determine the strain combination by using the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information as inputs of a second model.
[0038] The second model may include a latent factor collaborative filtering algorithm model, and the processor may use the target microorganism and the single strain as users, use the third metabolic information and the third growth index information of the data connecting the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information as items, and determine a microorganism strain combination for at least one of the third metabolic information and the third growth index information.
[0039] The processor may re-determine by determining respective weights of the estimated third metabolic information and the third growth indicator information.
[0040] The second model may include a transformer model, and the processor may learn the relationship between the target microorganism and the single strain based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information, and determine a strain combination based on the relationship between the target microorganism and the single strain.
[0041] The microbial strain combination determined or derived using the methods, computer devices, and / or processors of the present invention can be a strain combination used to improve the functionality of a target microorganism. The improved functionality can be achieved by improving the growth, activity (including physiological or pharmacological activity), stability, and / or intestinal colonization of the target microorganism.
[0042] Beneficial effects
[0043] According to an embodiment of the present invention, the best action for each individual suitable for its role and situation can be determined by effectively grasping and reflecting various types of individuals and their characteristics in the community.
[0044] The effects of the present invention are not limited to the above-described effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 FIG. 1 is a diagram schematically illustrating a strain combination determination process of a computer device according to an embodiment of the present invention.
[0046] Figure 2 FIG. 1 is a diagram schematically illustrating a strain combination determination process of a computer device according to an embodiment of the present invention.
[0047] Figure 3FIG. 1 is a flowchart illustrating an operation of a computer device for estimating growth indicator information according to an embodiment of the present invention.
[0048] Figure 4 FIG. 1 is a diagram schematically illustrating a process of obtaining genome analysis information by a computer device according to an embodiment of the present invention.
[0049] Figure 5 FIG. 1 is a diagram schematically illustrating a process of obtaining metabolic information by a computer device according to an embodiment of the present invention.
[0050] Figure 6 A conceptual diagram schematically illustrates a process of obtaining metabolic information by analyzing a composite strain into single strains using a computer device according to an embodiment of the present invention.
[0051] Figure 7 FIG. 1 is a diagram schematically illustrating a process of estimating growth indicator information by a computer device according to an embodiment of the present invention.
[0052] Figure 8 FIG. 1 is a flow chart illustrating an operation of a computer device for determining a strain combination according to an embodiment of the present invention.
[0053] Figure 9 A diagram schematically illustrates a process of determining strain combinations based on collaborative filtering by a computer device according to an embodiment of the present invention.
[0054] Figure 10 FIG. 1 is a diagram schematically illustrating a process of determining a strain combination based on a ranking model by a computer device according to an embodiment of the present invention.
[0055] Figure 11 is a block diagram illustrating the configuration of a computer device according to an embodiment. DETAILED DESCRIPTION
[0056] The terms used in the present invention are only used to describe specific embodiments and are not intended to limit the scope of other embodiments. Unless the context clearly indicates otherwise, a singular expression may include a plural expression. The terms used herein, including technical terms or scientific terms, may have the same meaning as that generally understood by those of ordinary skill in the art described in the present invention. The terms defined in general dictionaries among the terms used in the present invention may be interpreted as having the same or similar meaning as that in the context of the relevant technology, and unless clearly defined in the present invention, should not be interpreted in an idealized or overly formalized sense. Depending on the circumstances, even if a term is defined in the present invention, it cannot be interpreted as excluding embodiments of the present invention.
[0057] Hereinafter, various embodiments will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement the present invention. However, the technical ideas of the present invention can be implemented in various forms and are not limited to the embodiments described in this specification. When describing the embodiments disclosed in this specification, if it is determined that the specific description of the relevant known technology may obscure the main idea of the technical ideas of the present invention, the specific description of the known technology will be omitted. The same or similar components are marked with the same reference numerals, and their repeated descriptions are omitted.
[0058] The term "unit" used in this embodiment refers to a component that performs a specific function executed by software or hardware such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). However, the "unit" is not limited to being executed by software or hardware. The "unit" can exist in the form of data stored in an addressable storage medium, can be implemented through instructions, or can be configured as one or more processors to perform specific functions.
[0059] Software may include computer programs, code, instructions, or a combination of one or more of these, which can be configured to cause a processing device to perform a desired operation or can independently or collectively instruct a processing device. Software and / or data may be embodied, permanently or temporarily, in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave for interpretation by or provision of instructions or data to a processing device. Software may be distributed across networked computer systems and stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media. Software may be read into main memory from another computer-readable medium, such as a data storage device, or from another device via a communication interface. The software instructions stored in main memory may cause the processor to perform the processes or steps described in detail below. Alternatively, processes consistent with the principles of the present invention may be performed using hardwired circuitry instead of or in combination with software instructions. Therefore, embodiments consistent with the principles of the present invention are not limited to any specific combination of hardware circuitry and software.
[0060] The terms used in this application are only used to describe specific embodiments and are not intended to limit the present invention. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. The terms "first", "second", etc. can be used to describe various components, but components should not be limited by these terms. The terms are only used to distinguish one component from another.
[0061] The term "model" as used herein may include any algorithm or methodology for learning or understanding specific patterns or structures from data. Models may include machine learning models such as regression models, decision trees, random forests, support vector machines, K-nearest neighbors, naive Bayes, clustering algorithms, and deep learning models such as neural networks, convolutional neural networks, recurrent neural networks, transformer-based neural networks, generative adversarial networks (GANs), and autoencoders. A "model" may refer to a set of parameters or weights trained to predict or classify an output for a specific input, and the model may be trained using methods such as supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Furthermore, the model encompasses not only a single model but also various learning methods and structures, such as ensemble models, multimodal models, and models using transfer learning. These models can be pre-trained on a different computer device from the one used to predict the output of the input and then used on the other computer device.
[0062] Figure 1 and Figure 2 FIG. 1 is a diagram schematically illustrating a strain combination determination process of a computer device according to an embodiment of the present invention.
[0063] Reference Figure 1 and Figure 2 The computer device may obtain user input 101 from a user. The user input 101 may include the genome sequence data of the target microorganism and the number of constituent strains. The target genome sequence data may include genetic information of the microorganism, and the number of strains may indicate the number of constituent strains used in the strain combination. In other words, the user may input the genome sequence data of the target microorganism used to constitute the colony and the number of constituent strains included in the colony.
[0064] The computer device can obtain single strain sequence data and composite strain sequence data from database 103. The single strain sequence data can indicate the genome sequence data of each single strain. Furthermore, composite strain sequence data can indicate one or more co-occurring groups (morphologies) of two or more microbial species, as well as genome sequence data associated with each of the microorganisms belonging to the co-occurring group. Specifically, the co-occurring group can include a collection of microorganisms isolated or confirmed from various environments in the intestines and / or feces that actually coexist or form colonies. Based on user input 101, the computer device can obtain genome analysis information 113 associated with the target microorganism. As one example, this information can be obtained using a shotgun sequencing process 111. Specifically, the computer device can extract DNA from the target microorganism and, if necessary, amplify the DNA. The computer device can obtain sequences by randomly fragmenting the DNA into thousands or millions of small fragments. Using the obtained sequence data, the computer device can reconstruct the genome sequence and predict the locations of genes in the assembled genome, a list of all coding gene types, and / or functional annotations of these genes. The genome analysis information 113 can include species composition data, predicted gene data, and metabolite data associated with the target microorganism.
[0065] Furthermore, the computer device may obtain metabolic information 117 based on the database 103, user input 101, and the metabolite interaction simulation model 115. This metabolic information may include at least one of metabolic resource overlap (MRO) and metabolic interaction potential (MIP). For example, as one embodiment, the metabolite interaction simulation model 115 may include at least one of CarveMe, which automatically generates genome-based metabolic models, and SMETANA, which predicts metabolic interactions within microbial communities. The computer device may calculate MRO and MIP values by using the genomic sequence data of the target microorganism and the genomic sequence data of the strains included in the database in the metabolite interaction simulation model 115.
[0066] The computer device can input the genome analysis information and metabolic information into the first model 121, which is a regression model using the genome analysis information and metabolic information as features and using the growth index information as a label, so as to estimate the maximum optical density (maxOD) 123, which is an example of growth index information of the colony including the target microorganism.
[0067] The computer device may determine the first strain combination 133 by inputting the metabolic information and growth index information 105 of the composite strain and the metabolic information and growth index information of the colony including microorganisms stored in the database into the second model 131 .
[0068] In addition, the computer device may also re-input the first strain combination 133 into the ranking model 141 to re-determine the second strain combination 143 in a weighted manner.
[0069] Figure 3 FIG. 1 is a flowchart illustrating an operation of a computer device for estimating growth indicator information according to an embodiment of the present invention.
[0070] Reference Figure 3 , the computer device can obtain genome analysis information related to the target microorganism in step S210.
[0071] For example, Figure 4 As shown, a computer device can obtain genomic analysis information related to a target microorganism based on a shotgun sequencing pipeline 320. The computer device can extract DNA using the target microorganism's genomic sequence data 310 and break the DNA into small fragments. For example, the computer device can use various chemical and physical methods to disrupt the target microorganism's cell wall and isolate the DNA within the cell, or can use physical or enzymatic methods to break the DNA into small fragments. The computer device can determine the base sequence of the broken DNA fragments through DNA sequencing. For example, the computer device can use Illumina sequencing or Nanopore sequencing to identify the base sequence of the DNA fragments and then perform sequencing. The computer device can reassemble the short fragments of DNA sequence to reconstruct the original genomic sequence. The computer device can match the used and overlapping sequence fragments to generate contigs (contigs), which are the longest possible continuous DNA sequences. The computer device can predict the location and function of genes from the reconstructed genomic sequence. The predicted gene data 330 can include at least one of species composition data, predicted gene data, and metabolite data related to the target microorganism as genomic analysis information related to the target microorganism.
[0072] According to one embodiment, a computer device may obtain first metabolic information associated with each of a plurality of first microbial colonies including a target microorganism in step S220. Each of the plurality of first microbial colonies may be composed of any composite strain that combines sequence data of a single strain with genome data of the target microorganism. The first metabolic information may include at least one of MRO data and MIP data.
[0073] For example, Figure 5As shown, the computer device can generate a composite strain sample 420 by using the genome sequence data of the target microorganism, the number of constituent strains, and the single strain sequence data 410. The composite strain sample 420 is any composite strain that combines the sequence data of the single strain and the genome data of the target microorganism, which can indicate a plurality of first microbial colonies including the target microorganism. The composite strain sample 420 will refer to Figure 6 Detailed description.
[0074] The computer device can obtain first metabolic information 440 by using the composite strain sample 420 and metabolite interaction simulation 430. The computer device can use CarveMe to reconstruct a metabolic network from the genome sequence based on the sequence data of the composite strain sample 420. Alternatively, the computer device can use SMETANA to simulate metabolite exchange and competition from the genome sequence to analyze synergistic effects and competitive relationships between strains.
[0075] Figure 6 This is a conceptual diagram outlining the process of analyzing composite strain 403 into individual strains 405 and reconstructing them into composite strain 407, which is again reconstructed through simulations that include target microorganism 401, to obtain metabolic information. A computer device can use bioinformatics and systems biology methods to analyze microbial communities. Composite strain 403 is data included in a database, etc., and the computer device can individually analyze each microbial strain that constitutes composite strain 403. For example, the computer device can analyze the genomic information, metabolic pathways, and functional characteristics of each microorganism based on at least one of high-performance DNA sequencing, genome interpretation, and metagenomic analysis. The computer device can construct a new composite strain that includes target microorganism 401 based on the data obtained from individual strain 405. The computer device can then apply metabolite interaction simulations to the new composite strain that includes target microorganism 401 to obtain metabolic information.
[0076] The computer device according to one embodiment may estimate first growth indicator information related to each of the plurality of first microbial colonies by using the genome analysis information and the metabolic information as inputs of the first model in step S230 .
[0077] For example, Figure 7As shown, a computer device can train a first model 530 for estimating growth index information using composite strain sample data 520 stored in a database, and use the genome analysis information and metabolic information of each of the plurality of first microbial colonies 510, including target microorganisms reconstructed through simulation, as input, thereby estimating growth index information 540 for each of the plurality of first microbial colonies 510. In other words, the first model is a model for estimating growth index information, which can be a regression model trained using a dataset that uses genome analysis information related to the plurality of microorganisms, metabolic information related to each of the plurality of microbial colonies including each of the plurality of microorganisms, and growth index information as features, and uses the growth index information as a label. As one example, the first model can be a regression model trained using at least one of the second metabolic information and second growth index information related to each of the plurality of second microbial colonies, and genome analysis information related to the individual strains constituting each of the plurality of second microbial colonies, stored in the database, as a dataset.
[0078] Figure 8 FIG. 1 is a flow chart illustrating an operation of a computer device for determining a strain combination according to an embodiment of the present invention.
[0079] Reference Figure 8 The computer device may obtain second metabolic information and second growth indicator information related to each of the plurality of second microbial colonies stored in the database in step S610.
[0080] According to one embodiment, the computer device may determine, in step S620, a strain combination including the target microorganism by using a) first metabolic information related to each of a plurality of first microbial colonies including the target microorganism and first growth indicator information related to each of the plurality of first microbial colonies, and b) second metabolic information and second growth indicator information related to each of a plurality of second microbial colonies stored in a database.
[0081] For example, Figure 9 As shown, the computer device can determine the strain combination based on a collaborative filtering algorithm model. Figure 9 As shown, the user-item evaluation matrix of the latent factor model can be trained by utilizing the metabolic information and growth index information of the composite strain 720 and the metabolic information and growth index information of the strain sample 730 including the target microorganism stored in the database. The loss function of the latent factor model can be shown in Mathematical Formula 3.
[0082] [Mathematical formula 3]
[0083]
[0084] r (u,i) is the actual evaluation value between user u and item i, that is, the value of row u and column i of the R matrix, p u When decomposing the actual R matrix into the P matrix and Q matrix including the latent factors, the user u row vector of the P matrix, q i T When decomposing the actual R matrix into the P matrix and the Q matrix including the latent factors, the i-th row transposed vector of the Q matrix is represented, and λ may represent the regularization parameter multiplied by the regularization term.
[0085] The predicted R matrix value is calculated by taking the inner product of the P matrix and the Q matrix. The principle of the latent factor collaborative filtering model is to repeatedly optimize the cost function to minimize the error with the actual R matrix value. In addition, a regularization term can be added to prevent overfitting of the data.
[0086] Specifically, a latent factor model is learned by using the target microorganism and a single strain as user u, and using the third metabolic information and third growth index information, which are data concatenated with the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information, as item i. Separately, a latent factor model is learned by using the target microorganism and a single strain as item i, and using the third metabolic information and third growth index information, which are data concatenated with the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information, as user u.
[0087] The latent factor vector represents hidden characteristics of the user (target microorganism) or item (metabolite). These characteristics cannot be directly observed but can be inferred through the model. A computer device obtains first metabolic information and first growth index information for a plurality of first microbial colonies, including the target microorganism. Additionally, second metabolic information and second growth index information for another microbial colony stored in the database can be obtained, and the metabolic information and growth index information can be used as input to the model.
[0088] In addition, the computer device can search for the optimal model hyperparameters by using GridSearchCV in the Surprise library. In this case, the optimized hyperparameters can be n_factors (the number of latent factor dimensions), lr_all (the learning rate), and reg_all (the regularization parameter) of the latent factor model.
[0089] According to another embodiment, the computer device may determine the strain combination in step S620 based on a collaborative filtering algorithm model and adjust the weights of growth indicator information and metabolic information according to the intended purpose. For example, the computer device may determine the strain combination by using a ranking model that adjusts the weights of growth indicator information, MIP, and MRO values, thereby giving a higher weight to growth indicator information when rapid growth is desired, a higher weight to MIP values when cooperative survival is desired, and a higher weight to MRO values when inhibition of a specific target strain is desired.
[0090] According to another embodiment, the computer device may determine a strain combination in step S620 based on a second model. The second model learns the relationship between the target microorganism and a single strain based on the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information. For example, the computer device may determine the strain combination based on a transformer model. The computer device may use an attention mechanism to assign weights based on the importance of data points. Based on the metabolic information and growth index information, the computer device may learn whether two microorganisms are in a competitive relationship where growth is inhibited or in a symbiotic relationship where synergistic effects are exhibited through mutually beneficial interactions.
[0091] Figure 10 FIG. 1 is a diagram schematically illustrating a process of determining a strain combination based on a ranking model by a computer device according to an embodiment of the present invention.
[0092] Reference Figure 10 , the computer device can use the growth index information value as an item, the growth index information recommended strain list and the MIP value as an item, and the MIP recommended strain list and the MRO value as an item through a strain-based collaborative filtering algorithm model to output an MRO recommended strain list. The computer device can use these lists as input 810 of the ranking model 820 to determine the final recommended strain list 830 as a strain combination. The ranking model 820 can give weights to each input data set according to the user's purpose. After applying the weights, the computer device can calculate the comprehensive score of each strain according to the ranking model, thereby selecting the strain combination with the highest score.
[0093] Figure 11 is a block diagram illustrating the configuration of a computer device according to an embodiment.
[0094] The computer device is shown as being composed of a memory 910 and a processor 920, but is not necessarily limited thereto. The memory 910 and the processor 920 may each exist as a physically independent component.
[0095] The memory 910 may store various data used for the overall operation of the computer device, such as a program for processing or controlling the processor 920 in the computer device.
[0096] According to one embodiment, the memory 910 can store complex strain data, single strain data, large-scale datasets, algorithms, analytical models, and the like. Furthermore, the memory 910 can store multiple application programs, data, and instructions for operating the computer device. The memory 910 can be implemented as internal memory included in the processor 920, such as ROM, RAM, or a solid-state drive (SSD), or as a separate memory from the processor 920. According to one embodiment, the memory 910 can store models and learning data.
[0097] The processor may be a component for controlling the entire computer device. For example, the processor 920 may control the computer device to perform operations according to an embodiment of the present invention.
[0098] According to one embodiment, the processor 920 may obtain genome analysis information related to the target microorganism, obtain first metabolic information related to each of a plurality of first microbial colonies including the target microorganism, estimate first growth index information related to each of the plurality of first microbial colonies by using the genome analysis information and the metabolic information as inputs of a first model, and determine a strain combination based on the growth index information.
[0099] The processor 920 according to one embodiment may obtain second metabolic information and second growth index information related to each of the plurality of second microbial colonies stored in the database, and determine a strain combination by using the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information as inputs of the second model.
[0100] According to one embodiment, the processor 920 may receive genome sequence data of a target microorganism constituting a colony and the number of constituent strains to be included in the colony, use the target microorganism and the single strain as users, use third metabolic information and third growth index information, which are data connecting the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information, as items, and determine a microorganism strain combination for at least one of the third metabolic information and the third growth index information. In this case, each microorganism strain combination may be a combination of strains consisting of the number of combined strains input by the user.
[0101] The processor 920 according to an embodiment may re-determine the determined strain combination by determining respective weights of the estimated third metabolic information and the third growth indicator information.
[0102] According to one embodiment, the processor 920 may learn the relationship between the target microorganism and the single strain based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information, and may determine a strain combination based on the relationship between the target microorganism and the single strain.
[0103] Specifically, the processor 920 can control the operation of the computer device by utilizing various programs stored in the computer device's memory 910. The processor 920 may include a CPU, RAM, ROM, and a system bus. The processor 920 may be implemented as a single CPU or multiple CPUs (or DSPs, SoCs). In one embodiment, the processor 920 may be implemented as a digital signal processor (DSP), a microprocessor, or a time controller (TCON). However, the present invention is not limited to this and may also include one or more of a central processing unit (CPU), a microprocessor unit (MCU), a controller, an application processor (AP), a communication processor (CP), or an ARM processor, or may be defined by such terms. Furthermore, the processor 920 may be implemented as a system on chip (SoC) with built-in processing algorithms, a large-scale integration (LSI), or a field programmable gate array (FPGA). To efficiently execute deep learning and data processing algorithms, the processor 920 may also include components specialized for high-performance computing. These components may include Neural Processing Units (NPUs), Graphics Processing Units (GPUs), Tensor Processing Units (TPUs), etc.
[0104] Although described with reference to the limited embodiments and figures including the above embodiments, those skilled in the art will appreciate that various modifications and variations can be made from the above description. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a manner different from that described, or other components or equivalents are replaced or substituted.
[0105] Therefore, other implementations, other examples, and equivalents of the claims are also included within the scope of the claims to be described later.
[0106] Explanation of symbols
[0107] 101: User input 103: Database
[0108] 105: Metabolic and Growth Index Information 111: Shotgun Sequencing Process
[0109] 113: Genome Analysis Information 115: Metabolite Interaction Simulation Model
[0110] 117: Metabolic Information 121: First Model as a Regression Model
[0111] 123: Growth indicator information 131: Second model
[0112] 133: First strain combination 141: Sorting model
[0113] 143: Second strain combination 310: Genome sequence data of target microorganism
[0114] 320: Shotgun Sequencing Process 330: Predicting Gene Data
[0115] 401: Target microorganism 403: Composite strain
[0116] 405: Single strain 407: Composite strain reconstructed through simulation
[0117] 410: Single strain sequence data 420: Composite strain sample
[0118] 430: Metabolite Interaction Simulation 440: First Metabolism Information
[0119] 510: Multiple first microbial colonies 520: Composite strain sample data
[0120] 530: First Model 540: Growth Index Information
[0121] 720: Composite strain 730: Strain sample
[0122] 810: Input 820: Sorting Model
[0123] 830: Final Recommended Strain List 910: Memory
[0124] 920: Processor.
Claims
1. A method for determining a combination of microbial strains, comprising: a step of obtaining genome analysis information related to the target microorganism; a step of obtaining first metabolic information related to each of a plurality of first microbial colonies including the target microorganism; a step of estimating first growth indicator information associated with each of the plurality of first microbial colonies by using the genome analysis information and the metabolic information as inputs to a first model; as well as The step of determining a strain combination based on at least one of the metabolic information and the growth index information.
2. The method according to claim 1, wherein Each of the plurality of first microbial colonies is composed of any composite strain that combines sequence data of one or more single strains with genome data of the target microorganism.
3. The method according to claim 1, wherein The genome analysis information includes species composition data, predicted gene data, and metabolite data related to the target microorganism.
4. The method according to claim 1, wherein The first metabolic information includes at least one of metabolic resource overlap data and metabolic interaction potential data.
5. The method according to claim 1, wherein The first model is a regression model that is trained using second metabolic information related to each of the plurality of second microbial colonies, second growth indicator information, and genome analysis information related to a single strain constituting each of the plurality of second microbial colonies stored in a database as a data set.
6. The method according to claim 1, further comprising: a step of obtaining second metabolic information and second growth indicator information related to each of a plurality of second microbial colonies stored in a database; as well as The step of determining the strain combination by using the first metabolic information, the second metabolic information, the first growth indicator information and the second growth indicator information as inputs of a second model.
7. The method according to claim 6, wherein: The second model includes a latent factor collaborative filtering algorithm model, and The step of determining the strain combination comprises: The step of using the target microorganism and single strain as a user; a step of using third metabolism information and third growth index information, data connecting the first metabolism information, the second metabolism information, the first growth index information, and the second growth index information, as an item; and The step of determining a microbial strain combination corresponding to at least one of the third metabolic information and the third growth indicator information.
8. The method according to claim 7, wherein: The step of determining the strain combination further includes the step of re-determining the determined strain combination by determining respective weights of the estimated third metabolic information and the third growth indicator information.
9. The method according to claim 6, wherein: The second model includes a converter model, and The step of determining the strain combination comprises: The step of determining a strain combination based on a second model, wherein the second model learns the relationship between the target microorganism and one or more single strains based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information.
10. A method for generating a model for estimating growth indicator information, comprising: The step of obtaining genome analysis information associated with a plurality of microorganisms; a step of obtaining metabolic information and growth index information related to each of a plurality of microbial colonies including each of the plurality of microorganisms; generating a data set using the genome analysis information and the metabolic information as features and the growth indicator information as a label; as well as The step of generating a model based on the data set, the model estimating growth indicator information associated with each of the plurality of microbial colonies included.
11. The method according to claim 10, wherein: The metabolic information includes at least one of metabolic resource overlap data and metabolic interaction potential data.
12. A computer device comprising a processor, in, The processor obtains genome analysis information related to a target microorganism, obtains first metabolic information related to each of a plurality of first microbial colonies including the target microorganism, estimates first growth indicator information related to each of the plurality of first microbial colonies by using the genome analysis information and the metabolic information as inputs of a first model, and determines a strain combination based on at least one of the metabolic information and the growth indicator information.
13. The computer device according to claim 12, wherein: Each of the plurality of first microbial colonies is composed of any composite strain that combines sequence data of a single strain with genome data of the target microorganism.
14. The computer device according to claim 12, wherein: The genome analysis information includes species composition data, predicted gene data, and metabolite data related to the target microorganism.
15. The computer device according to claim 12, wherein: The first metabolic information includes at least one of metabolic resource overlap data and metabolic interaction potential data.
16. The computer device according to claim 12, wherein: The first model is trained using second metabolic information related to each of the plurality of second microbial colonies, second growth indicator information, and genome analysis information related to a single strain constituting each of the plurality of second microbial colonies stored in a database as a data set.
17. The computer device according to claim 12, wherein: The processor obtains second metabolic information and second growth index information related to each of a plurality of second microbial colonies stored in a database, and determines the strain combination by using the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information as inputs of a second model.
18. The computer device according to claim 17, wherein: The second model includes a latent factor collaborative filtering algorithm model, and The processor uses the target microorganism and the single strain as users, uses third metabolic information and third growth indicator information that connect the data of the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information as items, and determines a microorganism strain combination for at least one of the third metabolic information and the third growth indicator information.
19. The computer device according to claim 18, wherein: The processor re-determines by determining respective weights of the estimated third metabolic information and the third growth indicator information.
20. The computer device of claim 17, wherein: The second model includes a converter model, and The processor learns the relationship between the target microorganism and the single strain based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information, and determines a strain combination based on the relationship between the target microorganism and the single strain.