Method and apparatus for determining combinations of microbial strains

The method leverages genomic and metabolic data with reinforcement learning to optimize microbial strain combinations, addressing inefficiencies in existing experimental methods by enhancing productivity and metabolite production.

JP2026510363APending Publication Date: 2026-04-02BIOMATZ CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods for selecting and combining microbial strains for industrial applications are time-consuming, labor-intensive, and lack reproducibility, with experimental conditions significantly affecting results.

Method used

A method using genomic analysis, metabolic information, and growth index information to determine optimal microbial strain combinations through reinforcement learning, employing models like regression and transformer models to estimate and optimize strain interactions.

Benefits of technology

Facilitates the determination of optimal microbial strain combinations for improved productivity and metabolite production by effectively understanding and reflecting the relationships within complex microbial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510363000001_ABST
    Figure 2026510363000001_ABST
Patent Text Reader

Abstract

Embodiments of this specification relate to methods for optimizing microbial strain combinations and predicting the growth state of said strains. A method for determining microbial strain combinations according to one embodiment of this specification for achieving the aforementioned objectives may include: obtaining genomic analysis information about a target microorganism; obtaining first metabolic information about each of a plurality of first microbial colonies containing the target microorganism; estimating first growth indicator information about each of the plurality of first microbial colonies using the genomic analysis information and metabolic information as input to a first model; and determining a strain combination based on at least one of the metabolic information and the growth indicator information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for optimizing a combination of microbial strains and predicting the growth state of the strains, and in particular, to determining a specific combination of microbial strains using genomic analysis, metabolic information, and growth index information, and through this, it can be applied to optimize productivity, efficiency, or the production of specific metabolites.

Background Art

[0002] The selection and combination of microbial strains play a fundamental role in various fields such as biotechnology, pharmaceuticals, environmental engineering, and food engineering. Such strains are utilized for various purposes such as the production of specific compounds, the decomposition of harmful substances, and the expression of special biological effects. In particular, in industrial fermentation processes, an appropriate combination of microbial strains is essential for mass-producing specific substances, and in biological treatment processes, the selection of strains for effectively removing environmental pollutants is important.

[0003] Generally, the selection and combination of microbial strains have mainly relied on experimental methods. This includes the process of evaluating the growth rates, types and amounts of produced substances, interactions, etc. of various strains through large-scale culture experiments. However, such a method not only requires a long time and labor, but also has limitations in finding the optimal combination among numerous possible combinations. In addition, minute differences in experimental conditions can greatly affect the results, and there may be problems with reproducibility.

[0004] The above-described background art is technical information that the inventor had in order to derive the present invention or acquired during the derivation process of the present invention, and it cannot necessarily be said that it is a technology publicly disclosed to the general public before the filing of the present invention.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The problem that this invention aims to solve is to provide a reinforcement learning method and apparatus that takes into account relationships such as cooperation and competition between objects in a complex crowd environment, as proposed to solve the aforementioned problems. [Means for solving the problem]

[0006] A method for determining a combination of microbial strains according to the present invention to achieve the aforementioned objectives may include the steps of: obtaining genomic analysis information relating to a target microorganism; obtaining first metabolic information relating to each of a plurality of first microbial colonies containing the target microorganism; estimating first growth indicator information relating to each of the plurality of first microbial colonies using the genomic analysis information and the metabolic information as input to a first model; and determining a combination of strains based on at least one of the metabolic information and the growth indicator information.

[0007] Each of the aforementioned plurality of first microbial colonies may consist of any composite strain formed by combining the sequence data of one or more single strains with the target microbial genome data.

[0008] The genome analysis information may include species composition data, predicted gene data, and metabolite data relating to the target microorganism.

[0009] The species composition data includes phylogenetic information of the target microorganism, and specifically may include at least one of the following: species classification information, strain classification information, phylogenetic tree information, and phylogenetic distance information of the target microorganism.

[0010] The predicted gene data includes information relating to all genes derived from the whole genome sequence of the target microorganism, and may include a coded list of all gene types and / or functional annotation data for those genes. The predicted gene data also includes at least one of the following: information on metabolism-related genes, information on antibiotic-related genes, information on toxin genes, information on optimal culture medium composition, and information on biosynthetic gene groups of the target microorganism.

[0011] The aforementioned metabolite data also includes information relating to metabolites identified by estimating biochemical and metabolic pathways derived from information relating to all genes identified from the whole genome sequence of the target microorganism, and information relating to metabolites identified from biochemical and metabolic pathways derived from the protein (enzyme) information encoded by those genes.

[0012] The aforementioned metabolic information also includes the results of analyzing metabolic interactions between microorganisms in two or more microbial colonies. The results of the metabolic interaction analysis are analyzed considering the overlap of metabolic resources between microorganisms, the possibility of metabolic interaction, metabolic inconsistencies, and / or the prediction of minimum nutrient diversity for growth, and specifically include at least one of the following: metabolic resource overlap (MRO) and metabolic interaction potential (MIP).

[0013] The aforementioned Metabolic Resource Overlap (MRO) is a value calculated by determining the similarity of metabolites required when microorganisms within a microbial community exist independently. It is an indicator of the degree of competition among microorganisms for a given nutrient within the community and can be calculated using the following formula 1.

[0014]

number

[0015] The aforementioned Metabolic Interaction Potential (MIP) is a value that indicates the maximum number of metabolites that can be exchanged between microorganisms within a microbial community, and is an indicator of the metabolic dependence between microorganisms constituting the community. It can be calculated using the following formula 2.

[0016]

number

[0017] The aforementioned MRO and MIP were calculated using the SMETANA tool, and a detailed explanation of the algorithm and other aspects of the SMETANA analysis tool is described in the document "Metabolic dependencies drive species co-occurrence in diverse microbial communities (PNAS May 19, 2015 112(20) 6449-6454)".

[0018] The growth indicator information refers to information relating to indicators used to confirm or evaluate the growth, development, and / or proliferation level of a microbial colony (complex strain population) containing one or more microorganisms. In one example, the growth indicator information includes one or more selected from the group consisting of optical density (OD) information measured using spectrophotometry, time to maxOD, growth rate, and growth rate for each stage of microbial growth (lag phase, exponential growth phase, stationary phase, death phase), and is not limited to any information that allows confirmation of the growth and growth level of the microorganism. Furthermore, the growth indicator information may include a comparative value [△(single / community)] between the growth indicator information of any single microorganism and the growth indicator information of a complex strain population containing the single microorganism.

[0019] The first metabolic information may include at least one of the following: metabolic resource overlap (MRO) data and metabolic interaction potential (MIP) data.

[0020] The first model is also a regression model that features the genome analysis information and the first metabolic information.

[0021] The method may further include the steps of: obtaining second metabolic information and second growth indicator information for each of a plurality of second microbial colonies stored in a database; and determining the strain combination using the first metabolic information, the second metabolic information, the first growth indicator information and the second growth indicator information as input to a second model.

[0022] The second model includes a potential factor collaborative filtering algorithm model. The step of determining the strain combination may include: using the target microorganism and single strains as users; using the third metabolic information and the third growth index information of the data obtained by concatenating the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information as items; and determining a microorganism strain combination for at least one of the third metabolic information and the third growth index information.

[0023] The step of determining the strain combination may further include determining and re-determining the weighted values of the estimated third metabolic information and the third growth index information for the determined strain combination.

[0024] The second model includes a Transformer model. The step of determining the strain combination may include determining the strain combination based on a second model that has learned the relationship between the target microorganism and one or more single strains based on the first metabolic information, the second metabolic information, the first growth index information, and the second growth index information.

[0025] A method for generating a model for estimating growth index information according to an embodiment of the present specification to achieve the above-described problems may include: obtaining genomic analysis information regarding a plurality of microorganisms; obtaining metabolic information and growth index information regarding each of a plurality of microorganism colonies each including one of the plurality of microorganisms; generating a dataset using the genomic analysis information and the metabolic information as features and the growth index information as labels; and generating a model for estimating growth index information regarding each of the plurality of microorganism colonies included based on the dataset.

[0026] A computer device according to one embodiment of this specification for achieving the aforementioned problems includes a processor which can acquire genomic analysis information relating to a target microorganism, acquire first metabolic information relating to each of a plurality of first microbial colonies including the target microorganism, estimate first growth indicator information relating to each of the plurality of first microbial colonies using the genomic analysis information and the metabolic information as input to a first model, and determine a strain combination based on the growth indicator information.

[0027] Each of the aforementioned plurality of first microbial colonies may consist of any composite strain formed by combining the sequence data of one or more single strains with the target microbial genome data.

[0028] The genome analysis information may include species composition data, predicted gene data, and metabolite data relating to the target microorganism.

[0029] The first metabolic information may include at least one of the following: metabolic resource overlap (MRO) data and metabolic interaction potential (MIP) data.

[0030] The first model is also a regression model that features the genome analysis information and the first metabolic information.

[0031] The processor can acquire second metabolic information and second growth indicator information for each of the multiple second microbial colonies stored in the database, and use the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information as input to the second model to determine the strain combination.

[0032] The second model includes a latent factor collaborative filtering algorithm model, wherein the processor uses the target microorganism and single strain as users, and the third metabolic information and third growth indicator information of data concatenated from the first metabolic information, second metabolic information, first growth indicator information, and second growth indicator information as items, and can determine a combination of microbial strains for at least one of the third metabolic information and third growth indicator information.

[0033] The processor may determine and re-determine the weighted values ​​of the estimated third metabolic information and the third growth indicator information, respectively.

[0034] The second model includes a transformer model, and the processor can learn the relationship between the target microorganism and a single strain based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information, and determine a strain combination based on the relationship between the target microorganism and the single strain.

[0035] The microbial strain combinations determined or derived using the methods, computer devices, and / or processors of the present invention are also strain combinations for improving the functionality of target microorganisms. This improvement in functionality also includes improving the growth, activity (including physiological or pharmacological activity), stability, and / or colonization ability of the target microorganisms in the intestines. [Effects of the Invention]

[0036] According to one embodiment of the present invention, the diverse types of objects and their characteristics within a crowd can be effectively understood and reflected, allowing each object to determine the optimal action according to its own role and situation.

[0037] The effects of the present invention are not limited to those mentioned above. [Brief explanation of the drawing]

[0038] [Figure 1]This diagram schematically illustrates the process of determining bacterial strain combinations using a computer device according to one embodiment of the present invention. [Figure 2] This diagram schematically illustrates the process of determining bacterial strain combinations using a computer device according to one embodiment of the present invention. [Figure 3] This flowchart illustrates the operation of estimating growth indicator information for a computer device according to one embodiment of the present invention. [Figure 4] This diagram schematically illustrates the process of acquiring genome analysis information using a computer device according to one embodiment of the present invention. [Figure 5] This diagram schematically illustrates the process of acquiring metabolic information from a computer device according to one embodiment of the present invention. [Figure 6] This is a conceptual diagram illustrating the process of analyzing a complex bacterial strain into a single bacterial strain using a computer device according to one embodiment of the present invention, in order to obtain metabolic information. [Figure 7] This diagram schematically illustrates the process of estimating growth indicator information for a computer device according to one embodiment of the present invention. [Figure 8] This is a flowchart illustrating the operation of a computer device that determines a combination of bacterial strains according to one embodiment of the present invention. [Figure 9] This diagram schematically illustrates the process of determining a combination of bacterial strains based on collaborative filtering by a computer device according to one embodiment of the present invention. [Figure 10] This diagram schematically illustrates the process of determining a combination of bacterial strains based on a ranking model of a computer device according to one embodiment of the present invention. [Figure 11] This is a block diagram illustrating the configuration of a computer device according to one embodiment. [Modes for carrying out the invention]

[0039] The terms used in this invention are used solely to describe specific embodiments and are not intended to limit the scope of other embodiments. Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have meanings that are generally understood by those ordinaryly skilled in the art described herein. Terms used herein that are defined in general dictionaries shall be interpreted as having the same or similar meaning as in the context of the relevant art, and not as ideal or overly formal unless explicitly defined herein. Where applicable, terms defined herein should not be construed as excluding embodiments of the invention.

[0040] Hereinafter, various embodiments will be described in detail with reference to the accompanying drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present invention pertains. However, the technical idea of ​​the present invention can be transformed and embodied in various forms, and is not limited to the embodiments described herein. In describing the embodiments disclosed herein, if it is determined that specifically describing related prior art would obscure the gist of the technical idea of ​​the present invention, then specific descriptions relating to such prior art will be omitted. Identical or similar components will be given the same reference numeral, and redundant descriptions thereof will be omitted.

[0041] In this embodiment, the term "~part" refers to a component that performs a specific function, whether performed by software or hardware such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). However, "~part" is not limited to being performed by software or hardware. "~part" can also exist in data form stored on an addressable recording medium and can be embodied by instructions, configured to allow one or more processors to perform a specific function.

[0042] Software comprises computer programs, code, instructions, or one or more combinations thereof, which can configure a processing unit to operate as desired, or independently or collectively, instruct the processing unit. Software and / or data can be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer recording medium or device, or transmitted signal wave, in order to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software can be distributed, stored in a distributed manner, or executed on a networked computer system. Software and data can be stored in one or more computer-readable media. Software can be read into main memory from other computer-readable media, such as data storage devices, or from other devices via communication interfaces. Software instructions stored in main memory can cause the processor to perform processes or steps described in detail later. On the other hand, processes consistent with the principles of the present invention can be performed using fixed wiring circuits instead of, or in combination with, software instructions. Therefore, embodiments consistent with the principles of the present invention are not limited to any particular combination of hardware circuitry and software.

[0043] The terms used in this application are used solely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. Hereinafter, terms such as “includes” or “having” are intended to indicate the existence of features, numbers, stages, operations, components, parts, ingredients, materials, or combinations thereof described in the specification, and should not be understood to presuppose the existence or possibility of adding one or more other features, numbers, stages, operations, components, parts, ingredients, materials, or combinations thereof. Terms such as “first,” “second,” etc., may be used to describe a variety of components, but the components should not be limited by such terms. Such terms are used solely for the purpose of distinguishing one component from another.

[0044] The term "model" as used in this invention may include any form of algorithm or methodology used to learn or understand specific patterns or structures from data. Models may include not only machine learning models such as regression models, decision trees, random forests, support vector machines, K-nearest neighbors, Naive Bayes, and clustering algorithms, but also deep learning models such as neural networks, convolutional neural networks, circulatory neural networks, Transformer-based neural networks, GANs (Generative Adversarial Networks), and autoencoders. A "model" specifies a set of learned parameters or weights used to predict or classify outputs related to a given input, and this model may be learned through methods such as supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning. Furthermore, it may include not only single models but also diverse learning methods and structures such as ensemble models, multimodal models, and models learned through transfer learning. Such models may be pre-trained on one computer device and another computer device to predict outputs for a given input, and then used on the other computer device.

[0045] Figures 1 and 2 are schematic diagrams illustrating the strain combination determination process of a computer device according to one embodiment of the present invention.

[0046] Referring to Figures 1 and 2, the computer device may obtain user input 101 from the user. Such user input 101 may include the genetic sequence data of a target microorganism and the number of constituent strains. Such target genetic sequence data includes the genetic information of the microorganism, and the number of strains may indicate the number of constituent strains for a strain combination. In other words, the user may input the genetic sequence data of a target microorganism to constitute a colony and the number of constituent strains to be included in the colony.

[0047] The computer device can acquire single-strain sequence data and complex-strain sequence data from database 103. Such single-strain sequence data can indicate the genetic sequence data of each individual single strain. Complex-strain sequence data can indicate one or more coexisting groups (morphologies) consisting of two or more microbial species, and the genetic sequence data of the microorganisms belonging to each of the coexisting groups. Specifically, the coexisting groups may include collections of microorganisms that are isolated or identified, actually coexist, or form colonies in diverse environments, including the intestines and / or feces. Based on user input 101, the computer device can acquire genomic analysis information 113 about the target microorganism, which in one embodiment may be acquired using a shotgun analysis pipeline 111. Specifically, the computer device can extract the DNA of the target microorganism and, if necessary, amplify the DNA. The computer device can randomly divide the DNA into thousands to millions of small fragments and obtain sequences. The computer device can use the acquired sequence data to reconstruct the genome sequence and predict the location of genes, a list of coded whole gene types, and / or functional annotations of such genes in the assembled genome. The genome analysis information 113 may include species composition data, predicted gene data, and metabolite data for the target microorganism.

[0048] Furthermore, the computer device can acquire metabolic information 117 based on the database 103, user input 101, and the metabolite interaction simulation model 115. Such metabolic information may include at least one of the following: metabolic resource overlap (MRO) data and metabolic interaction potential (MIP) data. For example, in one embodiment, the metabolite interaction simulation model 115 may include at least one of CarveMe, which automatically generates genome-based metabolic models, and SMETANA, which predicts metabolic interactions within a microbial community. The computer device can calculate MRO and MIP values ​​using the target microorganism gene sequence data and the gene sequence data of strains included in the database in the metabolite interaction simulation model 115.

[0049] The computer device can estimate the maximum optical density (maxOD) 123 as an example of growth indicator information for a colony containing a target microorganism by inputting the genome analysis information and metabolic information into a first model 121, which is a regression model that features genome analysis information and metabolic information and labels growth indicator information.

[0050] The computer device can input metabolic information and growth indicator information 105 of the composite strains stored in the database, as well as metabolic information and growth indicator information of the colonies containing microorganisms, into the second model 131 to determine the first strain combination 133.

[0051] Furthermore, the computer device can re-determine the second strain combination 143 by inputting the first strain combination 133 back into the ranking model 141 to determine the weighted values.

[0052] Figure 3 is a flowchart illustrating the operation of estimating growth indicator information for a computer device according to one embodiment of the present invention.

[0053] Referring to Figure 3, the computer system can acquire genomic analysis information about the target microorganism at the S210 stage.

[0054] For example, a computer system can obtain genomic analysis information about a target microorganism based on a Shotgun Sequencing Pipeline 320, as illustrated in Figure 4. The computer system can extract DNA using the target microorganism's genetic sequence data 310 and divide it into smaller sections. For example, the computer system can use various chemical and physical methods to disrupt the target microorganism's cell wall and separate the DNA inside the cell, and can degrade the DNA into smaller sections using physical or enzymatic methods. The computer system can determine the base sequence of the degraded DNA sections through DNA sequencing. For example, the computer system can use Illumina sequencing or Nanopore sequencing to identify and sequence the base sequence of the DNA sections. The computer system can reconstruct the original genetic sequence by reconstructing short DNA sequence sections. The computer system can use overlapping sequence sections to match and thereby generate contigs, which are the longest continuous DNA sequences. The computer device can predict the location and function of genes in the reconstructed gene sequence. Such predicted gene data 330 may include at least one of the following as genomic analysis information for the target microorganism: species composition data, predicted gene data, and metabolite data.

[0055] A computer device according to one embodiment may acquire first metabolic information for each of a plurality of first microbial colonies containing a target microorganism in step S220. Each of these plurality of first microbial colonies may consist of any composite strain formed by combining sequence data of a single strain and genome data of the target microorganism. The first metabolic information may include at least one of metabolic resource overlap (MRO) data and metabolic interaction potential (MIP) data.

[0056] For example, as shown in Figure 5, a computer device can generate a composite strain sample 420 using the genetic sequence data, number of constituent strains, and single strain sequence data 410 of the target microorganism. The composite strain sample 420 is an arbitrary composite strain formed by combining the sequence data of a single strain and the target microorganism genome data, and can indicate multiple first microbial colonies containing the target microorganism. Such a composite strain sample 420 is described in detail in Figure 6.

[0057] The computer system can acquire first metabolic information 440 using a composite strain sample 420 and a metabolite interaction simulation 430. Based on the sequence data of the composite strain sample 420, the computer system can reconstruct the metabolic network from the gene sequence using CarveMe. Alternatively, the computer system can analyze synergistic effects and competitive relationships between strains by simulating metabolite exchange and competition from the gene sequence using SMETANA.

[0058] Figure 6 is a conceptual diagram illustrating the process of obtaining metabolic information by analyzing the complex strain 403 into a single strain 405, and then reconstructing it into a complex strain 407 that includes the target microorganism 401 in a simulation. The computer system can analyze the microbial community using approaches from bioinformatics and systems biology. The complex strain 403 is data contained in a database, and the computer system can individually analyze each microbial strain that makes up the complex strain 403. For example, the computer system can analyze the genomic information, metabolic pathways, and functional characteristics of each microorganism based on at least one of the following methods: high-performance DNA sequencing, genome interpretation, and metagenomic analysis. The computer system can construct a new complex strain including the target microorganism 401 based on the data obtained through the single strain 405. Subsequently, the computer system can obtain metabolic information by applying metabolite interaction simulations to the new complex strain including the target microorganism 401.

[0059] In one embodiment, the computer device can estimate first growth indicator information for each of several first microbial colonies by taking genome analysis information and metabolic information as input to the first model in step S230.

[0060] For example, a computer device can train a first model 530 that estimates growth indicator information using composite strain sample data 520 stored in a database, as shown in Figure 7, and estimate growth indicator information 540 for each of the first multiple microbial colonies 510, each containing a target microorganism reconstructed in a simulation, using genomic analysis information and metabolic information for each of the first multiple microbial colonies 510 as input. That is, the first model is a model for estimating growth indicator information and is also a regression model trained as a dataset with genomic analysis information for multiple microorganisms, metabolic information and growth indicator information for each of the multiple microbial colonies containing each of the multiple microorganisms as features, and growth indicator information as labels. In one embodiment, the first model is also a regression model trained as a dataset with at least one of the following as a dataset: second metabolic information for each of the multiple second microbial colonies stored in a database, second growth indicator information, and genomic analysis information for each of the single strains constituting each of the multiple second microbial colonies.

[0061] Figure 8 is a flowchart showing the operation of a computer device that determines a combination of bacterial strains according to one embodiment of the present invention.

[0062] Referring to Figure 8, the computer device can acquire second metabolic information and second growth indicator information for each of the multiple second microbial colonies stored in the database at step S610.

[0063] In one embodiment, a computer device can determine a combination of strains containing a target microorganism in step S620 using a) first metabolic information and first growth indicator information for each of a plurality of first microbial colonies containing the target microorganism, and b) second metabolic information and second growth indicator information for each of a plurality of second microbial colonies stored in a database.

[0064] For example, a computer device can determine strain combinations based on a collaborative filtering algorithm model, as illustrated in Figure 9. As illustrated in Figure 9, a user-item evaluation matrix of a latent factor model can be trained using metabolic and growth indicator information of 720 complex strains stored in a database and metabolic and growth indicator information of a strain sample 730 containing the target microorganism. The loss function of such a latent factor model is expressed as shown in Equation 3.

[0065]

number

[0066] The predicted R-matrix value is calculated by the inner product of the P-matrix and the Q-matrix. The principle of the latent factor collaborative filtering model lies in iterating through this process to optimize the cost function so that it has the minimum error between the predicted R-matrix value and the actual R-matrix value. Furthermore, a normalization term may be added to prevent data overfitting.

[0067] Specifically, a latent factor model can be trained using the target microorganism and a single strain as user u, and the third metabolic information and third growth indicator information of data concatenated from the first metabolic information, second metabolic information, first growth indicator information, and second growth indicator information as item i. Alternatively, a latent factor model can be trained using the target microorganism and a single strain as item i, and the third metabolic information and third growth indicator information of data concatenated from the first metabolic information, second metabolic information, first growth indicator information, and second growth indicator information as user u.

[0068] The latent factor vector represents the hidden characteristics of the user (target microorganism) or item (metabolites), which are not directly observed but can be inferred through a model. The computer system acquires primary metabolic information and primary growth indicator information for multiple primary microbial colonies, including the target microorganism. It also acquires secondary metabolic information and secondary growth indicator information for other microbial colonies stored in a database, and the metabolic and growth indicator information can be used as input to the model.

[0069] Furthermore, the computer can use the GridSearchCV library from Surprise to search for the optimal model hyperparameters. In this process, the hyperparameters to be optimized are also the n_factors (number of latent factor dimensions), lr_all (learning rate), and reg_all (normalization parameters) of the latent factor model.

[0070] In other embodiments, the computer device may determine strain combinations based on a collaborative filtering algorithm model in step S620, and adjust the weighting values ​​of growth indicator information and metabolic information according to the purpose to determine the strain combinations. For example, the computer device may determine strain combinations using a ranking model that adjusts the weighting values ​​of growth indicator information, MIP, and MRO values, assigning a high weighting to the growth indicator information value when rapid growth is desired, assigning an even higher weighting to the MIP value if the purpose is cooperative engraftment, and assigning a high weighting to the MRO value if the purpose is to suppress a specific target fungus.

[0071] Furthermore, in other embodiments, the computer device may, in step S620, determine strain combinations based on a second model that has learned the relationship between the target microorganism and a single strain, based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information. For example, the computer device may determine strain combinations based on a transformer model. The computer device may use an attention mechanism to assign weighted values ​​to the relationship with data points according to their importance, and learn, based on the metabolic information and growth indicator information, whether the two microbial species are in a competitive relationship that inhibits each other's growth, or in a symbiotic relationship that shows synergistic effects through mutually beneficial interactions.

[0072] Figure 10 is a schematic diagram illustrating the process of determining a combination of bacterial strains based on a ranking model of a computer device according to one embodiment of the present invention.

[0073] Referring to Figure 10, the computer device may output a list of recommended strains based on growth indicator information values ​​as items, a list of recommended strains based on metabolic interaction potential (MIP) values ​​as items, and a list of recommended strains based on metabolic resource superposition (MRO) values ​​as items, based on a collaborative filtering algorithm model for strains. The computer device may use such lists as inputs 810 to a ranking model 820 to determine a final recommended strain list 830 as a strain combination. The ranking model 820 may assign weights to each input dataset according to the user's purpose. After the weights are applied, the computer device can calculate an overall score for each strain using the ranking model and select strain combinations that received high scores.

[0074] Figure 11 is a block diagram illustrating the configuration of a computer device according to one embodiment.

[0075] Although the computer device is illustrated as consisting of memory 910 and processor 920, it is not necessarily limited to this configuration. Memory 910 and processor 920 can each exist as a single, physically independent component.

[0076] The memory 910 can store various data for the operation of the computer device as a whole, such as programs for processing or controlling the processor 920.

[0077] In one embodiment, memory 910 can store complex strain data, single strain data, large datasets, algorithms, analytical models, etc., and memory 910 can also store numerous application programs that are driven, data for the operation of computer devices, and instruction words. Memory 910 can be implemented as internal memory such as ROM, RAM, or SSD included in processor 920, or as memory separate from processor 920. In one embodiment, memory 910 can store models and training data.

[0078] The processor 920 is also a component for controlling the computer device overall. For example, the processor 920 can control the computer device to perform operations according to one embodiment of the present invention.

[0079] A processor 920 according to one embodiment can acquire genomic analysis information about a target microorganism, acquire first metabolic information for each of a plurality of first microbial colonies containing the target microorganism, estimate first growth indicator information for each of the plurality of first microbial colonies using the genomic analysis information and metabolic information as input to a first model, and determine a strain combination based on the growth indicator information.

[0080] In one embodiment, the processor 920 can acquire second metabolic information and second growth indicator information for each of a plurality of second microbial colonies stored in a database, and determine a strain combination using the first metabolic information, second metabolic information, first growth indicator information, and second growth indicator information as input to a second model.

[0081] In one embodiment, the processor 920 receives the genetic sequence data of a target microorganism to constitute a colony and the number of constituent strains included in the colony as input. The target microorganism and a single strain are designated as the user, and the third metabolic information and third growth indicator information, which are concatenated from the first metabolic information, second metabolic information, first growth indicator information, and second growth indicator information, are designated as the item. The processor can then determine a combination of microbial strains for at least one of the third metabolic information and third growth indicator information. In this case, each combination of microbial strains is also composed of as many strains as the number of constituent strains input by the user.

[0082] In one embodiment, the processor 920 can determine the determined strain combination by determining the weighted values ​​of the estimated third metabolic information and third growth indicator information, respectively.

[0083] A processor 920 according to one embodiment can learn the relationship between the target microorganism and a single strain based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information, and can determine a strain combination based on the relationship between the target microorganism and the single strain.

[0084] Specifically, the processor 920 can control the operation of the computer device using various programs stored in the computer device's memory 910. The processor 920 may include a CPU, RAM, ROM, system bus, etc. The processor 920 may be embodied as a single CPU or multiple CPUs (or DSP, SoC). In one embodiment, the processor 920 may be embodied as a digital signal processor (DSP), microprocessor, or TCON (Time controller) that processes digital signals. However, it is not limited to these, and may include or be defined as one or more of the following: central processing unit (CPU), MCU (Micro Controller Unit), MPU (micro processing unit), controller, application processor (AP), or communication processor (CP), or ARM processor. Furthermore, the processor 920 can be implemented as a System on Chip (SoC) or Large Scale Integration (LSI) with integrated processing algorithms, or as a Field Programmable Gate Array (FPGA). To efficiently perform deep learning and data processing algorithms, the processor 920 may further include components specialized for high-performance computing. Such components may include a Neural Processing Unit (NPU), a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), and so on.

[0085] As stated above, even if embodiments are described by limited embodiments and drawings, a person with ordinary skill in the art can make various modifications and variations from the above description. For example, the described technique may be performed in a different order than described, and / or the components of the described system, structure, apparatus, circuit, etc. may be combined or combined in a different manner than described, or substituted by other components or equivalents, and the appropriate results may be achieved.

[0086] Therefore, other embodiments, other embodiments, and those equivalent to the claims described below also fall within the scope of the claims. [Explanation of Symbols]

[0087] 101 User Input 103 Databases 105 Metabolic Information and Growth Indicator Information 111 Shotgun Analysis Pipeline 113 Genome Analysis Information 115 Metabolite Interaction Simulation Models 117 Metabolic information 121 The first regression model 123 Growth index information 131 Second Model 133 First strain combination 141 Ranking Models 143 Second strain combination 310 Genetic sequence data of target microorganisms 320 Shotgun Sequence Analysis Pipeline 330 Predicted Gene Data 401 Target microorganism 403 Compound strain 405 Single strains 407 Composite strains reconstructed in simulations 410 Single-Strain Sequence Data 420 combined bacterial strain samples 430 Metabolite Interaction Simulation 440 1st metabolic information 510 First Multiple Microbial Colonies 520 Composite Strain Sample Data 530 First Model 540 Growth index information 720 complex strains 730 bacterial strain samples 810 Input 820 Ranking Models 830 Final Recommended Strain List 910 memory 920 Processor

Claims

1. In the method for determining combinations of microbial strains, The stage of obtaining genomic analysis information about the target microorganism, The steps include obtaining first metabolic information for each of a plurality of first microbial colonies, including the aforementioned target microorganism, The steps include: estimating first growth indicator information for each of the multiple first microbial colonies using the genome analysis information and the metabolic information as input to the first model; A method comprising the step of determining a strain combination based on at least one of the metabolic information and growth indicator information.

2. Each of the aforementioned plurality of first microbial colonies is The method according to claim 1, comprising any composite strain formed by combining sequence data of one or more single bacterial strains and genome data of the target microorganism.

3. The method according to claim 1, wherein the genome analysis information includes species composition data, predicted gene data, and metabolite data relating to the target microorganism.

4. The first metabolic information mentioned above is: The method according to claim 1, comprising at least one of the following: metabolic resource overlap (MRO) data and metabolic interaction potential (MIP) data.

5. The method according to claim 1, wherein the first model is a regression model trained using a dataset of second metabolic information, second growth indicator information, and genomic analysis information for each of the multiple second microbial colonies stored in a database.

6. The steps include obtaining secondary metabolic information and secondary growth indicator information for each of the multiple secondary microbial colonies stored in the database, The method according to claim 1, further comprising the step of determining the strain combination using the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information as input to a second model.

7. The second model includes a latent factor collaborative filtering algorithm model. The step of determining the aforementioned strain combination is, The steps include: using the aforementioned target microorganism and single strain as users; The steps include: setting the third metabolic information and the third growth indicator information of the data obtained by concatenating the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information as items; The method according to claim 6, comprising the step of determining a combination of microbial strains for at least one of the third metabolic information and the third growth indicator information.

8. The step of determining the aforementioned strain combination is, The method according to claim 7, further comprising the step of determining and re-determining the weighted values ​​of the estimated third metabolic information and the third growth indicator information, respectively, for the determined strain combination.

9. The aforementioned second model includes a Transformer model, The step of determining the aforementioned strain combination is, The method according to claim 6, further comprising the step of determining a strain combination based on a second model that has learned the relationship between the target microorganism and one or more single strains, based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information.

10. In a method for generating a model to estimate growth indicator information, The stage of obtaining genome analysis information for multiple microorganisms, The steps include obtaining metabolic information and growth indicator information for each of the multiple microbial colonies, each of which contains the aforementioned multiple microorganisms, A step of generating a dataset in which the genome analysis information and the metabolic information are features and the growth indicator information is used as labels, A method for generating a model that estimates growth indicator information for each of several microbial colonies included in the aforementioned dataset.

11. The aforementioned metabolic information is, The method according to claim 10, comprising at least one of the following: metabolic resource overlap (MRO) data and metabolic interaction potential (MIP) data.

12. In a computer device, Including the processor, The aforementioned processor, A computer device that obtains genome analysis information about a target microorganism, obtains first metabolic information about each of a plurality of first microbial colonies containing the target microorganism, estimates first growth indicator information about each of the plurality of first microbial colonies using the genome analysis information and the metabolic information as input to a first model, and determines a strain combination based on at least one of the metabolic information and the growth indicator information.

13. Each of the aforementioned plurality of first microbial colonies is The computer device according to claim 12, comprising any composite strain formed by combining sequence data of a single bacterial strain and the genome data of the target microorganism.

14. The computer apparatus according to claim 12, wherein the genome analysis information includes species composition data, predicted gene data, and metabolite data relating to the target microorganism.

15. The first metabolic information mentioned above is: The computer device according to claim 12, comprising at least one of metabolic resource overlap (MRO) data and metabolic interaction potential (MIP) data.

16. The computer device according to claim 12, wherein the first model is trained using a dataset of second metabolic information, second growth indicator information, and genome analysis information for each of the multiple second microbial colonies stored in a database.

17. The aforementioned processor, The computer device according to claim 12, which obtains second metabolic information and second growth indicator information for each of a plurality of second microbial colonies stored in a database, and determines the strain combination using the first metabolic information, the second metabolic information, the first growth indicator information and the second growth indicator information as input to a second model.

18. The second model includes a latent factor collaborative filtering algorithm model. The aforementioned processor, The computer device according to claim 17, wherein the target microorganism and single strain are used as users, the third metabolic information and third growth indicator information of data obtained by concatenating the first metabolic information, the second metabolic information, the first growth indicator information and the second growth indicator information are used as items, and a combination of microbial strains is determined for at least one of the third metabolic information and the third growth indicator information.

19. The aforementioned processor, The computer device according to claim 18, which determines and re-determines the weighted values ​​of the estimated third metabolic information and the third growth indicator information, respectively.

20. The aforementioned second model includes a Transformer model, The aforementioned processor, The computer device according to claim 17, which learns the relationship between the target microorganism and a single strain based on the first metabolic information, the second metabolic information, the first growth indicator information, and the second growth indicator information, and determines a strain combination based on the relationship between the target microorganism and the single strain.