Determining Microbiota Samples and Predicting Mixes to Produce Target Mix Products

The method addresses inefficiencies in predicting and producing complex microbial community mixes by using a mix interaction model and evolutionary algorithms to simulate and determine initial samples, achieving accurate and cost-effective production of microbiome products.

JP2025536596APending Publication Date: 2025-11-07マーファルマ
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525622
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-08
Filing Date
2023-11-06
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Current methods for predicting and producing complex microbial community mixes are inefficient, consuming scarce materials and time due to empirical approaches and unsatisfactory linear prediction models, which do not accurately account for the nonlinear behavior of microbiota mixing.

Method used

A computer-aided method using a mix interaction model and evolutionary algorithms to predict and determine initial samples for producing targeted microbiota mixes, accounting for nonlinear behavior through matrix calculations and iterative co-culture strategies.

Benefits of technology

Enables rapid and accurate simulation of microbiota mixes at low cost, allowing for efficient production of microbiome products that meet therapeutic or environmental needs without material waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536596000001_ABST
    Figure 2025536596000001_ABST
Patent Text Reader

Abstract

An evolutionary algorithm is used to determine parameters for a complex microbial community (CMC) product production process based on a target profile of the CMC product. The CMC mixing and co-cultivation operations for the CMC product are modeled using a trained matrix-based model. The evolutionary algorithm iteratively refines candidate parameters, including the set and mixing ratio of complex microbial community samples in an initial sample collection, for one or more mixing operations in the production process. The determined set of samples and associated mixing ratios are then used to control the actual extraction and processing of the complex microbial community samples according to the mix production process, resulting in a CMC product that is as close as possible to the target profile in terms of profiling characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to modeling processes or pipelines that produce mixes of microbiomes, and more particularly to methods for predicting mix compositions and determining which complex microbial communities (microbiomes) should be mixed to produce a target mix composition. [Background technology]

[0002] Complex microbial communities (also known as microbiota) play an important role in health and disease. In particular, it has been discovered that administering or transplanting complex microbial communities, for example through fecal microbial transplantation (FMT), may have the potential to treat infections and diseases.

[0003] When administering or transplanting complex microbial communities, it is important that the administered or transplanted sample has an appropriate profile in terms of viability, functionality, and diversity of microorganisms and components, such as bacteria, archaea, viruses, phages, protozoa, metabolites, yeast, RNA, and / or fungi.

[0004] Some administration and transplantation methods, such as FMT, are often empirical and do not take special precautions to ensure the diversity of microorganisms present in the samples used or to optimize microbial viability.

[0005] Furthermore, not all samples collected from donors provide a satisfactory profile of complex microbial communities and their derived products (metabolites, RNA, etc.) for efficient treatment.

[0006] Therefore, to increase the diversity of samples that can be used as inoculum for administration or transplantation, a mix of complex microbial community samples collected from multiple donors has been considered, such as the Native Microbiome Ecosystem Therapies product (MET-N)® manufactured by MaaT Pharma.

[0007] As disclosed in WO 2022 / 136694, expanding a complex microbial community by microbial co-cultivation in separate bioreactors helps expand the microbial community while maintaining high diversity. An exemplary expansion method includes: (a) culturing the complex microbial community in at least two bioreactors, preferably at least three bioreactors, to obtain at least two co-culture samples, each bioreactor having at least one different parameter selected from pH, temperature, pressure, incubation time, retention time, gas supply conditions, redox potential, culture medium, light source, and combinations thereof; and (b) mixing the at least two co-culture samples obtained in step (a) to obtain an expanded complex microbial community. When formulated for administration, this type of product corresponds to a Co-cultivated Microbiome Ecosystem Therapies product (MET-C), such as that developed by MaaT Pharma.

[0008] To test various mixes or expanded complex communities, samples are currently mixed or expanded randomly, and the resulting product is then sequenced to obtain a final mix or expanded profile, from which preventative and therapeutic properties are inferred. This test-based approach has several drawbacks. In particular, it consumes scarce materials, given the significant difficulty in obtaining samples from donors, and requires weeks to complete due to the lengthy analytical process. Furthermore, significant co-cultivation time (e.g., hours or days) is required in addition to expansion, particularly when expanding complex communities, lengthening the testing period.

[0009] Therefore, the prediction of the mix composition, i.e., the prediction of the mix product profile, has been attempted through modeling of the process or pipeline of the production of the microbiota mix, often as a linear prediction model. However, known models are often unsatisfactory, and there is a need to propose an improved modeling of the mix production process.

[0010] Invertible prediction models can be advantageously used to determine the initial sample to be processed using a mix production process with the goal of obtaining a desired final mix composition (e.g., one with a desired therapeutic effect). However, since many of the prediction models are irreversible, a new mechanism for determining the initial sample must be proposed. Summary of the Invention

[0011] The present invention seeks to overcome some of the problems mentioned above.

[0012] One aspect of the present invention focuses on improving the prediction (and therefore modeling) of mixing co-culture samples obtained after growth in a bioreactor. The improved prediction is based on computer-aided design to account for the deviation between traditional linearly predicted profiles of such mixtures and the true profiles (obtained by profiling the resulting expanded complex microbial communities) when predicting mix composition.

[0013] In this regard, this aspect of the invention provides a computer-assisted method for predicting the mix composition resulting from the mixing of Co-cultured Complex Microbial Community (CCMC) products obtained by co-cultivation of one or more starting Complex Microbial Community (CMC) products, which may be different, preferably by different co-cultivation in each bioreactor of the same starting CMC product, comprising: (a) predicting intermediate mix profiles for CCMC product mixes using a linear approach; and (b) correcting the intermediate mix profile to a predicted mix profile using a mix interaction model trained from a reference linear predictive mix profile and a corresponding reference true mix profile; This paper proposes a method including:

[0014] Thus, this aspect of the invention models the nonlinear behavior of microbiota mixing, which advantageously allows for rapid and accurate simulation of various mix compositions of CCMC products at low cost, especially without consuming real materials.

[0015] Therefore, prior to routine production, co-culture strategies can be defined according to the needs of the intended application (therapeutic, preventive, environmental, etc.).

[0016] The modeling can be performed based on matrices. For example, predicting the intermediate mix profile may include calculating a matrix product between a first matrix defining the mixture ratio of CCMC products and a second matrix defining the individual profiles of CCMC products. Then, correcting the intermediate mix profile may include calculating a matrix product between a matrix representing the intermediate mix profile and a square mix interaction matrix of a trained mix interaction model. Here, the mix interaction model may be a square mix interaction matrix trained from a reference linear predicted mix profile and a corresponding reference true mix profile.

[0017] The use of matrices to perform microbiota mix predictions advantageously takes into account a large number of profiling features, allowing for rapid calculation to obtain one or more predicted mixture profiles for the resulting mix product(s).

[0018] In some embodiments, the mix composition is obtained from repeated mixing of CCMC products obtained by separate co-cultures in each bioreactor of the same starting CMC product, and the method includes multiple repetitions of steps (a) and (b), where the predicted mix profile obtained in step (b) of the previous repetition is used to obtain the profile of each CCMC product for step (a) of the next repetition. This approach models multiple loops of co-culture and mixing, i.e., the starting CMC product for the next repetition is the mixing result of the previous repetition. The same CMC product produces the CCMC product mixed in the next repetition. By chaining co-cultures from the same starting CMC product, a significant amount of material (final mix composition) can be obtained. Therefore, this configuration allows for better prediction of chained co-cultures.

[0019] In some embodiments, the same mix interaction model is used across iterations, meaning that the same mix interaction model (e.g., the same matrix) is used in each step (b) across iterations, thereby reducing processing and memory costs since only one model training is required to run multiple iterations and only one model needs to be stored.

[0020] In some embodiments, the starting CMC product is obtained from a mix of complex microbial community samples selected from an initial sample collection, and the method comprises: Predicting the CMC profile of the starting CMC product based on the sample profile of the selected complex microbial community sample; and Predicting the CCMC profile of the CCMC product based on the predicted CMC profile (used in the first iteration of steps (a) to (b) if multiple iterations are contemplated); Including, The intermediate mix profile is linearly predicted from the CCMC profile in step (a).

[0021] Mixing samples has the advantage that it can average out variability between samples from different donors or between samples from the same donor (variability between two donations made on different days) and can also make the starting CMC product for co-culture richer (in terms of diversity) than just the samples.

[0022] In some embodiments, predicting the CMC profile comprises: Predicting intermediate CMC profiles for selected complex microbial community sample mixes using a linear approach; and and correcting the intermediate CMC profile to a predicted CMC profile using a CMC interaction model learned from the reference linear predicted CMC profile and the corresponding reference true CMC profile. The CMC interaction model models the nonlinear behavior of sample mixtures. Advantageously, this modeling allows for instant and accurate simulation of various mix compositions from selected samples in a sample collection at low cost, particularly without consuming actual materials.

[0023] In some embodiments, the CMC interaction model and the mixed interaction model are one and the same model, e.g., the same matrix, which reduces processing and memory costs by requiring only one model training run to perform the entire prediction and storing only one model.

[0024] A different aspect of the present invention focuses on a new mechanism for determining appropriate initial samples to process, which does not rely on the invertibility of the mix generation model.

[0025] In this regard, this aspect of the present invention proposes a computer-assisted method for determining the population and mixture ratio of complex microbial community (CMC) samples in an initial sample collection, and producing a mix result product from the CMC samples using a mix production process configured with the mixture ratio, the method comprising: Obtaining an initial candidate population, each candidate representing a set and mixture proportion of CMC samples in the initial sample collection; applying an evolutionary algorithm to iteratively modify the candidate population based on a model of the mix production process and a target mix profile representing the target mix result product; and Selecting candidates for the modified population obtained from the evolutionary algorithm.

[0026] The inventors have found that evolutionary algorithms, such as genetic algorithms, provide accurate determination of production process parameters (initial samples and blending ratios). Therefore, this algorithm can be advantageously used when the production process model is irreversible. It also works with non-convex fitness score metrics, such as Bray-Curtis dissimilarity, and is easily adaptable (with small incremental complexity) to multiple iterations of the same model (i.e., production subcycle) within the overall production pipeline.

[0027] The determined set of samples and associated mixing ratios (forming the selected candidates) may then be used to control the actual extraction and processing of complex microbial community samples according to a mix production process, resulting in a mix result product that is as close as possible to the target mix profile (in terms of profiling characteristics).

[0028] The present invention also provides a method for producing a co-cultured complex microbial community (CCMC) result product, comprising: Using the above determination method based on the target mix profile, obtain candidates that represent the assemblage and mixture ratio of complex microbial community (CMC) samples in the initial sample collection; The actual extraction of the CMC samples of the set from the initial sample collection; and Processing the extracted CMC sample using a mix production process configured with the obtained blend ratio to obtain a CCMC result product; It can be seen that the present invention provides a method comprising:

[0029] The present invention allows for the production of CCMC resultant products that meet a target profile (e.g., configured to prevent or treat a disease or restore a necessary function after drug-induced dysbiosis) without unnecessary use of materials and in large quantities.

[0030] Therefore, prior to the production routine, a mixed production strategy can be defined according to the needs of the intended use (therapeutic, preventive, environmental, etc.).

[0031] The resulting CCMC products thus obtained can then be administered or implanted into humans or animals, or into plants as fertilizers, or into environmental media including water, soil, and subsurface materials for the treatment of contamination by bioremediation.

[0032] Preferably, the Microbiome Ecosystem Therapy product can be produced using the above method.

[0033] Optional features of embodiments of the invention are defined in the appended claims, some of which will be described herein below with reference to methods, but which may also be replaced by device / system features.

[0034] In some embodiments, the mix production process model comprises: Predicting the intermediate CMC profile for a mix of selected complex microbial community samples given the mixing ratio using a linear approach; and and correcting the intermediate CMC profile to a predicted CMC profile representing the complex microbial community (CMC) product obtained from the mix using a CMC interaction model trained from the reference linear predicted CMC profile and the corresponding reference true CMC profile.

[0035] This approach accurately models the first pooling stage of the mix production process, where samples are mixed, because the model includes a correction step that reflects the nonlinear behavior of mixing.

[0036] In some embodiments, each candidate includes a mixing ratio representing the respective proportions of CMC samples to be mixed. These embodiments thus allow such ratios to be determined (through an evolutionary algorithm) given a desired target mix profile.

[0037] In some embodiments, the mix production process model comprises one or more loops (or iterations) of: Predicting a plurality of CCMC profiles representing co-cultured CMC products obtained by separate co-cultures of the same starting CMC product from the predicted CMC profile or from the predicted mix result profile of the preceding loop; and Using a linear approach and from the CCMC profile, predicting an intermediate mix result profile representing a second mix of CCMC products given the mix ratio; and and correcting the intermediate mix result profile using a mix interaction model learned from the reference linear predicted mix profile and the corresponding reference true mix profile to form a predicted mix result profile representing a mixed CCMC product obtained from the second mix.

[0038] This approach accurately models the second stage of expanding the CMC product resulting from the mixing of the initial samples using parallel co-culture in a bioreactor.

[0039] In some embodiments, each candidate further comprises a blend ratio representing the respective proportions of CCMC products to be blended. The present invention thus allows such ratios to be determined (through an evolutionary algorithm) given a desired target blend profile.

[0040] In some embodiments, the same mixing ratios representing the respective proportions of CCMC products to be mixed are used throughout the loop, thereby reducing complexity at the candidate level. Of course, more complex algorithms involving different mixing ratios for successive loops may also be used.

[0041] In some embodiments, the CMC interaction model and the mixed interaction model are one and the same model, e.g., the same matrix, which reduces processing and memory costs by requiring only one model training run to perform the entire prediction and storing only one model.

[0042] Reflecting the above model, the mix production process may include a first pooling step in which extracted CMC samples are mixed to obtain a CMC product. The mixing is preferably performed according to the obtained candidate mixing ratio. The mix production process may also include one or more iterations of expanding the starting CMC product, where the iterations include (i) co-cultivating the CMC product or the resulting mix product obtained from the previous iteration in a bioreactor with the respective operating parameters to obtain a CCMC product, and (ii) mixing the CCMC products to obtain a resulting mix product. The mixing step (ii) is preferably performed according to the obtained candidate mixing ratio.

[0043] In some embodiments, each candidate is defined by a sequence of genes that includes a sample identifier and a mixture ratio, with each sample identifier and mixture ratio defining a separate gene.

[0044] In some embodiments, the iterations within the evolutionary algorithm include: assessing a score for each candidate in the current population based on the mix production process model and the target mix profile; Selecting a portion of the current population based on the assessed score; and and generating a new candidate population based on the selected candidates using genetic crossover between genes and / or genetic mutation within the gene array of the selected candidates.

[0045] In some embodiments, evaluating the score includes calculating a distance between the target mix profile and a mix result profile predicted from the candidate using the mix production process model.

[0046] In some embodiments, the profile of a complex microbial community (sample or CMC product or CCMC product) comprises the relative abundance of profiling features in the complex microbial community.

[0047] In certain embodiments, relative abundance represents the mass or volume ratio of the profiling features in a complex microbial community.

[0048] In some embodiments, the profiling features that constitute the profile of a complex microbial community include one or more features from taxon, gene, antibiotic resistance gene, function, metabolite trait, metabolite, RNA and protein production, preferably taxon.

[0049] In some embodiments, the individual profiles of complex microbial community samples are obtained using profiling techniques such as 16S rRNA gene amplicon sequencing, NGS shotgun sequencing, non-16S rRNA gene-based amplicon sequencing, NGS amplicon-based targeted sequencing, PhyloChip-based profiling, whole metagenomic sequencing (WMS), polymerase chain reaction (PCR) identification, mass spectrometry (e.g., LC / MS, GC / MS, or MS / MS mass spectrometry), near-infrared (NIR) spectroscopy, and nuclear magnetic resonance (NMR) spectroscopy. Preferably, 16S rRNA gene amplicon sequencing or NGS is used.

[0050] In some embodiments, the profile of a complex microbial community defines profiling features for one or more microorganisms from bacteria, archaea, viruses, phages, protozoa, yeast, and fungi present in the complex microbial community, preferably defining profiling features for bacteria and / or archaea.

[0051] In some embodiments, the profile of a complex microbial community defines profiling features that define the relative abundance of microorganisms considered at one or more taxonomic levels from strain, species, genera, families, orders, classes, and phyla, preferably at one or more taxonomic levels from genera, families, and orders.

[0052] In some embodiments, the profile of a complex microbial community comprises the relative abundance in the complex microbial community of bacterial and / or archaeal taxa considered at any taxonomic level within genus, family, order, class, and phylum.

[0053] In some embodiments, the profile of a complex microbial community comprises the relative abundance within the complex microbial community of bacterial and / or archaeal taxa defined by the presence / absence or expression of certain genes and / or functions (e.g., production of butyrate, antibiotic resistance genes, production of enzymes such as organophosphate hydrolases, phosphodiesterases, superoxide dismutases, production of antimicrobial peptides, organophosphate hydrolyases, or other enzymes useful in bioremediation processes, etc.).

[0054] In some embodiments, the initial sample collection comprises samples selected from the group consisting of unprocessed complex microbial community samples, manipulated / processed complex microbial community samples, artificial complex microbial community samples (e.g., bacterial consortia obtained by mixing isolated strains), samples containing genetically modified organisms (e.g., bacteria, archaea, phages, viruses), and hypothetical complex microbial community samples.

[0055] In some embodiments, the initial sample collection includes one or more of a fecal sample, a skin sample, an oral sample, a vaginal sample, a nasal sample, a tumor sample, a human sample, an animal sample, a plant sample, a water sample, a soil sample, etc. For example, the initial sample collection may include one or more fecal samples obtained from at least one donor, preferably from at least two donors.

[0056] Another aspect of the invention relates to a computing device including at least one microprocessor configured to perform the steps of any of the above methods. Thus, the computing device may be configured to emit signals to control a production device to actually extract and process a complex microbial community (CMC) sample from an initial sample collection using a mix production process to obtain a co-cultured complex microbial community (CCMC) result product.

[0057] Another aspect of the invention relates to a non-transitory computer readable medium storing a program which, when executed by a microprocessor or computer system in a device, causes the device to perform any of the methods defined above.

[0058] At least part of the methods according to the present invention may be computer-implemented. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "circuit," "module," or "system." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium of expression and having computer-usable program code embodied therein.

[0059] Because the present invention can be implemented in software, the present invention can be embodied as computer-readable code on any suitable recording medium for provision to a programmable device. Tangible recording media can include storage media such as hard disk drives, magnetic tape drives, or solid-state memory devices. Transient recording media can include signals such as electrical, electronic, optical, acoustic, magnetic, or electromagnetic signals, e.g., microwave or RF signals. [Brief explanation of the drawings]

[0060] [Figure 1] FIG. 1 illustrates a complex microbial community mixing platform for implementing an embodiment of the present invention. [Figure 1A] Figure 1A shows the behavior of the error measure as a function of the regularization term hyperparameter during modeling. [Figure 2A] FIG. 2A shows a schematic diagram of the sample first mix production process. [Figure 2B] FIG. 2B shows a schematic diagram of the second mix production process of the sample. [Figure 2C] FIG. 2C shows a schematic diagram of the third mix production process of the sample. [Figure 3A] FIG. 3A shows a schematic modeling of the first mix production process of FIG. 2A. [Figure 3B] FIG. 3B shows a schematic modeling of the second mixed production process of FIG. 2B. [Figure 3C] FIG. 3C shows a schematic modeling of the third mixed production process of FIG. 2C. [Figure 4] FIG. 4 illustrates a genetic algorithm designed for candidate determination given a target mix profile, according to an embodiment of the present invention. [Figure 5] FIG. 5 shows a schematic diagram of a computing device according to an embodiment of the present invention. [Figure 5A]FIG. 5A shows the results of a first experiment of the present invention, based on a mixture of untreated complex microbial community samples. [Figure 5B] FIG. 5B shows the results of a first experiment of the present invention based on a mixture of untreated complex microbial community samples. [Figure 5C] FIG. 5C shows the results of a first experiment of the present invention, based on a mixture of untreated complex microbial community samples. [Figure 6A] FIG. 6A shows the results of another experiment of the present invention based on the mixing of co-cultured complex microbial community samples. [Figure 6B] FIG. 6B shows the results of another experiment of the present invention based on the mixing of co-cultured complex microbial community samples. [Figure 6C] FIG. 6C shows the results of another experiment of the present invention based on the mixing of co-cultured complex microbial community samples. [Figure 7A] FIG. 7A shows the results of yet another experiment of the present invention in which untreated and co-cultured samples were mixed. [Figure 7B] FIG. 7B shows the results of yet another experiment of the present invention in which untreated and co-cultured samples were mixed. [Figure 8] Figure 8 shows a PCA based on the relative abundance of genera obtained from NGS shotgun sequencing of samples from experiment 2. [Figure 9] Figure 9 shows the PCA-based approach used in Experiment 2. [Figure 10A] Figure 10A shows the average baseline predictions from Experiment 3 using the leave-one-out method. [Figure 10B] Figure 10B shows the TO baseline predictions in Experiment 3. [Figure 11A] FIG. 11A shows the distribution of MSE (left) and Bray-Curtis (BC) distance (right) between each pair of target profile and corresponding predicted profile after genetic algorithm processing in Experiment 4. [Figure 11B]Figure 11B shows a comparison of one of the 100 target profiles with the associated predicted profile in Experiment 4. DETAILED DESCRIPTION OF THE INVENTION

[0061] The present invention relates to processes involving the mixing or "pooling" of complex microbial communities or CMCs representing complex microbial communities, or "microbiota" or "microbiota samples." More specifically, the present invention is directed to methods and devices that use trained predictive models to predict mixed co-culture complex microbial community (CCMC) products, as well as methods and devices that use trained predictive models to determine an initial complex microbial community sample given a target mix profile.

[0062] As used herein, the terms "microbiota," "microbiota composition," and "complex microbial community" or "CMC" can be used interchangeably to refer to any microbial population that includes a large number of different species of microorganisms living in symbiotic and potentially interactive relationships. Microorganisms that may be present in a complex microbial community include yeast, bacteria, archaea, viruses, fungi, algae, phages, and any protozoan of different origins, such as protozoans of soil, water, plants, animals, or humans.

[0063] The microbiome, as used herein, includes complex microbial communities of natural origin (e.g., gut microbiomes, i.e., the population of microorganisms living in the intestines of animals) as well as "engineered complex microbial communities," i.e., complex communities resulting from transformation processes such as the removal of potentially harmful microorganisms by additional processing of isolated beneficial strains (e.g., by using genes targeting rare-cutting endonucleases specific to pathogenic symbionts) or expansion by culturing under specific conditions (e.g., co-cultivation in an appropriate medium). "Isolated beneficial strains," as used herein, refer to naturally occurring strains known to have beneficial effects under certain conditions (e.g., Akkermansia muciniphila), as well as genetically modified strains, including strains in which potentially harmful genes have been knocked out (e.g., using rare-cutting endonucleases such as Cas9) and strains in which transgenes have been introduced (e.g., by using bacteriophages or CRISPR systems).

[0064] The complex microbial communities and microbiota referred to herein include "unprocessed" or "native" complex communities or microbiota, i.e., those obtained directly from a source, a donor, or multiple donors, that have not been processed by downstream processing, as well as engineered complex communities or microbiota, and "processed complex microbial communities", which include any complex microbial community resulting from processing or downstream processing of, or transformation of, one or more natural, unprocessed complex microbial communities (e.g., complex communities or microbiota that have been filtered, frozen, thawed, and / or lyophilized and / or extracted, isolated, or separated from their initial matrix by techniques well known to those skilled in the art, such as those described in WO 2016 / 170285 and WO 2017 / 103550).

[0065] The expressions "sample", "complex microbial community sample", "CMC sample", and "microbiota sample" can be used synonymously and refer to an initial complex community or microbiota in the sense of the present invention, i.e., available for processing, including mixing. The expressions "product", "mix", "complex microbial community product", "complex microbial community mix", "CMC product", and "microbiota mix" can be used synonymously and refer to an intermediate complex community or microbiota in the sense of the present invention, i.e., available for further processing, such as co-cultivation or further mixing. The expressions "microbiota" and "complex microbial community" refer to either the initial or intermediate complex community described above and can be used synonymously.

[0066] The term "Microbiome Ecosystem Therapy product" as used herein refers to any composition containing a complex microbial community (naturally occurring or engineered, raw or processed), provided that it is in a form suitable for administration to an individual in need thereof. Microbiome Ecosystem Therapy (MET) aims to achieve a health benefit by modifying an individual's microbiome (e.g., preventing or alleviating disease symptoms, increasing the likelihood that the individual will respond to treatment, etc.). Typically, Microbiome Ecosystem Therapy is performed in a subject in need thereof by replacing at least part of a dysfunctional and / or damaged ecosystem with a different complex microbial community. Microbiome Ecosystem Therapy includes fecal microbial transplantation (FMT). Unless otherwise specified, the term "FMT" is used broadly herein to refer to any type of Microbiome Ecosystem Therapy.

[0067] As shown in Figure 1, which illustrates a platform 1 for processing complex microbial communities implementing an embodiment of the present invention, samples 100 are available through an initial sample bank or collection 10. Although one collection or bank is shown, samples may be stored in multiple sub-banks that collectively form the collection or bank 10.

[0068] A sample of the present invention may comprise or consist of microorganisms obtained from one or more sources and / or from one or more donors 101.

[0069] Samples of the present invention may be obtained from: One source, At least two sources, One donor, At least two donors, one source and one donor, one source and at least two donors; At least two sources and one donor, or At least two sources and at least two donors.

[0070] As used herein, the term "source" refers to any environment from which a sample is obtained, such as soil, water, a part of a plant, a part of an animal's body or bodily fluid, or a part of a human body or bodily fluid. In the case of a human or animal, the source can refer to any body part (skin, nasal mucosa, etc.) or bodily fluid such as the contents of the intestine (e.g., a stool sample).

[0071] As used herein, the term "donor" refers to a plant, a physical location (a source such as soil or water), an animal, or a human, preferably a human.

[0072] Donors may be pre-selected according to methods and criteria described in the prior art, for example in WO 2019 / 171012 A1.

[0073] In the example shown, some samples, designated 100d, 100e, 100f, 100g, are unprocessed complex microbial communities or microbiota, i.e., obtained directly from one donor or multiple donors without subsequent processing.

[0074] Other samples, designated 100a, 100b, and 100c, are "processed samples," i.e., engineered complex microbial communities resulting from treatment or subsequent processing or transformation of one or more naturally occurring, unprocessed complex communities. As noted above, processing may include filtration, centrifugation, co-cultivation, freezing, lyophilization, or even mixing of the initial complex communities, but may also include treatments aimed at isolating spores and spore-forming bacteria, such as the use of ethanol, chloroform, or heat.

[0075] As shown, the initial composite community may be one of the samples 100d, 100e, 100f, 100g from the initial sample collection 10 or may be an external sample 99.

[0076] The initial sample collection 10 can include one or more samples from any source (fecal, skin, nasal, oral, vaginal, tumor, etc.) of any origin (human, animal, plant, soil, etc.), with one or more fecal samples obtained from at least one donor, preferably at least two donors.

[0077] According to certain embodiments, the samples in collection 10 include fecal samples.

[0078] The fecal sample collected from the donor can be controlled according to methods and quality criteria described in the prior art, such as WO 2019 / 171012(A1). For example, sample quality criteria can include sample consistency of 1 to 6 on the Bristol scale; absence of blood and urine in the sample; and / or absence of specific bacteria, parasites, and / or viruses, as described in WO 2019 / 171012(A1).

[0079] The fecal sample can be collected according to any method described in the prior art, such as WO 2016 / 170285 (A1), WO 2017 / 103550 (A1), and / or WO 2019 / 171012 (A1). Preferably, the sample is placed under anaerobic conditions after collection. For example, as described in WO 2016 / 170285 (A1), WO 2017 / 103550 (A1), and / or WO 2019 / 171012 (A1), the sample can be placed in an oxygen-tight collection device within 5 minutes of collection.

[0080] The sample may be prepared according to methods described in the prior art, for example in WO 2016 / 170285 (A1), WO 2017 / 103550 (A1), and / or WO 2019 / 171012 (A1).

[0081] All samples 100a-100g shown are actual samples stored in at least one bank.

[0082] Samples 100y-100z, represented by dotted lines, are hypothetical samples that were not actually collected from donors or created and therefore are not actually stored in storage bank(s) 10. As explained below, these "virtual" samples 100y-100z are drawn to illustrate theoretical composite community profiles 110z imagined by an entity (e.g., a computer, a worker, a researcher, etc.).

[0083] The initial sample collection 10 may include only unprocessed samples 100d-100g, only processed samples 100a-100c, only virtual samples 100y-100z, or a combination thereof.

[0084] The samples in the initial sample collection 10 are used to input into a mix production process with the goal of producing a mix result product, for example, in a microbiome-based therapeutic product process, to produce a substance with a mix result profile in terms of taxonomy, functionality aligned with a therapeutic target.

[0085] FIG. 2A shows a schematic diagram of the sample first mix production process.

[0086] The process begins by selecting (and extracting) a set Γ of 1 to k microbiota samples from an initial sample collection 10. The selected samples are denoted γ1 to γ k where γ i may simply be an identifier for the samples in the initial sample collection 10. Each sample in the initial sample collection has a known microbiota profile.

[0087] "Profile" refers to a description of the composition of a given complex microbial community or microbiota composition (sample or mix or product). For example, a profile specifies the relative abundance of profiling features in a complex community or microbiota composition. "Relative" means that the sum of the abundances equals 1. Relative abundance can be expressed as the mass ratio (or weight ratio) or volume ratio of the profiling features in a complex microbial community.

[0088] Depending on the application (e.g., depending on the target disease in the therapeutic field, or depending on the pollutant to be removed in the bioremediation field), the profiling features may be of different types. Typically, the profiling features are selected from the group comprising taxa, genes, antibiotic resistance genes, RNA, function, and metabolite traits, as well as metabolite and protein production. A profile may also mix different types of profiling features, e.g., taxa and antibiotic resistance genes. In certain embodiments, only taxa are considered for profiling complex microbial communities.

[0089] Functions can describe the known actions of a protein or protein family (defined phylogenetically, e.g., in databases such as KEGG KO, NCBI COG, or EC numbers), or they can define a metabolic context (e.g., BiGG models (reaction level) or KEGG pathways (metabolic pathway level)); some databases are specialized, e.g., the CaZy database is a catalog of carbohydrate-active enzymes. Any of these functional categories (or a combination of them) can be used as features in matrix models.

[0090] KEGG stands for "Kyoto Encyclopedia of Genes and Genomes," KO stands for "KEGG Orthology," NCBI stands for "National Center for Biotechnology Information," COG stands for "Cluster of Orthologous Groups," and BiGG stands for "Biochemical Genetic and Genomic."

[0091] Various profiling techniques for obtaining complex community profiles are known, including 16S rRNA gene amplicon (i.e., metagenomic) sequencing, NGS shotgun sequencing, non-16S rRNA gene-based amplicon sequencing, NGS amplicon-based targeted sequencing, 18S / ITS gene sequencing, metagenomic sequencing, PhyloChip-based profiling, polymerase chain reaction (PCR) identification, mass spectrometry (e.g., LC / MS, GC / MS, or MS / MS type mass spectrometry), near-infrared (NIR) spectroscopy, and nuclear magnetic resonance (NMR) spectroscopy.

[0092] In the process of Figure 2A, k samples are mixed with each other in a mixture ratio M=[α1...α k ] (Σα=1) to obtain a mixed result product.

[0093] By "mixing" is meant any actual mixing of samples resulting in a new CMC product or a new microbiota composition. This result is also referred to as a mixed result product, as it may be used for administration or transplantation as described above. The mixed result product may be used, for example, as an MET inoculum.

[0094] Figure 2B shows a schematic diagram of a second sample mix production process, which is a co-culture-based mix production process in that it includes a co-culture operation as described below.

[0095] The process begins with selecting (and extracting) a set Γ of 1 to k microbiota samples from the initial sample collection 10, as described above.

[0096] k samples are mixed in the following ratios M=[α1...α k ] (Σα=1), resulting in a CMC product. The CMC product is therefore an intermediate product, or "pooling product," in the overall mix production process.

[0097] The CMC product is then transferred to j bioreactors F1, ..., F j (j is an integer greater than 1), j co-culture products or CCMC products are obtained.

[0098] The term "bioreactor" refers to any device containing a vessel useful for co-cultivating microorganisms, allowing for control of culture parameters (temperature, pH, retention time, aeration, feeding, etc.). In the present invention, the terms "fermenter" and "bioreactor" have the same meaning. A bioreactor can reconstruct the environment from which a complex microbial community was obtained by integrating key environmental parameters. For example, if a complex microbial community is obtained from an in vivo human colon environment, the bioreactor can integrate in vivo human colon environment parameters such as pH, temperature, ileal waste feeding, retention time, and anaerobic conditions (anaerobiosis), as described in the prior art, for example, in Cordonnier et al., 2015 (Dynamic In Vitro Models of the Human Gastrointestinal Tract as Relevant Tools to Assess the Survival of Probiotic Strains and Their Interactions with Gut Microbiota, Microorganisms. 2015 December; 3(4): 725-745).

[0099] The same starting CMC product is fed to each of the j bioreactors. "j" can be 2 or more, 3 or more, or equal to 4. The bioreactors have at least one different operating parameter. The operating parameter is preferably selected from pH, temperature, pressure, incubation time, retention time, gassing conditions, redox potential, culture medium, light source, and combinations thereof. As an example, the j bioreactors have different pH setpoints. The co-cultivation process in the bioreactors can last for multiple hours or even multiple days, for example, 3 days.

[0100] In the process of Figure 2B, the CCMC products are synthesized in a controlled ratio B = [β1, β2, ..., β j ], resulting in a mixed result product.

[0101] Further details regarding the co-cultivation process followed by mixing are provided in WO 2022 / 136694.

[0102] Figure 2C shows a schematic representation of a third sample mix production process, which, compared to the second process, loops two or more co-cultivation steps, where the final mix result product is obtained from the repeated mixing of CCMC products obtained by separately co-cultivating the same starting CMC product (for each repetition) in each bioreactor.

[0103] The process begins with selecting (and extracting) a set Γ of 1 to k microbiota samples from the initial sample collection 10, as described above.

[0104] k samples are mixed in the following ratios M=[α1...α k ] (Σα=1), resulting in a CMC product. The CMC product is therefore an intermediate product, or "pooling product," in the overall mix production process.

[0105] The CMC product is generated from j bioreactors F1, ..., F j (j is an integer greater than 1) are fed to a multi-iteration (or multi-loop) co-culture process. Preferably, the same j bioreactors are used during the co-culture iterations. However, this is not required; different numbers of bioreactors (j1, j2, ..., jN, where N is the number of iterations) can be used. Also, the bioreactors can have different characteristic operating parameters or different characteristic parameter settings during the iterations.

[0106] In the first co-culture iteration, the CMC product is fed to j bioreactors to obtain j first-iteration co-culture products or CCMC products. The CCMC products of the first iteration are fed to j bioreactors to obtain j first-iteration co-culture products or CCMC products at the controlled respective ratios B1 = [β 1,1 , β 1,2 , ..., β 1,j1], resulting in a first iteration intermediate mix result product or "mixed CMCC product."

[0107] In the ith co-culture iteration (i=2...N), the intermediate mix result product or "mixed CMCC product" obtained from the (i-1)th iteration becomes the new starting CMC fed into the bioreactor used in this iteration, resulting in the ith iteration CCMC product. The ith iteration CCMC product is then mixed with the controlled respective ratios B i =[β i,1 , β i,2 , ..., β i,ji ], resulting in an intermediate mix result product or "mixed CMCC product" of the iteration.

[0108] Preferably, the same j bioreactors and mixing ratios are used during the iterations, i.e., regardless of (i, m), B i =B m (denoted by the symbol "B"), which is the example in Figure 2C.

[0109] The description of FIG. 2B regarding the bioreactor and co-cultivation steps also applies to each iteration of the third mix production process.

[0110] The intermediate mix result product or "mixed CMCC product" of the Nth iteration is the final mix result product.

[0111] In some implementations of the third mix production process, additional samples from the initial collection may be added in one or more iterations of the process. For example, a new sample is added during the mixing of the CMCC products in iteration i. This new sample may have its own ratio β, which is added to the other mixing ratios. The new sample may be added in each or some of the iterations, for example, in the third or subsequent iterations, or in every other iteration, or in the final iteration. The same additional sample may be added throughout the iterations, or different samples may be added throughout the iterations. Of course, more than one additional sample may be added in a given iteration.

[0112] The present invention uses models, such as trained models of the mix production process. In particular, models may be designed and trained "block-by-block," i.e., for the mixing and co-cultivation operations separately, as shown in Figures 3A-3C.

[0113] 3A shows a schematic representation of a model of the first mix production process shown in FIG. 2A. The model includes only one trained model 130, called the pool predictor, which is used to predict the resulting mix product in terms of its profile from the samples Γ used in the mix and their respective proportions M (mixing ratios) in the mix. As described below, this model can advantageously predict multiple profiles in a single calculation.

[0114] The known profile of the sample is shown schematically through reference 110, while the predicted profile of the resulting mix product (here a CMC product) is indicated by the symbol R'.

[0115] This modeling involves predicting the mix composition resulting from the mixing of samples 100a-100z belonging to the initial sample collection 10. The prediction involves two aspects: Predicting the intermediate CMC profile for a selected complex microbial community sample Γ mix using a linear approach; and and correcting the intermediate CMC profile to a predicted CMC profile using an interaction model trained from the reference linear predicted CMC profile and the corresponding reference true CMC profile, preferably a square interaction matrix trained from the reference linear predicted CMC profile and the corresponding reference true CMC profile.

[0116] Such trained interaction models, more specifically matrix-based models, provide accurate prediction results once the interaction model or matrix is ​​trained, thus providing precise hints about the final product without consuming any material from the initial sample collection.

[0117] Because the prediction is computer-implementable, predicted CMC profiles can be obtained quickly no matter how many mixes are predicted, how many samples are available in the initial sample collection 10, or how many features are used to profile the complex microbial community (samples and mixes).

[0118] As shown in Figure 1, a profiler (or sequencer) 12 is preferably used to provide profiles (e.g., 16S sequencing) of actual samples 100a-100g. The corresponding individual profiles thus obtained are designated 110a-110g and form an initial profile collection or bank 11. Of course, 16S rRNA sequencing is not required, and other methods, as defined above, can be used alone or in combination to provide the profile 110.

[0119] The individual profiles are converted to the same format, whatever the analytical method used, into a matrix or vector a x The coefficients a of the individual profile "x" are stored in the computer memory (not shown). x (j) indicates the relative abundance of profiling feature "j" in the sample of interest.

[0120] As mentioned above, several individual profiles 110z may be artificially constructed by an operator, for example, by a coefficient a that represents the relative abundance of profiling feature “j” in a theoretical sample. x It may be constructed by defining (i).

[0121] The initial profile collection 11 may therefore include only the individual profiles 110d-110g corresponding to the unprocessed samples 100d-100g, or only the individual profiles 110a-110c corresponding to the processed samples 100a-100c, or only the virtual profiles 110y-110z corresponding to the virtual samples 100y-100z, or any combination thereof.

[0122] Any other profiles treated subsequently (e.g., so-called intermediate or mixed profiles) follow the same profile format (e.g., a vector of the same profiling features "j" in the same order).

[0123] Preferably, a bacterial abundance profile is obtained, meaning that this profile defines the relative abundance of profiling features related to bacteria. More broadly, a profile of a complex microbial community may define profiling features related to one or more microorganisms (bacteria, archaea, viruses, phages, protozoa, yeast, algae, and fungi) present in the complex microbial community, preferably related to bacteria and / or archaea. Of course, profiling features within the same profile may relate to different microorganisms, such as those listed above.

[0124] Preferably, a genus-based bacterial abundance profile is obtained, i.e., the profiling features describe the relative abundance of bacteria at the genus level in a complex microbial community. More broadly, a profile of a complex microbial community may define profiling features that define the relative abundance of microorganisms considered at one or more taxonomic levels from strain, species, genera, families, orders, and phyla, preferably at one or more taxonomic levels from genera, families, orders, and phyla.

[0125] The prediction operation is performed by predictor module 13 (implementing pool predictor 130) in the first mix production process under the control of module 14. Module 14 is called the "test and decision module" or "decision module" and operates platform 1 with the objective of predicting a mix profile and / or determining a set of samples given a target mix profile and / or producing at least one mix or CCMC result product.

[0126] Modules 13 and 14 are preferably implemented via a computer having an input / output or user interface (e.g., keyboard, mouse, screen) to interact with platform 1. The target mix profile may be defined by an operator interacting with module 14 using the user interface.

[0127] As shown in FIG. 3A, the pool predictor 130 is matrix-based and includes two steps to predict the resulting mix profile from the initial profiles of the mixed samples.

[0128] Matrix A defines the individual profiles of all samples available in collection 10. Matrix A can be formed by the profiler or sequencer 12, or at least by the individual profiles obtained from the profiler. Furthermore, any virtual individual profiles can be added to the matrix. Preferably,

number

[0129] The square matrix W is the pooling or CMC interaction matrix defined above, which models the interactions between microorganisms. A description of the modeling matrix W (including the training method) is provided in more detail below. The pooling interaction matrix is ​​intended to represent the nonlinear interactions between the various profiling features of samples when they are mixed.

[0130] The prediction operation includes a first matrix-based step of predicting, for at least one mix of selected samples, an intermediate mix profile formed by matrix I using matrix A: I=P*A, where P is a matrix representing at least one mix of selected samples from collection 10.

[0131] The matrix P can define each mix in terms of mass or volume ratios of the samples in the initial sample collection. P may also define a single mix if desired. For example,

number

number

[0132] If sample r is not used in mix x, then p x (r) = 0. In other words, for the mix, for each sample index r that is not in Γ, p x (r)=0.

[0133] The matrix-based approach has the advantage that it allows us to jointly predict a variable number of mixes: each row of P defines a mix to predict (so in the example above there are t mixes defined), and this number 't' can vary from prediction to prediction.

[0134] The prediction operation I=P*A is, for example, computer-implemented.

number

[0135] Therefore, according to the invention, the prediction operation includes a second step of correcting the intermediate CMC profile (i.e., matrix I) using an interaction model, in particular the pooling interaction matrix W, to obtain a predicted CMC profile, which is represented by the following matrix:

number

[0136] Predicted CMC profiles can thus be obtained quickly for a variety of mixes without consuming any of the material in the collection 10 .

[0137] Relative abundance r xIt is expected that the (j) are non-negative and together form a perfect composition (i.e., their sum equals 1 for a given mix "x"). However, this may not be the case in matrix multiplication. Embodiments of the present invention therefore include post-processing the result matrix R into R' to satisfy biological constraints.

[0138] For example, each negative value in R is clipped, meaning that its negative abundance is set to 0. The relative abundances rx(j) are then normalized, i.e., adjusted (e.g., using linear interpolation) to r′x(j) so that their sum equals 1:

number

number

[0139] The efficiency of pooling prediction comes from modeling the actual positive and negative interactions between microorganisms in a mixed sample into a matrix called the pooling interaction matrix W. A two-step matrix-based process is then efficiently used to predict the observed CMC profile.

[0140] The pooling interaction matrix W is trained for a given set of m profiling features. If the profiling features in a profile are reordered, the coefficients of the pooling interaction matrix W should be reordered accordingly.

[0141] The m profiling features may evolve over time, for example, as new features are discovered, some features are removed because they are no longer meaningful, and / or some features are split into multiple features to become more accurate. Evolution of profiling features can also occur due to enhancements in profiling / sequencing methods and profilers / sequencers12 that provide new profiling data, as well as improvements in bioinformatics methods that combine algorithms and reference databases of features.

[0142] Different sets of profiling features may also be considered, such as for different diseases or targeted treatments.

[0143] Not only the profiling features themselves, but also the number of features in the set can evolve and change.

[0144] Thus, each time a new set of profiling features is considered, the pooling interaction matrix W may be calculated anew, similar to the matrix A describing the initial profile collection 11. The calculated pooling interaction matrix W may be stored in the memory of the pooling predictor 130 so that it can be reused when the corresponding set of profiling features is newly used.

[0145] The pooling interaction matrix is ​​preferably obtained using machine learning, which is performed using a training data set, which is a set of multiple mixes {p ref (k)}.

[0146] The actual reference sample mix is ​​homogenized for a period of 10 minutes to 3 hours, preferably 30 minutes to 1.5 hours. Homogenization is carried out at a temperature between 0°C and 10°C, preferably between 2°C and 6°C, more preferably at about 4°C.

[0147] The mix is ​​then considered stable for several hours, at least up to 16 hours from mixing, and preferably up to 24 hours from mixing.

[0148] This means that the pooling interaction matrix represents the interactions that would occur between the microorganisms in a stabilized mix at 4°C.

[0149] Other pooling interaction matrices may be generated that represent other mixing conditions.

[0150] For sample x, the individual profile {a x (j)} j (j=1...m) are either known or obtained from a sequencer profiling sample x. Therefore, using the linear equation I=P*A above, we can obtain a reference linear predicted CMC profile {i ref (j)} j also becomes known.

[0151] The profile of the reference CMC product "ref" is the reference CMC profile {r true (j)} j , which is also known or obtained from a sequencer profiling a reference CMC product "ref."

[0152] Reference predicted CMC profile {r pred (j)} j is the reference linear prediction CMC profile {i ref (j)} j corresponds to the matrix product of R and the square pooling interaction matrix W (during training): for a single reference CMC product “ref”, pred =I ref *W or {r pred (j)}j ={i ref (j)} j *W.

[0153] Machine learning attempts to minimize the error in predicting the reference CMC profile, i.e., the true CMC profile for the reference.

number

number

[0154] Training data for machine learning is ref and R true In some embodiments, the equation to minimize is the residual vector {r true-i (j)} j -{r pred-i (k)} k ={r true-i (j)} j -{i ref (k)} k *W, or residual matrix R true -R pred =R true -I ref *W.

[0155] Any norm may be used: L1, L2, Lp, etc. Preferably, the sum of squared errors (SSD) or its derivative, the mean squared error (MSE), may be used. Alternatively, the least chi-square method may be used.

[0156] Machine learning may then attempt to solve the following convex optimization problem:

number

number

[0157] In an embodiment that avoids overfitting of W, the formula adds a regularization term, preferably a Ridge (L2-based) regularization term, to the difference. In a variant, a Lasso (L1-based) regularization term can be used. In another variant, an Elastic Net-based regularization term (a regularized regression method that linearly combines the Lasso and the L1 and L2 penalties of the Ridge methods) can be used. The Ridge approach advantageously allows for a larger number of non-zero coefficients in W, which helps to more accurately model interactions between profiling features.

[0158] Therefore, machine learning attempts to solve the following convex optimization problem:

number

number

[0159] In addition, R pred Constraints may be imposed during machine learning such that I do not have negative relative abundances and the sum of the relative abundances of each reference predicted CMC profile is 1. In other words, I ref * Clipping negative relative abundances in W, and then each reference predicted CMC profile (i.e., I ref *The modified matrix R' corresponds to normalizing the sum of the relative abundances of each element (each row in W) to 1. pred is preferably used. Corrected I ref *W is

number

number

[0160] The set of training data (say N reference CMC products) is split into two subsets: one for optimizing the hyperparameter λ and the other for optimizing W.

[0161] Various methods are known for optimizing λ, including in particular approaches that minimize an information criterion (e.g., minimizing the Akaike Information Criterion or the Bayesian Information Criterion) and approaches that minimize cross-validation residuals, which use an initial subset of the training data. For this optimization, W may be set to be different from ID by default.

[0162] For example, if λ is 10 -5 From 10 3 The MSE of the above formula when varying between is calculated for the training and test datasets (splitting the subset for optimizing the hyperparameter λ). The resulting MSE is shown in Figure 1A.

[0163] As shown, when λ is small, the MSE of the training data set is close to 0, while the MSE of the test data set is very high. In this situation, the model is overfitted. On the other hand, when λ is high, the model is underfitted. Therefore, λ may be selected to minimize the MSE of the test data set.

[0164] Once λ is known, a second subset of the training data is used to learn W by minimizing the cross-validation residual: a k-fold algorithm is performed.

[0165] A subset of the training data (i.e., {r true-i (j)} j and {i ref-i (j)} j ) is divided into k subsets, preferably k is selected from an integer between 3 and 20, preferably between 4 and 10, and more preferably is equal to 5.

[0166] Each of the k subsets is selected in turn in a round-robin fashion (cyclic order) to define a test subset, and the k-1 remaining subsets define training subsets.

[0167] In each of the k rounds, the model is trained using the training subset, i.e., to find W,

number

[0168] The learned pooling interaction matrix W is then validated on a test subset, which is used to model the matrix-based model R. true =I ref *Applies to W. Score based on any norm, e.g., MSE

number

[0169] This operation is repeated for each of the k test subsets, resulting in k scores.

[0170] The learned pooling interaction matrix W corresponding to the best score (i.e., the lowest score) may then be selected to construct the pooling predictor 130.

[0171] Of course, other machine learning methodologies can also be used as long as they result in a learned pooling interaction matrix W.

[0172] In some embodiments, the profiling features of sample 100 (i.e., those used to form matrix A) are the same as the profiling features of the final mix result (i.e., those used to form matrix R). As described above, the profiling features can be taxa, genes, antibiotic resistance genes, functions, RNA, metabolite traits, and metabolite and protein production.

[0173] In other embodiments, the profiling features of sample 100 (i.e., those used to form matrix A) are different (partially or wholly) from the profiling features of the final mix result (i.e., those used to form matrix R). Any of the profiling features described above (taxon, gene, function, etc.) may be used.

[0174] As an example, when a profiling technique such as NGS shotgun sequencing is used, a larger number of profiling features are obtained per sample 100 compared to 16S sequencing. Thus, the sample 100 may be profiled using NGS shotgun sequencing (hence matrix A is formed of the NGS shotgun profiling features), while the final mix result may be maintained with a reduced number of profiling features, such as those obtained using 16S sequencing (hence matrix R is formed of the 16S profiling features). In that case, matrix I is formed of the NGS shotgun profiling features, and the pooling interaction matrix W is not a square matrix and still models interactions between microorganisms, but in this example, as relationships between the NGS shotgun profiling features and the 16S profiling features.

[0175] In certain embodiments seeking to reduce the large number of NGS shotgun profiling features, principal component analysis (PCA) is performed to project the large number of features into k principal components (k PCs). In one embodiment, PCA is performed on the features profiling the sample, i.e., to construct matrix A. In another embodiment, matrix I is generated with the large number of profiling features, and PCA is performed on matrix I.

[0176] As mentioned above, the pool predictor 130

number

number

[0177] Additional details regarding Poole predictors can be found in co-pending International Application No. PCT / EP2022 / 062226. In particular, experimental results regarding modeling Poole predictors are provided in this application, which are incorporated herein by reference.

[0178] 3B schematically illustrates modeling of the second mix production process of FIG. 2B. In this case, the predictor module 13 includes a first pool predictor model 130, a fermenter predictor model 131, and another mixture model (typically another pool predictor model 130). The other model may differ from the first pool predictor 130 because it reflects the mixture of CCMC products. For example, the other model may be trained using only the set of CCMC product profiles (predicted and true) and therefore may result in a different pooling interaction matrix W. However, for ease of implementation (saving memory and training costs), the two pool predictor models may be the same trained model.

[0179] For example, the initial mix {p1(k)} with 100 initial samples k is fed into the pooled predictor model 130 to obtain the predicted CMC profile {r′1(l)} l This predicted CMC profile is fed to the fermenter predictor model 131 to obtain j CCMC profiles: {r2 i (l)} l (i=1...j). This CCMC profile is then fed into a second pool predictor 130 to obtain the predicted mix composition of the final mix result product.

[0180] The fermenter predictor model 131 is used to predict CCMC products from a profile perspective from an input or starting complex microbial community, particularly from the CCMC products of an initial mix of samples selected from an initial sample collection.

[0181] As shown, the fermentor predictor 131 is matrix-based.

[0182] The input matrix Q describes the profiles of the starting CMCs (one per row if there are multiple), together with the corresponding growth conditions, i.e., the operating parameters of the bioreactor.

[0183] Submatrix Q1 describes the profile (e.g. taxon abundance, values ​​between 0 and 1):

number

number

[0184] The CCMC prediction procedure uses the co-culture interaction matrix Y to predict one or more starting CMC products and their respective growth conditions (q1 to q t ), the matrix

number

[0185] In the second mix production process, the same starting CMC product is fed to j bioreactors with different operating parameters. This means that the j rows of matrix Q are the same starting CMC (q = q = ... = q j ) is used to describe the different operational parameters (m, n) in any case belonging to 1...j. m ≠q2 n ) Of course, a different number of rows may be used in a single calculation, for example to predict the CCMC product (and therefore its profile) of different starting CMC products passing through a bioreactor.

[0186] Similar to the pooled predictor, the relative abundance r x (j) are expected to be non-negative and to collectively form a perfect composition (i.e., their sum equals 1 for a given mix "x"). Therefore, each negative value in R2 may be clipped. Non-zero relative abundances (non-zero values ​​in R2) of profiling features not present in the starting CMC product may be set to zero. Relative abundance r2 x (j) may be normalized.

[0187] The co-culture interaction matrix Y represents the interactions between the profiling features (taxa) given the operational parameters. The co-culture interaction matrix Y is preferably obtained using machine learning, which is performed using a training data set. The training data is constructed based on a reference CCMC product "ref2" obtained from one or more reference complex microbial communities combined with growth conditions (i.e. operational parameters): {q ref}.

[0188] The following mix production process can be used to train Y. Co-culture runs of the starting CMC product can be performed in at least two bioreactors, preferably three, with the bioreactors having at least one different parameter selected from pH, temperature, pressure, incubation time, retention time, gas supply conditions, redox potential, culture medium, light source, and combinations thereof. For example, single different setpoints can be used. The pH setpoints of the bioreactors can be selected between 4.5 and 8.0, preferably between 5 and 7.6, and more preferably between 5.2 and 7.4. Other parameters of the different bioreactors are fixed. Co-culture runs can be performed simultaneously or with a time delay, and the resulting CMC products can be mixed to feed a new co-culture run, serving as the CMC product once or multiple times. After the co-culture run is completed, the resulting CMC products are sequenced (e.g., 16S, shotgun sequencing, etc.). The data thus generated is used to train the associated Y matrix.

[0189] Other co-culture interaction matrices may be generated that represent other growth conditions.

[0190] Individual vector q ref is constructed by simply concatenating the known individual profiles of the starting CMC products and the growth conditions.

[0191] The CCMC profile of the reference CCMC product “ref2” is the true CCMC profile for reference {r2true (j)} j , which is also known or obtained from a sequencer profiling a reference CCMC product "ref2."

[0192] Reference predicted CCMC profile {r2 pred (j)} j is the individual vector {q x (j)} j and the co-culture interaction matrix W (during training): For a single reference CCMC product “ref2”, R2 pred =Q ref *Y or {r2 pred (j)} j ={q x (j)} j *Y.

[0193] Similar to the pooled predictor, the machine learning attempts to minimize the error (e.g., MSE) in predicting the reference CCMC profile. For example, the error to be minimized is:

number

number

number

number

[0194] Training data for machine learning is Q ref and R2 trueand may be split into subsets to learn Y, e.g., in a round-robin fashion, as explained above.

[0195] The learned co-culture interaction matrix Y corresponding to the best score (i.e., the lowest score) may then be selected to construct the fermenter predictor 131.

[0196] In use, the output R2 of the fermentor predictor 131 is fed to the second pool predictor 130. The second pool predictor 130 thus predicts the mix composition resulting from the mixing of the same complex microbial community (CMC) products obtained by separately co-cultivating the CMC products in respective bioreactors, the method comprising: (a) predicting intermediate mix profiles for CCMC product mixes using a linear approach; and (b) correcting the intermediate mix profile to a predicted mix profile using a mix interaction model learned from the reference linear predictive mix profile and the corresponding reference true mix profile.

[0197] Upfront, cascading the pool predictor 130, the fermenter predictor 131, and then the pool predictor 130 ensures that the overall prediction includes: predicting the CMC profile of a CMC product based on the sample profile of a complex microbial community (CMC) sample selected from the initial sample collection (by the first pool predictor 130), and predicting the CCMC profile of a CCMC product based on the CMC profile (by the fermenter predictor 131).

[0198] Considering this cascade, the intermediate mix profile (by the second pooling predictor 130) is linearly predicted from the CCMC profile.

[0199] Thus, the method predicts the mix composition resulting from the second mix production process of Figure 2B.

[0200] 3C shows a schematic representation of the modeling of the third mixed production process of FIG. 2C. In that case, the predictor module 13 includes a first pool predictor model 130 and a loop formed by a fermenter predictor model 131 and another mixed model (typically another pool predictor model 130). As explained above, the two pool predictors may be different. However, for ease of implementation (saving memory and training costs), the two pool predictor models may be one and the same trained model. In a variant, instead of having a loop between the fermenter predictor model 131 and the pool predictor model 130, the modeling may be performed by a loop formed by the fermenter predictor model 131 and the pool predictor model 130. i and pooled predictor model 130 i This may allow different pooling (or fermentor) interaction matrices to be learned.

[0201] As shown, during calculation of the predicted mix composition of the final mix result product, the predicted product profile obtained from the second pool predictor 130 is fed back to the fermenter predictor 131 for the next loop until N (predetermined) loops are reached. In other words, the method for predicting the mix composition includes multiple iterations of steps (a) and (b) above, where the predicted mix profile obtained in step (b) of the previous iteration is used to obtain the profile of each CCMC product for step (a) of the next iteration.

[0202] In some embodiments not shown, the second pooling 130 may optionally add one or more profiles of CMC samples from the initial collection if it is anticipated that other (e.g., all) initial CMC samples will be added during one or more co-culture iterations, with this or these additional samples also having their own index γ and mixing ratio β.

[0203] Returning to Figure 1, as mentioned above, predictor module 13 is controlled by module 14. Module 14 can therefore configure predictor module 13 according to any of the models of Figures 3A-3C depending on which mix production process (Figures 2A-2C) is implemented. The user interface of platform 1 may allow an operator to input the mix production process to be used.

[0204] Module 14 includes a sample and ratio determination module 140. This module is designed to determine a set of complex microbial community (CMC) samples within an initial sample collection and their mixture ratios, so that a mix result product can be produced from the CMC samples using a mix production process configured with the mixture ratios. The determination is based on a target mix profile that represents a target mix result product.

[0205] Module 14 also includes a product generator control module 141. This module is designed to control the product generator module 15, which is used to actually produce the mix or CCMC result product from the determined sample and mix ratios using the mix production process.

[0206] The sample and ratio determination module 140 attempts to estimate or determine production parameters of the mix production process to obtain a final CCMC result product having a profile that is as close as possible to the defined target mix profile in terms of the defined profiling features (such as classification hierarchy or functional features).

[0207] As is clear from the examples of FIGS. 3A to 3C, the generation parameters to be determined include the set of initial samples Γ={γ1...γ k}, the mixture ratio for each pooling M=[α1...α k ] (preferably the same for all poolings), and the mixing ratio B = [β1...β j] (if multiple, preferably the same for all co-culture steps: Figure 3C). The constraints on these data are as follows: γ i is an integer ⊂ [1...N], where N is the total number of available initial samples 100; α i ⊂[0...1] and β i ⊂[0...1].

[0208] In operation, Σα i and Σβ i is equal to 1. However, during the calculation of candidates, as explained below, these sums may deviate from such a condition. In this case, an optional constraint can be provided to add a penalty to the calculation of the fitness score, lowering the fitness score. To do so, these sums (Σα i or Σβ i ) is strictly less than 1, a corresponding penalty is calculated. For example, penalty = 0.001 * Σα i (or Σβ i ). Similarly, the sum of these (Σα i or Σβ i ) is strictly greater than 1, the corresponding penalty is also calculated. For example, penalty = 0.001 * (2 - Σα i ) (or 2-Σβ i ). This penalty may be adjusted as a hyperparameter of the production parameter determination model. For example, its value may be adjusted by conducting a variation study of such hyperparameter with respect to the MSE used (e.g., in a manner similar to that performed above with reference to FIG. 1A).

[0209] Hereafter, a possible set of generation parameters is called a "candidate" and is denoted by c i Therefore, the candidate is a sequence of genes, and the gene is formed by the sample identifier and the mixture ratio: for example, c i =[γ k ]|[α k ]|[β k ]. γ k is coded as an integer, and α k and βk may be coded as a float.

[0210] The sample and ratio determination module 140 therefore searches for one good candidate (given the target), or in some embodiments, the "best" candidate. If desired, the module 140 may obtain multiple candidates that are satisfactory for controlling the product generator module 15. Indeed, depending on the amount of material (samples 100) available in the initial sample collection 10, it may be worthwhile to drive the product generator module 15 with multiple different sample sets to obtain a sufficient amount of final mix / CCMC result products that all have fairly similar profiles (close to the target profile).

[0211] Preferably, the sample and ratio determination module 140 implements an evolutionary algorithm, for example a genetic algorithm. In practice, the module 140 first obtains or builds an initial candidate population. Each candidate represents a set and mixture ratio of CMC samples in the initial sample collection. The initial population has a predetermined size SIZ (number of candidates). The module 140 therefore accesses the initial sample collection 10 with a characterization profile for each sample. Each sample is assigned an index γ i The sequence of genes in each candidate corresponds to a "chromosome" in the theoretical definition of evolutionary algorithms, and each data (γ i , α i , β i ) can be considered a "gene."

[0212] In some embodiments, the initial population is constructed in a random manner, where SIZ sample sets are randomly selected in the range [1...N], where each set is n min ~n max samples), and the associated mixture ratio α i , β i This means that the values ​​are also randomly selected within a range of possible values.

[0213] Module 140 then applies an evolutionary algorithm to iteratively modify the candidate population based on a model of the mix production process and the target mix profile. Each candidate can breed with another candidate to form offspring candidates for the next population, which can then undergo mutation and modification. After each new population is generated, the fitness of each candidate in the new population is evaluated based on a fitness score that reflects its closeness to the target mix profile.

[0214] The module 140 then selects candidates for the modified population that results from the evolutionary algorithm.

[0215] FIG. 4 shows the genetic algorithm designed for candidate determination given a target mix profile.

[0216] Algorithm settings include: The initial population constructed above, (The population size SIZ may be predefined, e.g., 100, 200, 500) Sample {a i (j)} j 110 individual profiles of Optionally, a correspondence matrix mapping profiling feature j to higher taxonomic ranks (this matrix is ​​used below to evaluate the distance between the predicted mix result profile and the target profile, when the target profile defines the condition for high taxonomic rank), A model of the mix production process, i.e., the CMC interaction matrix(ies) W and the CCMC interaction matrix(ies) Y, if any; the objective function to be minimized by the algorithm (mean squared error (MSE) can be used), Goal profile, Algorithm constraints, e.g. *The minimum number of initial samples that form each candidate, n min and the maximum number n max(For example, they may be 2 and 100, respectively. Of course, other minimum numbers from 3, 4, or 5 to 10 may also be used, and other maximum numbers such as 10, 20, 30, 50, or more may also be used.) * Algorithm termination criteria (e.g., maximum number of algorithm iterations or fitness score threshold) (maximum number of iterations may be 100, 500, 1000, or 1500); *Any gene of any candidate in each individual solution (γ i , α i , β i ) is changed (e.g., replaced with a random value), the mutation probability (by default, this probability may be set to 0.3, i.e., 30%), * Crossover probability, which sets the proportion of the next generation population that will be produced by crossover operations within the selected candidates, i.e., by mixing the genes of two selected candidates (by default, this probability may be set to 0.5, i.e., 50%); * Parent ratio, which sets the number of candidates selected for breeding to build the next generation (by default, this ratio may be set to 0.3, i.e. 30%), * The crossover type, which sets how the gene exchange will proceed (for example, it may be a "uniform" crossover, which randomly selects the genes to be exchanged between pairs of candidates; in particular, each gene in a new offspring candidate may be selected from either parent candidate with equal probability); *Optionally, an elite ratio (by default, this ratio may be set to 0.01, i.e., 1%), which sets the number of candidates (usually the best candidates given the fitness score mentioned below) that will remain in the next generation population.

[0217] As shown in Figure 4, the algorithm involves five steps that are iteratively repeated on the current population (initially the initial population) to produce the next population.

[0218] First, for each candidate c in the current population i Based on the mix production process model and the target mix profile, the fitness score SCO(ci ) is evaluated (step 400).

[0219] To do this, candidate c i A sample profile 110 defined by the mixing ratio {α k}, {β k}, and the predicted mix profile PRED(c i ) is obtained.

[0220] For illustrative purposes, in the second mixed production process (FIGS. 2B and 3B), the variable Γ={γ k Starting from a set P of profiles corresponding to the selected samples of}, calculate: First mixture: Mixing ratio M = {α k}, then I=PA, followed by R'=IW, Co-culture: R2 = Q1|Q2.Y (where Q1 is j repetitions of R' and Q2 characterizes j bioreactors), Final mixture: Mixing ratio B={β k}, then I=R2.A, followed by IW, which gives the candidate c i Predicted mix profile PRED(c i ) is obtained, and is executed.

[0221] Relevance score SCO(c i ) is PRED(c i ) and the target profile TARG. The penalty value described above may be added to the score initially evaluated by the evaluation function.

[0222] In some embodiments, the fitness score is calculated using PRED(c i In other embodiments, if a TARG is defined with reference to only some profiling features (e.g., some taxa), the entire PRED(c i ) is used.

[0223] For example, PRED(c i ) and TARG, we can use MSE:

number

[0224] In some embodiments, TARG is PRED(c i ) may have profiling features defined at a higher taxonomic rank (e.g., genus or phylum taxonomic rank) than the taxonomic rank used in PRED(c i ) is post-processed to convert lower-ranked features of the same higher taxonomic rank into a single value corresponding to the higher rank. For example, the relative abundances of taxa associated with higher taxonomic ranks are summed. This may be done using the correspondence matrix mentioned above. The above MSE is then used to find the candidate c i can be used to obtain a relevance score for

[0225] If TARG defines a specific target abundance for a profiling feature, then MSE can be easily calculated. However, in some embodiments, the target for one or more profiling features may differ from an exact value, for example, it may be an inequality (< or >) or a range of values ​​(which can be assimilated into a double inequality).

[0226] In that case, the inequality(s) defined by TARG are first evaluated by PRED(c i) for each profiling feature. If a profiling feature does not satisfy the respective inequality constraint, the fitness score is calculated including the distance between the profiling feature and the inequality threshold. On the other hand, if a profiling feature satisfies the respective inequality constraint, the corresponding distance in the fitness score is set to 0. This can be achieved by setting TARG with its exact value and the inequality threshold for profiling features with inequalities, and then setting PRED(c) to the corresponding inequality threshold for each profiling feature that satisfies the inequality constraint. i ) to enforce zero distance for this feature. i ) and TARG to obtain a goodness-of-fit score.

[0227] By way of example, the inequality can reflect a target diversity criterion, such as a bacterial diversity criterion.

[0228] "Diversity" or "bacterial diversity" refers to the variety or variability of a complex microbial community (mix or sample), measured, for example, at the genus, species, gene, function, RNA, or metabolite level. Diversity can be expressed in terms of richness (the number of observed species or genera or genes), alpha diversity parameters for describing complex communities, such as the Shannon index, Simpson index, and inverse Simpson index, and beta diversity parameters for comparing complex communities, such as the Bray-Curtis index, UniFrac index, and Jaccard index.

[0229] Thus, the inequality can represent a minimum or maximum relative abundance for one or more specific profiling features. For example, it may be desirable for a given bacterial genus to be present in the resulting mix product at a ratio (by mass) that is at least 5% of other bacteria (as specified by other profiling features). The diversity criterion can also define a range within which the relative abundance of one or more specific profiling features should fall. Of course, various diversity criteria can be combined: a minimum or maximum relative abundance for one profiling feature with a range for another feature and / or a maximum relative abundance for a third feature, etc.

[0230] All candidates c in the current population i About the score SCO(c i ) is known, step 410 consists in the module 140 selecting a sub-portion of the current population based on the evaluated scores.

[0231] In some embodiments, the number of candidates selected is predefined. For example, the number of candidates selected corresponds to the parent ratio applied to the population. Applying a parent ratio of 0.3 to a population of 200 candidates will result in an average of 60 candidates being selected.

[0232] In some embodiments, the best-fit candidate (here the one with the lowest SCO (c i ) are selected.

[0233] In some embodiments, candidates are selected based on probability, with the probability of each candidate being based on its fitness score. Indeed, the most fit candidates should be selected with a higher probability than candidates with lower fitness scores.

[0234] Following selection, a new candidate population is generated.

[0235] In some embodiments, an elite group may be selected from the population based on the elite ratio and may include a corresponding number of the fittest candidates. For example, if the elite ratio is 0.01, the two fittest candidates from the population of 200 candidates will constitute the elite group and therefore will necessarily be selected in step 410. The elite group initially constitutes the new generation.

[0236] Other candidates that make up the new generation are obtained based on the selected candidates using genetic crossover between genes of the selected candidates and / or genetic mutation within gene sequences, steps 420, 430, and 440. The crossover probability defines how many candidates for the new (or next generation) population will be constructed through crossover between pairs of candidates, while the other part of the new population will be made of the selected candidates unchanged.

[0237] Step 420 consists in pairing the selected candidates and then, possibly, exchanging some of their "chromosomes", i.e. genes, in order to breed new candidates for the next population.

[0238] For example, new pairings are performed until the new population is completely constructed (i.e., until it has size SIZ), which means that pairs of selected candidates are obtained.

[0239] In some embodiments, each pairing is made randomly among all selected candidates, while in other embodiments, each selected candidate in a pair is assigned a decreasing probability of being randomly selected in a new pairing.

[0240] The candidate pairs may then exchange parts of their chromosomes based on the crossover probability. If crossover occurs for this pair, they will exchange their genes (γ i , α i , β i) to create one or two new candidates for the new population. This is a crossover operation on the candidates. If no crossover is performed on this pair, one or both of the selected candidates are kept as new candidates in the new population.

[0241] In some embodiments, genes are randomly selected from the gene sequence and swapped between the two candidates in a pair, with each pair then breeding to two new candidates. Alternatively, new candidates are constructed by randomly selecting each gene from either candidate in a pair with equal probability, with each pair then breeding to only one new candidate.

[0242] This results in a new population of SIZ candidates that form the new population.

[0243] In step 430, these SIZ candidates are subject to genetic mutations within their genetic sequences (chromosomes). i , α i , β i The decision whether to perform a bit flip in γ (a new random value for the gene) is based on the mutation probability. If a candidate is determined to undergo a gene mutation, the gene involved in the mutation can be randomly selected. i can be changed to a new integer randomly chosen from [1...N]. i , β i can be changed to a new float randomly chosen from [0...1].

[0244] Mutations help maintain diversity within a population and prevent premature convergence.

[0245] In step 440, a new population is constructed from the SIZ candidates (mutated or not) obtained from step 430. The new population can be used for the next iteration of the algorithm by looping back to step 400, or it can be used as the final population in terms of selecting candidate(s) as solutions for the decision process given the target mix profile (step 460).

[0246] Thus, step 450 determines whether the algorithm termination criteria are met, in order to terminate the decision process when the population has converged.

[0247] In some embodiments, termination occurs after a predetermined number of iterations, for example 1000.

[0248] In other embodiments, the termination is i or a set of candidates, or a proportion of candidates in a population, that are sufficiently well-matched. "Sufficiently well-matched" means having a fitness score that exceeds a predetermined fitness score threshold (e.g., is below the threshold if the fitness score is the distance MSE defined above).

[0249] Generally, the algorithm terminates either when the maximum number of generations has been generated or when a sufficient level of fitness has been reached in the population.

[0250] Once the iterations are complete, candidates for the population resulting from the evolutionary algorithm are selected (step 460), for example, based on the target mix profile. In an embodiment, the best candidates, i.e., those with lower SCO (c i ) is selected as the candidate solution for the decision process. This candidate is then the sample {γ i}, and the blend ratio {α i}, {β i} to produce a mix result product whose profile is close to the target mix profile.

[0251] Alternatively, one of the best-fit candidates (eg, one of the best 10) may be randomly selected.

[0252] In some embodiments, multiple candidates, e.g., two or three, are selected as the best candidates (the fittest ones, or the ones selected from the group of fittest ones). This applies particularly when attempting to generate a mix / CCMC result product from a set of multiple initial samples. Preferably, the multiple candidates do not share the same initial sample (i.e., the same index γ i ) in order to produce the resulting mix product without consuming the same material (initial sample). In this way, a larger amount of product can be produced with a mix profile closer to the target.

[0253] The candidate solutions c(i) obtained by module 140 can thus be used by module 141 to control the process of actually producing the CCMC result product 19. For example, the candidate solutions c(i) can be used to control, via signaling S1 and optionally signaling S2, a product generator 15 implementing a mix production pipeline, e.g., as in Figures 2A, 2B, or 2C. This means that a method for producing a complex microbial community (CMC) product can first include obtaining candidates representing the set and mixture ratios of CMC samples in an initial sample collection using the above-described decision method based on a target mix profile.

[0254] For ease of explanation, the following description focuses on a single candidate used to produce a target mix result product. Similar processes can be performed sequentially or in parallel for multiple candidate solutions.

[0255] Once the candidates are identified, the process of producing the mix result product 19 begins.

[0256] In an embodiment, module 141 uses S1 to provide product generator 15 with various samples {γ i} and mixing ratio {α i}, {β i}. In a variant, the signal S1 may be an indication to the operator. For example, i} and mixing ratio {α i}, {β i} is displayed on the screen to the operator, allowing the operator to manually carry out the actual sample extraction and mix production process.

[0257] For ease of explanation, the following description focuses on the second mixed production process, which includes the pooling and co-cultivation steps. Those skilled in the art will directly adapt this description to the first and third mixed production processes introduced above.

[0258] The product generator 15 is a machine with mechanical access (e.g., via a controlled articulated arm) to an initial sample collection 10 of samples, and may include a pooling bioreactor for performing mixing of complex microbial communities, and j bioreactors for performing co-cultures according to different growth conditions.

[0259] In response to signal S1, product generator 15 selects samples {γ i}, i.e., search or obtain, and extract the respective mixture ratios {α i} and the total volume or mass targeted for the resulting mix product 19, and take an amount of each sample given the known amplification factor of each co-cultivation step.

[0260] If the amount of one of the samples is insufficient, product generator 15 may send an error message indicating the missing sample back to module 141. In this case, module 141 may restart the mix production process using another candidate (e.g., the next best candidate given the fitness score) or may restart the decision process of FIG. 4 to obtain such another candidate before triggering the mix production process again.

[0261] All sample volumes obtained are poured into a pooling bioreactor where they are actually mixed.

[0262] Preferably, the sample is homogenized for 10 minutes to 3 hours, preferably 30 minutes to 1.5 hours. Homogenization is carried out at a temperature of 0° C. to 10° C., preferably 2° C. to 8° C., more preferably about 4° C. The resulting CMC product is considered stable for several hours thereafter, and is considered stable for at least 16 hours, preferably 24 hours, from mixing.

[0263] The CMC product can be frozen before the co-cultivation takes place.

[0264] The product generator 15 then feeds the CMC product to each of the j bioreactors in equal proportions.

[0265] The CCMC product is obtained by co-cultivation using at least two bioreactors, preferably three bioreactors, with the bioreactors having at least one different parameter, the parameter being selected from pH, temperature, pressure, cultivation time, retention time, gassing conditions, redox potential, culture medium, light source, and combinations thereof.

[0266] Then, j CCMC products with different profiles are obtained.

[0267] The product generator 15 then calculates the respective mixing ratios {β i} and obtain an amount of each CCMC product given the total volume or mass targeted for the mix result product 19. The obtained amount of CCMC product is poured into a pooling bioreactor (optionally together with an amount of new sample from the initial collection) where it is actually mixed, for example using the parameters described above. A mix / CCMC result product is then obtained that has a mix composition that approximates the target mix profile.

[0268] As mentioned above, some samples 100y to 100z may be hypothetical. i} is a virtual sample, then that sample must be actually produced from its virtual definition (i.e., the corresponding individual profile).

[0269] When module 141 detects such a virtual sample 100y-100z corresponding to a bacterial composition, it signals, using S2, the need for the production of said artificial sample to sample generator 16. S2 can identify the sample involved and indicate the amount of material required.

[0270] The sample generator 16 may be a machine that has mechanical access (e.g., via a controlled articulated arm) to a bank 160 of isolated strains and has storage access to a file 161 that defines the composition of the sample in terms of a mix of individual strains. The sample generator 16 also includes a bioreactor in which the mixing of strains is carried out.

[0271] In response to signal S2, the sample generator 16 retrieves the definition of the artificial sample (bacterial consortium) in terms of strains and, given the signaled amount of required material, retrieves the appropriate amount of each required strain from the strain bank 16. The retrieved amount of all required strains is poured into a bioreactor where it is actually mixed, for example, for 30 minutes at 4°C.

[0272] In an embodiment, the sample generator 16 may have access to the bank 10 and / or even to an external sample bank 99. When the module 141 detects a virtual sample corresponding to a manipulated or processed complex community (i.e., a mix containing the sample), it uses S2 to signal to the sample generator 16 the need for production of that manipulated or processed sample. S2 can identify each strain and / or each sample in the bank 10 and / or each external sample related to the mix and indicate the amount of material required.

[0273] In response to signal S2, the sample generator 16 retrieves or extracts the material and injects the material into the bioreactor where the material is actually mixed.

[0274] Once the mix has been performed and stabilized, the sample has been generated and is stored in an initial sample collection or bank 10 from which the product generator 15 can retrieve the sample to actually produce the mix result product 19.

[0275] Although signals S1 and S2 are described above as control signals for driving product generator 15 and sample generator 16, one or both of them may simply be signals displayed to the operator for the operator to actually manually perform the mixing.

[0276] Samples may die over time (either for the purpose of actually producing some products or because they deteriorate over time), while new samples may be taken from new donors. Consequently, after a mix definition is determined to produce a target mix result product, the collection 10 may evolve over time (hence the matrix A evolves). In accordance with the present invention, a pooling predictor 130 may be newly constructed using the evolved collection (A is redefined and W is learned) from which new candidate solutions can be determined given the target mix profile (FIG. 4).

[0277] 5 shows a schematic representation of a computing device 500 managing the production platform 1. The computing device 500 may, for example, implement a predictor module 13 and a test and decision module 14, and may control the sequencer 12, the product generator 15, and the sample generator 16 via adapted signaling (S1 and S2).

[0278] The computing device 500 is configured to implement at least one embodiment of the present invention. The computing device 500 may preferably be a device such as a microcomputer, a workstation, or a lightweight portable device. The computing device 500 includes a communication bus 501, which preferably includes: a central processing unit 502, such as a microprocessor, denoted CPU; a read-only memory 503, denoted ROM, for storing a computer program for implementing the present invention; a random access memory 504, denoted RAM, for storing executable code of the methods according to embodiments of the present invention, as well as registers configured to record variables and parameters necessary to implement the methods according to embodiments of the present invention; a communications interface 505 connected to the network 599 for communicating with a user or operator device and / or with other devices of the platform 1 (e.g., the sequencer 12, the product generator 15, and the sample generator 16); and Connected is a data storage means 506, such as a hard disk or flash memory, for storing computer programs for implementing the methods according to one or more embodiments of the present invention, as well as any data required for the embodiments of the present invention, including in particular the individual sample profiles (i.e. collections 11).

[0279] Optionally, the computing device 500 may include a screen 507 that serves as a graphical interface with an operator, for example, to configure the platform (e.g., to set a target mix profile) via a keyboard 508 or any other pointing means, and / or to display information (e.g., potential solutions).

[0280] The computing device 500 may optionally be connected to various peripherals not necessary to the present invention, as well as to a sequencer 12, each connected to an input / output card (not shown).

[0281] Preferably, a communication bus provides communication and interoperability between various elements included in or connected to computing device 500. The representation of a bus is not limiting, and in particular a central processing unit may communicate instructions to any element of computing device 500 directly or via another element of computing device 500.

[0282] The executable code may be stored, as desired, either in read-only memory 503, on hard disk 506 or on a removable digital medium (not shown). According to an optional variant, the executable code of the program may be received by communication network 599 via interface 505, for storage in any of the storage means of computer device 500, such as hard disk 506, before being executed.

[0283] The central processing unit 502 is preferably configured to control and direct the execution of instructions of the program(s) or parts of the software code according to the invention, these instructions being stored on any of the storage means mentioned above. On start-up, the program(s) stored in a non-volatile memory, e.g. the hard disk 506 or the read only memory 503, are transferred to the random access memory 504, which then contains registers for storing the executable code of the program(s) as well as variables and parameters necessary to implement the invention. [Example]

[0284] Experimental results Scope of the experiment The aim of the experiment was to investigate the efficiency of evolutionary algorithms in the task of determining initial samples and mixing ratios to obtain a mix result product with a target mix profile (Experiment 4).Preliminarily, experiments investigated the validity of the learned interaction matrix-based approach to model and therefore predict the mix profile of a mix of microbiota samples (Experiments 1 and 2) as well as the validity of the co-culture product profile of a product obtained from co-culture in a bioreactor with specific growth conditions (Experiment 3).

[0285] Experiment 1 - Protocol We considered an initial sample collection 10. By sequencing each microbiota sample using 16S-based microbiota taxon profiling, we obtained a corresponding initial profile collection 11. Therefore, 131 taxa (at the genus level) were evaluated as profiling features.

[0286] Next, the samples were mixed. Each mix product was a combination of three to six samples with different ratios. Mixing was performed at 4°C and homogenized for 30 minutes to 1 hour and 30 minutes after mixing. The mix products were sequenced using the same 16S-based microbiota taxon profiling method during their stable state (i.e., within several hours of homogenization, within 16 hours of mixing).

[0287] To construct the pooled predictor 13, i.e., to learn λ and the interaction matrix W, we adopted a k-fold cross-validation strategy with k = 5. The k-fold strategy ensured that none of the observations was used as both training data and test set during the same evaluation.

[0288] The modeling method was tested and applied at four different taxonomic levels: species, genus, family, and order. However, the species-level dataset was very sparse and therefore excluded from the testing procedure. Starting from the genus level, the assignment table was rich enough to allow for analysis, so lower resolution levels (family, order) were inferred from the genus table for visualization purposes only, if necessary, but were not used in the modeling procedure. The main reason is that it is not possible to infer the composition of higher resolution levels from the taxon level used in training, and from an application perspective, it is important to have genus information.

[0289] Models were trained on untreated samples (Figure 5) and co-cultured samples (Figure 6) separately, as well as both combined (Figure 7). MSE was used to quantify the quality of modeling when applied to the data. MSE was systematically compared between machine learning models and linear models (which provide simple predictions).

[0290] Experiment 1 - Results Figure 5A shows the initial profile collection 11 corresponding to the initial sample collection 10, which contained only raw fecal microbiota samples. Twenty-seven microbiota samples were examined, and the individual profiles of these samples are shown.

[0291] Figure 5B shows the mix profiles of 24 mix products, each of which was made by mixing 3 to 6 of the 27 microbiota samples in Figure 5A in their respective ratios or proportions. Mix definition {p x (k)} k is saved.

[0292] Figure 5C shows the mix definition {p x (k)} k and individual sample profiles {a x (j)} j , which corresponds to the step I=A*P.

[0293] The figure also shows, on the right, the error resulting from the POOL prediction, i.e., the error including the product interaction matrix W. W was machine-learned using only the samples and mix profiles (raw samples) from Figures 5A and 5B using a k-fold cross-validation strategy.

[0294] Model-based methods return better performance than linear methods on raw data sets.

[0295] Figure 6A shows the initial profile collection 11 corresponding to the initial sample collection 10, which included only co-cultured fecal microbiota samples. Thirty-six microbiota samples were examined, and the individual profiles of these samples are shown.

[0296] Figure 6B shows the mix profiles of 48 mix products, each of which was made by mixing three to six of the 36 microbiota samples in Figure 6A in their respective ratios or proportions. Mix definition {p x (k)} k is saved.

[0297] Figure 6C shows the mix definition {p x (k)} k and individual sample profiles {ax (j)} j , which corresponds to the step I=A*P.

[0298] This figure also shows, on the right, the error resulting from the POOL prediction, i.e., the error including the mix interaction matrix W. W was machine-learned using only the samples and mix profiles (co-culture samples) from Figures 6A and 6B using a k-fold cross-validation strategy.

[0299] The model-based method returns dramatically better performance than the linear method on the coculture dataset (median MSE is 5x lower with the ML model predictions).

[0300] For Figures 7A and 7B, the interaction matrix W was machine-learned using both the sample and mix profiles (i.e., untreated and co-cultured samples) from Figures 5A, 5B, 6A, and 6B as training data. Again, a k-fold cross-validation strategy was used.

[0301] FIG. 7A shows the results when the datasets of FIGS. 5A and 5B (i.e., the raw samples and their mix) are applied to the pooled predictor 13 configured in this way.

[0302] The left side of the figure shows the mix definition {p x (k)} k and individual sample profiles {a x (j)} j 1 shows the error resulting from linear prediction of the mix profile given

[0303] The right side illustrates the error resulting from the POOL prediction, i.e., the error including the interaction matrix W so learned.

[0304] Similar to the single dataset model, the combined dataset model slightly improves estimation when applied to the raw dataset.

[0305] FIG. 7B shows the results when the datasets of FIGS. 6A and 6B (i.e., the co-culture samples and their mixes) are applied to the pooled predictor 13 configured in this way.

[0306] The left side of the figure shows the mix definition {p x (k)} k and individual sample profiles {a x (j)} j 1 shows the error resulting from linear prediction of the mix profile given

[0307] The right side illustrates the error resulting from the POOL prediction, i.e., the error including the interaction matrix W so learned.

[0308] Similar to the single dataset model, the combined dataset model dramatically improves estimation when applied to the co-culture dataset (median MSE is 4x lower for ML model predictions).

[0309] Experiment 1 - Observations and Conclusions In all cases, model-based predictions improve the estimates of naive (linear) methods. This improvement is important, especially in the coculture dataset, because naive approaches perform poorly, especially for some taxon groups. The model-based correction approach was more efficient, likely because it had more room for improvement. If more data are added to train the model, we can assume that the overall performance and robustness will improve. The training method allows for such model evolution.

[0310] Experiment 2 - Protocol In this experiment, we used NGS shotgun sequencing to profile 100 samples. Metagenomic sequence data were obtained from 76 pools and 69 samples (donor-derived or individual co-cultures).

[0311] Due to the large number of NGS shotgun profiling features (compared to 16S sequencing, especially when looking at the species level instead of the genus level, or when looking at specific functions), PCA was used to reduce the dimensionality of each sample profile to k PCs.

[0312] Figure 8 illustrates PCA based on genus relative abundance obtained from NGS shotgun sequencing of untreated samples (untreated:sample or mix) and co-cultured samples (co-cultured:sample or mix). The co-cultured samples tended to cluster together, as did the untreated samples.

[0313] This PCA-based strategy is summarized in Figure 9, where it is clear that instead of learning a “taxon × taxon” interaction matrix W (as in Experiment 1), in Experiment 2 a “top k principal components × taxon” interaction matrix W is learned.

[0314] The methodology for learning this interaction matrix W is the same as the 16S analysis in Experiment 1.

[0315] Experiment 2 - Results [Table 1]

[0316] The interaction matrix W was trained using MSE. In addition, the predicted mix results (using W) were compared with the true mix results based on MSE or Bray-Curtis distance.

[0317] Both modeling approaches (with or without PCA) improve taxonomic profile predictions (according to MSE or BC indices) at the genus and species levels. The correction based on matrix W has a stronger impact in predicting mix from co-culture samples compared to untreated samples.

[0318] Profiling feature reduction using PCA appears to significantly improve prediction accuracy for co-culture samples, but only marginally for predictions from untreated samples.

[0319] Experiment 3 - Protocol Three co-culture procedures were used for this experiment, resulting in three datasets, designated EXP1, EXP2, and EXP3, respectively.

[0320] For the three procedures, three co-culture bioreactors were used, which simultaneously provided different environmental conditions, specifically differing only in terms of pH: Fermenter or bioreactor 1: pH=5.3, Fermenter or bioreactor 2: pH=6.3, and Fermenter or bioreactor 3: pH=7.3.

[0321] A single microbiota (starting CMC product) was cultured in three bioreactors at a time as follows.

[0322] In EXP1, the starting CMC product in three bioreactors was grown for 14 days and sampled on day 4, and three CCMC products were obtained.

[0323] Two other starting CMC products from different donors were also grown in the same bioreactor for 14 days.

[0324] Thus, the nine CCMC products sampled on day 4 were obtained through EXP1.

[0325] In EXP2, the starting CCMC product was grown in three bioreactors in parallel for three days. The resulting CCMC products were each frozen, thawed, and fed to each bioreactor for an additional three-day growth. The resulting CCMC products (after freezing and thawing) were fed again to each bioreactor for a third three-day growth. Periodic sampling was performed.

[0326] Thus, nine different propagation processes were performed in EXP2, resulting in three final and six intermediate CCMC products (hence, nine CCMC products).

[0327] In EXP3, the starting CCMC product was grown in three bioreactors in parallel for three days. The resulting CCMC product was mixed and fed to three bioreactors for an additional three days of growth. The resulting CCMC product was mixed again and fed to three bioreactors for an additional three days of growth. Periodic sampling was performed.

[0328] Thus, in EXP3, nine different propagation processes were performed, resulting in three final and six intermediate CCMC products (hence, nine CCMC products).

[0329] In total, 27 CCMC products (after 3 or 4 days of growth) were generated through three experiences, each involving three co-culture bioreactors.

[0330] The profiles of the starting CMC product and of these 27 CCMC products were obtained by sequencing using 16S-based microbiota taxon profiling.

[0331] These profiles were used to train fermenter predictor models. Six models were used, the first of which, as described above with reference to Figure 3B, is R2 = Q*Y, where Q = Q1|Q2 defines the growth conditions (here, pH values) and Y is the co-culture interaction matrix to be trained.

[0332] Other models included: Defined as (Q1*Y1)+(Q2*Y2), where Y1 and Y2 need to be trained,Model 2; It is defined as (Q2*Y2)·Q1, where Y2 needs to be learned and · is the Hadamard product, i.e., the element-wise product. Model 3; It is defined as (Q1*Y1)+(J*Y'2), where Y1 and Y'2 need to be learned and J is an n×1 matrix filled with ones (n is the number of rows in Q1, i.e., the number of starting CMC products). Model 4; It is defined as ((Q2*Y2)·Q1)*W), where Y2 needs to be trained and W is the matrix of pooled predictors, Model 5; Model 6 is defined as ((Q2*Y2)·Q1)*Y1), where Y1 and Y2 need to be trained.

[0333] Each model was trained using three datasets as described above for the pooled predictor (Experiment 1 and Experiment 2). Specifically, a k-fold cross-validation strategy was applied. MSE was used to evaluate the distance between the predicted and true CCMC profiles. Comparisons were made using the calculated distance between the baseline predicted CCMC profile (naive prediction) and the true CCMC profile.

[0334] The initial naive prediction, NAIV1, was the leave-one-out average of the relative abundances across all experiments for each taxon in the output of each bioreactor. Indeed, the taxonomic profile of each sample measured after cultivation was compared to the average profile of all other samples cultivated in the same bioreactor. Figure 10A illustrates the observed relative abundances (x-axis) and the naive prediction according to this leave-one-out average for each sample and each genus (different colors). There are three lines per genus because three different bioreactors were tested. The lines exhibit a slight slope because the more important the true relative abundances, the more removed from the leave-one-out average. A simple average would have resulted in a horizontal line for each genus / bioreactor combination.

[0335] The second naive prediction used, NAIV2, was that the starting CMC product would retain the same composition as if the cultivation process had not altered the relative abundances or, more broadly, the taxonomic profile (hence, it is a 100% naive prediction). Figure 10B suggests that this assumption is not valid, as the points do not appear particularly close to the y = x line that would correspond to a perfect prediction.

[0336] Experiment 3 - Results [Table 2]

[0337] Experiment 3 - Discussion and Conclusions Table 2 confirms that better predictions are obtained than the baseline predictions NAIV1 and NAIV2. Table 2 further shows that for the dataset used, Model 1 provides the best predictions.

[0338] Notably, the dataset used to train and evaluate the model was small and low in diversity. Using a larger dataset with more diversity between starting CMCs is expected to increase the benefit of using Model 1 compared to the baseline predictions.

[0339] Experiment 4 - Protocol In this experiment, we investigated the efficiency of evolutionary algorithms in the task of determining initial samples and mixing ratios to obtain a mix result product with a target mix profile.

[0340] Various sub-experiments were performed that differed in how the target mix profile was defined. The initial sample collection used contained n = 63 samples profiled with 16S sequencing methods, and therefore, for all of them, a table of relative abundances at the genus level was available.

[0341] In EXP4.1, we randomly generated a collection of 100 genus-level target profiles with k=8 (the number of samples to mix). We repeated the following steps 100 times: 1:n alpha ratios were randomly generated with the following constraints:

number

number

[0342] 4: The CCMC profile was obtained using the Fermenter Predictor 131 and the final mix product profile was calculated by application of the Pool Predictor 130 given the calculated B matrix and CMC profile as input.

[0343] 5: The calculated final mix product profile was saved for use as the target product profile. Therefore, 100 diverse target product profiles consistent with realistic production process data were generated. Then, the above evolutionary algorithm was run using each of the 100 targets as input with the following parameters: Maximum iterations = 500 Population size = 200 Mutation probability = 0.3 Elite ratio = 0.01 Crossover probability = 0.5

[0344] The MSE and Bray-Curtis distance between each target query and each pair of mix result product predictions (as processed in silico according to the settings shown in Figure 2B) were then calculated and stored.

[0345] In EXP4.2, the target profile was not fully defined. Seven genera were used to define the target profile. Their sum of relative abundances was 0.27. The minimization function was applied only to those seven genera. Other genera were allowed in the final mix. Their numbers and relative abundances were not constrained.

[0346] In EXP4.3, the target profile was not fully defined. Seven genera were defined as target profiles, three of which had inequality constraints and four had equality constraints. As explained above, the minimization function (MSE) was applied differently for each taxon depending on the operator. For example, for equality, the distance between the relative abundance of the taxon in the target profile and the predicted profile was minimized. For inequality, the same distance continued to be minimized while the predicted relative abundance did not satisfy the inequality. Once this condition was met, the taxon was no longer constrained. Other genera were allowed in the final mix; their numbers and relative abundances were not constrained.

[0347] In EXP4.4, the target profile was not fully defined, and this target definition used features from a different taxonomic hierarchy. Therefore, the target profile was defined by adding one order constraint to a random full target profile (as in EXP4.1) (orders as a taxonomic hierarchy include the family level, which itself includes the genus level).

[0348] In EXP4.5, the target profile was not fully defined, but rather constructed using features from several classification hierarchies and mixed operators.

[0349] Experiment 4 - Results For EXP4.1, Figure 11A shows the distribution of MSE (left) and Bray-Curtis (BC) distance (right) between each pair of target and corresponding predicted profiles after the genetic algorithm process. Therefore, 100 points are plotted.

[0350] The MSE and BC distances are very low. In fact, the BC distance between two profiles obtained from two sequences of one real sample is rarely below 0.20. In EXP4.1, the maximum distance obtained here was below 0.09.

[0351] Figure 11B shows a comparison of one of the 100 target profiles (number 58) with the associated predicted profile. Each point corresponds to one genus, where the x value is the expected relative abundance defined in the target and the y value is the predicted relative abundance of the best candidate obtained through genetic algorithm prediction.

[0352] Test 58 was chosen as an example because the associated prediction profile has an MSE close to the average of 100 experiments.

[0353] In the figure, the points are all distributed very close to the y=x axis, thereby indicating a high similarity between the target and predicted profiles for the average outcome.

[0354] For EXP4.2, Table 3 reports the target relative abundance and predicted relative abundance for each of the seven selected genera in the target profile.

[0355] [Table 3]

[0356] The best candidate obtained with EXP4.2 had an MSE of 5.84E-06 and a BC distance of 1.91E-02, demonstrating that it is possible to find promising candidates even with a less than perfect target profile.

[0357] For EXP4.3, Table 4 reports the target relative abundance and predicted relative abundance for each of the seven selected genera in the target profile.

[0358] [Table 4]

[0359] The best candidate obtained with EXP4.3 had an MSE of 5.15E-08 and a BC distance of 6.02E-02, demonstrating that it is possible to find promising candidates even with the inequality operator and an incomplete target profile.

[0360] For EXP4.4, Table 5 reports the target and predicted relative abundances for two taxa defined at several taxonomic levels among the multiple taxa measured.

[0361] [Table 5]

[0362] The best candidate obtained with EXP4.4 had an MSE of 5.32E-06 and a BC distance of 5.95E-02, demonstrating that it is possible to find promising candidates even with several taxon levels and a less-than-complete target profile.

[0363] For EXP4.5, Table 6 reports the target relative abundance and predicted relative abundance for each of the five taxa defined in several taxonomic hierarchies.

[0364] [Table 6]

[0365] The best candidate obtained with EXP4.5 had a high MSE of 6.8E-06 and a high BC distance of 32.9E-02. This was mainly due to the fact that the relative abundance value of the predicted profile was 0.21 compared to the target profile's relative abundance value of 0.50, even though the inequality constraint was satisfied (0.29 ≤ 0.50).

[0366] This successful example thus demonstrates that it is possible to find promising candidates even at several taxon levels and with less than complete target profiles.

[0367] Experiment 4 - Discussion and Conclusions Experiment 4 aimed to evaluate the validity of the genetic algorithm for solving the candidate selection problem using the implemented parameters. This evaluation was carried out by generating 100 realistic profiles at the genus level using the pool predictor and the fermenter predictor. The results of the main task (EXP4.1) were satisfactory, with all predictions showing a high similarity to their respective targets.

[0368] In practical situations, it makes sense not to define the goal comprehensively. There may not be enough information to enumerate all necessary features, to query them using inequalities, and to query using different feature levels if they are organized hierarchically (e.g., microbial taxa, but also functional features such as EC numbers). Experiments EXP4.2–EXP4.5 aimed to explore each possibility and their combinations, showing that genetic algorithms can function efficiently in these settings.

[0369] Experiment 4, which was based solely on in silico evaluation, was successful in showing that the genetic algorithm can efficiently find candidates that solve the task of determining the initial samples and mixing ratios to obtain a mix result product with a target mix profile.

[0370] Although the present invention has been described with reference to particular embodiments, it is not limited to those embodiments, and modifications within the scope of the invention will be apparent to those skilled in the art.

[0371] Many further modifications and variations will be suggested to those skilled in the art by reference to the preferred embodiments described above, but these embodiments are given by way of example only and are not intended to limit the scope of the invention, which is determined solely by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.

[0372] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

Claims

1. The set of complex microbial community (CMC) samples (100) and their mixing ratios (α i , β i ) and producing a mix result product from the CMC sample using a mix production process configured with the mix ratio, An initial candidate population (c), each candidate representing a set (Γ) and mixture ratio (M, B) of CMC samples in the initial sample collection. i ) to obtain applying (400-450) an evolutionary algorithm to iteratively modify the candidate population based on a model of the mix production process and a target mix profile (TARG) representing a target mix result product; and selecting (460) candidates for the modified population resulting from the evolutionary algorithm; A method comprising:

2. The mixed production process model is as follows: Predicting the intermediate CMC profile for a mix of selected complex microbial community samples given the mixing ratio using a linear approach; and correcting the intermediate CMC profile to a predicted CMC profile representing a complex microbial community (CMC) product obtained from the mix using a CMC interaction model trained from a reference linear predicted CMC profile and a corresponding reference true CMC profile; The method of claim 1 , comprising:

3. The method of claim 2 , wherein each candidate includes a mixing ratio representing a respective ratio of the CMC samples to be mixed.

4. The mixed production process model may include one or more of the following loops: From the predicted CMC profile or from the predicted mix result profile of the preceding loop, predicting a plurality of CCMC profiles representing co-cultured CMC products obtained by separate co-cultures of the same starting CMC product; and Using a linear approach and from the CCMC profile, predicting an intermediate mix result profile representing a second mix of the CCMC product given the blend ratio; and correcting the intermediate mix result profile to a predicted mix result profile representing a blended CCMC product obtained from the second mix using a mix interaction model learned from a reference linear predicted mix profile and a corresponding reference true mix profile; The method of claim 2 or 3, further comprising:

5. 5. The method of claim 4, wherein each candidate further comprises a blending ratio representing the respective ratios of CCMC products to be blended.

6. The method according to claim 4 or 5, wherein the same mixing ratio representing the respective ratios of the CCMC products to be mixed is used throughout the loop.

7. The method according to any one of claims 4 to 6, wherein the CMC interaction model and the mix interaction model are one and the same model.

8. 8. The method of claim 1, wherein each candidate is defined by a sequence of genes including a sample identifier and a mixture ratio, and each sample identifier and mixture ratio defines a separate gene.

9. The iterations within the evolutionary algorithm include: assessing a score for each candidate in the current population based on the mix production process model and the target mix profile; Selecting a portion of the current population based on the assessed score; and generating a new candidate population based on the selected candidates using genetic crossover between genes and / or genetic mutation within the gene array of the selected candidates; The method of claim 8, comprising:

10. The method of claim 9 , wherein evaluating the score comprises calculating a distance between a target mix profile and a mix result profile predicted from the candidate using the mix production process model.

11. 1. A method for producing a co-cultured complex microbial community (CCMC) result product, comprising: The method of any one of claims 1 to 10 based on a target mix profile is used to determine the set (100) of complex microbial community (CMC) samples and the mixing ratios (α i , β i ) i ) to obtain actually extracting the set of CMC samples from the initial sample collection; and Processing the extracted CMC sample using a mix production process configured with the obtained mixing ratio to obtain a CMC result product; A method comprising:

12. The method of claim 11, wherein the mix production process includes a first pooling step in which the extracted CMC samples are mixed to obtain a CMC product.

13. The method of claim 12 , wherein the mixing is performed according to the obtained candidate mixing ratios.

14. The method of claim 12 or 13, wherein the mix production process further comprises one or more second iterations of expanding the starting CMC product, the iterations comprising (i) co-cultivating the CMC product or a resulting mix product obtained from a previous iteration in a bioreactor having respective operating parameters to obtain a CCMC product, and (ii) mixing the CCMC products to obtain a resulting mix product.

15. A computing device comprising at least one microprocessor configured to carry out the method of any one of claims 1 to 14.

16. A non-transitory computer readable medium storing a program which, when executed by a microprocessor or computer system in a device, causes said device to carry out the method of any one of claims 1 to 14.