Methods of determining microbiota samples and predicting mixtures for production of target mixed products

Through the combination of computer-aided design and mixture interaction models, the prediction and production of complex microbial community mixture compositions are solved, and the problems of inefficiency and rare materials are achieved in the prior art, and the rapid and accurate mixture composition simulation is achieved.

CN120226086APending Publication Date: 2025-06-27MAAT PHARMA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380075799.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-08
Filing Date
2023-11-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency, rare materials and long analysis time when predicting and producing the composition of complex microbial community mixtures, especially in the process of amplifying and maintaining microbial viability.

Method used

Using a computer-aided design method, the intermediate mixture feature spectrum of co-cultivated complex microbial communities is predicted by linear method, and the learned mixture interaction model is used to correct it to predict the mixture feature spectrum to achieve nonlinear behavior modeling of microbial groups mixing.

Benefits of technology

The composition of complex microbial community mixtures is achieved at a low cost, instant and accurate manner, avoiding actual material consumption and significantly shortening the testing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226086A_ABST
    Figure CN120226086A_ABST
Patent Text Reader

Abstract

And under the condition that the target characteristic spectrum of the CMC product is given, the production process parameters of the CMC product are determined by utilizing an evolutionary algorithm. And modeling CMC mixing operation and co-culture operation of the CMC product by using a matrix-based learning model. The evolutionary algorithm iteratively modifies candidates representing parameters including a set of complex microflora samples in an initial sample collection library and a mixing ratio for one or more mixing operations in a production process. And then, according to a mixture production process, the group of determined samples and the related mixing ratio are utilized to control the actual picking and processing processes of the complex microbial community samples, and a CMC product which is close to a target characteristic spectrum as far as possible in the aspect of characteristic analysis is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the modeling of production processes or pipelines of microbiota mixtures, and more particularly, to methods for predicting mixture compositions and determining complex communities of microorganisms or microbiota to be mixed for producing a target mixture composition. Background Art

[0002] Complex microbial communities (also known as microbiota) play a key role in health and disease treatment. Specifically, it has been found that various infection and disease problems can be addressed by administering or transplanting complex microbial communities (e.g., via fecal microbiota transplantation (FMT)).

[0003] For the administration or transplantation of complex microbial communities, the samples to be administered or transplanted need to have an appropriate characteristic profile in terms of the viability, functionality, and diversity of microorganisms and their components (such as bacteria, archaea, viruses, bacteriophages, protozoa, metabolites, yeasts, RNAs, and / or fungi).

[0004] Some administration and transplantation methods (such as FMT) are usually empirical methods, and these methods do not take specific measures to ensure the diversity of the microorganisms present in the samples used or to maximize the viability of the microorganisms.

[0005] In addition, not all samples collected from donors may have a satisfactory characteristic profile of complex microbial communities and their derivatives (metabolites, RNAs, etc.), thus preventing effective treatment.

[0006] Therefore, the diversity of samples can be increased by collecting mixtures of complex microbial community samples from several donors, and these samples can be used as inocula for administration or transplantation, such as the native microbiome ecosystem therapy product (MET-N) produced by MaaT Pharma (registered trademark).

[0007] In addition, as described in WO2022 / 136694, amplifying complex microbial communities by microbial co-culture in separate bioreactors helps to amplify the microbial communities while also maintaining a high diversity. Exemplary amplification includes: (a) culturing the complex microbial community in at least two bioreactors, preferably in at least three bioreactors, to obtain at least two co-culture samples, wherein the bioreactors have at least one different parameter selected from the group consisting of pH value, temperature, pressure, culture time, residence time, aeration conditions, redox potential, culture medium, light source, and combinations thereof; (b) mixing the at least two co-culture samples obtained in step (a) to obtain an amplified complex microbial community. Once formulated and ready for administration, such products correspond to co-cultured microbiome ecosystem therapy products (MET-C) (such as products developed by MaaT Pharma).

[0008] To test various different mixtures or amplified complex communities, samples are currently randomly mixed or amplified, and then the resulting products are sequenced to obtain the final mixture or amplified characteristic spectrum, from which preventive and therapeutic characteristics are inferred. However, this test-based method has some drawbacks. Specifically, if it is difficult to obtain samples from donors, this method requires the consumption of rare materials, and due to the long analysis time, it takes several weeks to complete. In addition, especially for the amplification of complex communities, the amplification also requires a long co-culture time (e.g., several hours or days), thus prolonging the test duration.

[0009] Therefore, it has been considered to predict the mixture composition (i.e., the characteristic spectrum of the mixed product) by modeling the microbiota mixture production process or pipeline (usually as a linear prediction model). However, the known models usually do not meet the requirements, which requires an improved solution for modeling the mixture production process.

[0010] A reversible prediction model can be conveniently used to determine the initial samples to be processed in the mixture production process, so as to obtain the desired final mixture composition (e.g., the mixture composition with the expected therapeutic effect). However, prediction models are usually irreversible, so a new initial sample determination mechanism needs to be proposed. Summary of the Invention

[0011] The present invention aims to solve some of the above problems.

[0012] One aspect of the present invention focuses on discussing a prediction optimization scheme for mixing co-culture samples obtained after growth in a bioreactor (and thus focuses on modeling). When predicting the mixture composition, the prediction optimization depends on whether a computer-aided design is performed on the conversion between such mixing and the conventional linear prediction characteristic spectrum of the true characteristic spectrum (obtained by analyzing the amplified complex microbial community).

[0013] For this purpose, in this aspect, the present invention proposes a computer-aided method, that is, in respective bioreactors, co-culture complex microbial community (CCMC) products are obtained by co-culturing one or more starting complex microbial community (CMC) products (possibly different products) (preferably by performing different co-cultures on the same starting CMC product), these products are mixed to generate a mixture composition, and the mixture composition is predicted. The method includes:

[0014] (a) Predicting the intermediate mixture characteristic spectrum of the CCMC product mixture by a linear method;

[0015] (b) The intermediate mixture characteristic spectrum is corrected to a predicted mixture characteristic spectrum by using a mixture interaction model learned from the linearly predicted reference mixture characteristic spectrum and the corresponding reference true mixture characteristic spectrum.

[0016] Thus, in this regard, the present invention models the non-linear behavior of microbiota mixing. The advantage of this modeling method is that it can simulate various mixture compositions of CCMC products at low cost, instantaneously, and accurately, especially without consuming any actual materials.

[0017] Therefore, according to the intended use (such as treatment, prevention, and environmental requirements, etc.), the co-culture strategy can be defined before the production process is executed.

[0018] Modeling can be based on matrices. For example, predicting the intermediate mixture characteristic spectrum may include: calculating the matrix product between a first matrix and a second matrix, where the first matrix defines the mixture in terms of the CCMC product ratio, and the second matrix defines the individual characteristic spectra of the CCMC products. Next, correcting the intermediate mixture characteristic spectrum may include: calculating the matrix product between the matrix representing the intermediate mixture characteristic spectrum and the mixture interaction square matrix of the learned mixture interaction model. Here, the mixture interaction model can be a mixture interaction square matrix learned from the linearly predicted reference mixture characteristic spectrum and the corresponding reference true mixture characteristic spectrum.

[0019] The advantage of using matrices for microbiota mixture prediction is that a large number of analytical features can be considered and fast operations can be achieved, so as to obtain one or more predicted mixture characteristic spectra of the mixed result product.

[0020] In some embodiments, different co-cultures are performed on the same starting CMC product in respective bioreactors to obtain CCMC products, and the mixture composition is generated by iterative mixing of the CCMC products, and the method includes multiple iterations of steps (a) and (b), wherein the predicted mixture characteristic spectrum obtained in the previous iteration step (b) is used to obtain the characteristic spectrum of each CCMC product in the next iteration step (a). This method models multiple co-culture and mixing cycles: the starting CCMC product for the next iteration is the mixed result of the previous iteration. The CCMC products produced from the same CCMC product will be mixed in the next iteration. Starting from the same starting CCMC product for chain co-culture, a large amount of materials (final mixture composition) can be obtained. Therefore, this configuration facilitates better prediction of chain co-culture.

[0021] In some embodiments, the same mixture interaction model is used iteratively. This means that the same mixture interaction model (e.g., the same matrix) is used in each step (b) throughout the iteration process. Since multiple iterations require single-model learning and storage of the single model, processing and memory costs are saved.

[0022] In some embodiments, a starting CMC product is obtained from a mixture of complex microbial community samples selected from an initial sample collection library, the method comprising:

[0023] Predicting the CMC characteristic spectrum of the starting CMC product based on the sample characteristic spectrum of the selected complex microbial community sample;

[0024] Predicting the CCMC characteristic spectrum of the CCMC product (for the first iteration of steps (a)-(b) when multiple iterations are considered) based on the predicted CMC characteristic spectrum;

[0025] wherein, in step (a), the intermediate mixture characteristic spectrum is predicted by a linear algorithm according to the CCMC characteristic spectrum.

[0026] The advantage of sample mixing is that it averages the variations between samples from different donors as well as the variations between samples from the same donor (variations between two donations made on different dates), while also obtaining a starting CMC product for co-culture that is richer (in terms of diversity) than just one sample.

[0027] In some embodiments, the predicted CMC characteristic spectrum includes:

[0028] Predicting the intermediate CMC characteristic spectrum required for the mixture of the selected complex microbial community sample using a linear method;

[0029] Using a CMC interaction model learned from the linearly predicted reference CMC characteristic spectrum and the corresponding true reference CMC characteristic spectrum to correct the intermediate CMC characteristic spectrum to the predicted CMC characteristic spectrum. The CMC interaction model models the non-linear behavior of sample mixing. The advantage of this modeling is that it can simulate various mixture compositions of the selected samples from the sample collection library at low cost, instantaneously, and accurately, especially without consuming any actual materials.

[0030] In some embodiments, the CMC interaction model and the mixture interaction model are the same model, e.g., the same matrix. Since the entire prediction process requires single-model learning and storage of the single model, processing and memory costs are saved.

[0031] A unique aspect of the present invention lies in highlighting a new mechanism for determining relevant initial samples to be processed, independent of the reversibility of the mixture production model.

[0032] To this end, in this regard, the present invention proposes a computer-aided method for determining a set of complex microbial community (CMC) samples and their mixing ratios in an initial sample collection library (10), and using a mixture production process configured with the mixing ratios to produce a mixed result product from the CMC samples. The method includes:

[0033] Obtaining an initial candidate population, where each candidate represents a set of CMC samples and their mixing ratios in the initial sample collection library;

[0034] Based on the model of the mixture production process and the target mixture characteristic spectrum representing the target mixed result product, applying an evolutionary algorithm to iteratively modify the candidate population;

[0035] Selecting a candidate from the modified population generated by the evolutionary algorithm.

[0036] The inventors of the present invention have found that evolutionary algorithms such as genetic algorithms can accurately determine the production process parameters (initial samples and mixing ratios). Therefore, when the production process model is irreversible, it is advisable to use such algorithms. Such algorithms can also be used together with non-convex fitness scoring metrics such as Bray-Curtis dissimilarity and can easily adapt (with only a relatively low degree of increase in complexity) to multiple iterations (i.e., production sub-cycles) of the same model throughout the production pipeline.

[0037] Then, according to the mixture production process, using this set of determined samples and the relevant mixing ratios (constituting the selected candidate), controlling the actual picking and processing process of the complex microbial community samples to obtain a mixed result product that is as close as possible to the target mixture characteristic spectrum in terms of analysis characteristics.

[0038] It has been proven that the present invention also provides a method for producing a co-cultured complex microbial community (CCMC) result product, the method including:

[0039] Based on the target mixture characteristic spectrum, using the above determination method to obtain a candidate, where the candidate represents a set of complex microbial community (CMC) samples and mixing ratios in the initial sample collection library;

[0040] Actually picking the set of CMC samples from the initial sample collection library;

[0041] Using a mixture production process configured with the obtained mixing ratios to process the picked CMC samples to obtain a CCMC result product.

[0042] The present invention helps to obtain CCMC resulting products that conform to the target characteristic spectrum (for example, adjusted to prevent or treat diseases, or restore necessary functions after drug-induced dysbiosis, etc.), and achieve high yields without wasting materials.

[0043] Therefore, according to the intended use (such as treatment, prevention, and environmental requirements, etc.), a mixture production strategy can be defined before the production process is executed.

[0044] The CCMC resulting products obtained thereby can be administered or transplanted into the human body or animal body, or applied to plants as fertilizers, or even applied to environmental media (including water bodies, soil, and subsurface materials), for example, pollution can be treated through bioremediation techniques.

[0045] Preferably, the above method can be used to produce microbiome ecosystem therapy products.

[0046] Optional features of embodiments of the present invention are defined in the appended claims. Some of these features are explained below with reference to the method, and these features can also be transformed into device / system features.

[0047] In certain embodiments, the mixture production process model includes:

[0048] Predicting the required intermediate CMC characteristic spectrum of the mixture of selected complex microbial community samples using a linear method at a given mixing ratio;

[0049] Using a CMC interaction model learned from the linearly predicted reference CMC characteristic spectrum and the corresponding true reference CMC characteristic spectrum to correct the intermediate CMC characteristic spectrum to a predicted CMC characteristic spectrum, and the predicted CMC characteristic spectrum represents the complex microbial community (CMC) product produced by the mixture.

[0050] This method accurately models the first pooling stage of the mixture production process, in which the samples will be mixed together. The reason is that this model contains a correction step, and this correction step reflects the non-linear behavior in the mixing process.

[0051] In certain embodiments, each candidate includes a mixing ratio representing the respective proportions of the CMC samples to be mixed. Therefore, in these embodiments, given the required target mixture characteristic spectrum, these proportions can be determined (through an evolutionary algorithm).

[0052] In certain embodiments, the mixture production process model further includes one or more loops (or iterations) of the following steps:

[0053] Predict a plurality of CCMC characteristic spectra based on the predicted CMC characteristic spectrum or based on the predicted mixing result characteristic spectrum of the previous cycle, where these CCMC characteristic spectra represent co-cultured CMC products obtained by performing different co-cultures on the same starting CMC product;

[0054] Based on the CCMC characteristic spectra, use the linear method to predict the intermediate mixing result characteristic spectrum, where the intermediate mixing result characteristic spectrum represents the second mixture of the CCMC product at a given mixing ratio;

[0055] Utilize the mixing interaction model learned from the linearly predicted reference mixing characteristic spectrum and the corresponding true reference mixing result characteristic spectrum to correct the intermediate mixing result characteristic spectrum into the predicted mixing result characteristic spectrum, where the predicted mixing result characteristic spectrum represents the mixed CCMC product produced by the second mixture.

[0056] This method uses a bioreactor to perform parallel co-cultures and accurately model the second stage of CCMC product amplification, where the CCMC product is generated by mixing the initial samples.

[0057] In some embodiments, each candidate further includes a mixing ratio representing the respective proportions of the CCMC products to be mixed. Thus, given a desired target mixture characteristic spectrum, the present invention is capable of determining these proportions (through an evolutionary algorithm).

[0058] In some embodiments, throughout the cycle, the same mixing ratio representing the respective proportions of the CCMC products to be mixed is used, which reduces the complexity at the candidate level. Of course, more complex algorithms can also be used, which include different mixing ratios required for successive cycles.

[0059] In some embodiments, the CCMC interaction model and the mixture interaction model are the same model, such as the same matrix. Since the entire prediction process requires single-model learning and the single model needs to be stored, processing and memory costs are saved.

[0060] Corresponding to the above model, the mixture production process may include a first pooling stage, that is, mixing the picked CCMC samples to obtain a CCMC product. Preferably, the mixing is performed according to the mixing ratio of the obtained candidate. The mixture production process may further include a second stage, that is, one or more iterations of amplifying the starting CCMC product, where the iteration includes: (i) co-culturing the CCMC product or the mixed result product obtained from the previous iteration in a bioreactor with respective operating parameters to obtain a CCMC product; (ii) mixing the CCMC products to obtain a mixed result product. Preferably, the mixing step (ii) is performed according to the mixing ratio of the obtained candidate.

[0061] In some embodiments, each candidate is defined by a gene array that includes a sample identifier and a mixing ratio, and each gene array defines a separate gene.

[0062] In some embodiments, an iteration in the evolutionary algorithm includes:

[0063] Evaluating the score of each candidate in the current population based on a mixture production process model and a target mixture characteristic spectrum;

[0064] Selecting a portion from the current population based on the evaluated scores;

[0065] Generating a new population of candidates based on the selected candidates, using gene crossover between the respective genes of the selected candidates and / or gene mutations within the gene array.

[0066] In some embodiments, evaluating the score includes: calculating the distance between the target mixture characteristic spectrum and the characteristic spectrum of the mixing result predicted according to the candidate, using the mixture production process model.

[0067] In some embodiments, the characteristic spectrum of a complex microbial community (sample or CMC or CCMC product) includes the relative abundances of the analytical features in this complex microbial community.

[0068] In a specific embodiment, the relative abundance represents the mass or volume ratio of the analytical features in the complex microbial community.

[0069] In some embodiments, the analytical features constituting the characteristic spectrum of a complex microbial community include one or more of the following features: taxon, gene, antibiotic resistance gene, function, metabolite feature, metabolite, RNA, and protein production, preferably including taxon.

[0070] In some embodiments, analytical techniques such as 16S rRNA gene amplicon sequencing, NGS shotgun sequencing, amplicon sequencing of non-16S rRNA genes, NGS amplicon-based targeted sequencing, phylogenetic microarray-based analysis, whole metagenome sequencing (WMS), polymerase chain reaction (PCR) identification, mass spectrometry (such as LC / MS type, GC / MS type, or MS / MS type), near-infrared (NIR) spectroscopy, nuclear magnetic resonance (NMR) spectroscopy (preferably using 16S rRNA gene amplicon sequencing or NGS), etc. are used to obtain the individual characteristic spectrum of a complex microbial community sample.

[0071] In some embodiments, the characteristic spectrum of a complex microbial community defines the analytical features associated with one or more microorganisms present in the complex microbial community from bacteria, archaea, viruses, bacteriophages, protozoa, yeasts, and fungi (preferably associated with bacteria and / or archaea).

[0072] In certain embodiments, the characteristic profile of a complex microbial community defines analytical features that specify the relative abundances of microorganisms considered at one or more of the following taxonomic levels: strain, species, genus, family, order, class, and phylum, preferably genus, family, and order.

[0073] In certain embodiments, the characteristic profile of a complex microbial community includes the relative abundances of bacterial and / or archaeal taxa in the complex microbial community considered at one of the following taxonomic levels: genus, family, order, class, and phylum.

[0074] In certain embodiments, the characteristic profile of a complex microbial community includes the relative abundances of bacterial and / or archaeal taxa in the complex microbial community, which are defined by the presence / absence or expression of certain genes and / or functions (e.g., production of butyrate, production of antibiotic resistance genes, production of enzymes (such as organophosphate hydrolase, phosphodiesterase, superoxide dismutase, etc.), production of antimicrobial peptides, production of organophosphate hydrolase, or production of other enzymes useful in bioremediation processes, etc.).

[0075] In certain embodiments, the initial sample collection library includes samples selected from the following groups: samples of original complex microbial communities, engineered / treated complex microbial community samples, artificial complex microbial community samples (e.g., bacterial communities obtained by mixing isolated strains), samples containing genetically modified organisms (e.g., bacteria, archaea, phages, viruses), and virtual complex microbial community samples.

[0076] In certain embodiments, the initial sample collection library includes one or more of feces, skin, oral cavity, vagina, nasal cavity, tumor, human, animal, plant, water body, and soil samples. For example, the initial sample collection library may include one or more fecal samples from at least one donor (preferably from at least two donors).

[0077] Another aspect of the present invention relates to a computer device including at least one microprocessor that can execute the steps of any of the above methods. Thus, the computer device can send signals to control a production device to actually pick and process complex microbial community (CMC) samples from the initial sample collection library using a mixture production process to obtain a co-cultured complex microbial community (CCMC) resultant product.

[0078] Another aspect of the present invention relates to a non-transitory computer-readable medium for storing a program, which, when executed by a microprocessor or computer system in a device, enables the device to execute any of the above methods.

[0079] At least part of the method according to the present invention can be implemented by a computer. Accordingly, the present invention may take the form of a complete hardware embodiment, a complete software embodiment (including firmware, embedded software, microcode, etc.) or an embodiment combining both software and hardware, which may be collectively referred to as "circuit", "module" or "system" in the present invention. In addition, the present invention may take the form of a computer program product contained in any tangible expression medium, and the medium contains computer-usable program code.

[0080] Since the present invention can be implemented in software, the present invention can be embodied as computer-readable code and provided to a programmable device on any suitable carrier medium. The tangible carrier medium may include storage media such as hard disk drives, tape devices or solid state memory devices, etc. The transient carrier medium may include signals such as electrical signals, electronic signals, optical signals, acoustic signals, magnetic signals or electromagnetic signals (such as microwave or radio frequency (RF) signals). BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 Shows a complex microbial community mixing platform for implementing embodiments of the present invention;

[0082] Figure 1a Shows that the behavior of error measurement during modeling depends on the hyperparameters of the regularization term;

[0083] Figure 2A Schematically shows the production process of the first mixture of samples;

[0084] Figure 2B Schematically shows the production process of the second mixture of samples;

[0085] Figure 2C Schematically shows the production process of the third mixture of samples;

[0086] Figure 3A Schematically shows Figure 2A the modeling of the production process of the first mixture of

[0087] Figure 3B Schematically shows Figure 2B the modeling of the production process of the second mixture of

[0088] Figure 3C Schematically shows Figure 2C the modeling of the production process of the third mixture of

[0089] Figure 4 Shows a genetic algorithm designed specifically for candidate determination given the characteristic spectrum of a target mixture according to an embodiment of the present invention;

[0090] Figure 5Shows a schematic diagram of a computer device according to an embodiment of the present invention;

[0091] Figure 5a , 5b and 5c show the results of the first experiment of the present invention based on a mixed native complex microbial community sample;

[0092] Figure 6a , 6b and 6c show the results of other experiments of the present invention based on a mixed co-culture complex microbial community sample;

[0093] Figure 7a and 7b show the results of other experiments of the present invention in mixing native samples and co-culture samples;

[0094] Figure 8 Shows PCA based on genus relative abundances, where the genus relative abundances are obtained from NGS shotgun sequencing of samples in Experiment 2;

[0095] Figure 9 Shows the PCA-based method used in Experiment 2;

[0096] Figure 10a Shows the leave-one-out baseline prediction means in Experiment 3;

[0097] Figure 10b Shows the T0 baseline prediction results in Experiment 3;

[0098] Figure 11a Shows the distribution of MSE (left figure) and Bray-Curtis (BC) distance (right figure) between each pair of target feature spectra and corresponding predicted feature spectra after being processed by a genetic algorithm in Experiment 4;

[0099] Figure 11b Shows the comparison between one target feature spectrum and the related predicted feature spectrum among 100 target feature spectra in Experiment 4. Detailed implementation manners

[0100] The present invention relates to the processing of complex microbial communities or CMC (meaning complex microbial communities) or "microbiota" or "microbiota samples", which includes mixing or "pooling". More specifically, it relates to methods and devices for predicting the products of mixed co-culture complex microbial communities (CCMC) using a learned prediction model, and methods and devices for determining an initial complex microbial community sample given a target mixture feature spectrum using a learned prediction model.

[0101] The terms "microbiota", "microbiota composition", and "complex microbial community" or "CMC" used in the present invention are used interchangeably and refer to any microbial population that includes a large number of different species of microorganisms that live together and may interact. Microorganisms that may be present in a complex microbial community include yeasts, bacteria, archaea, viruses, fungi, algae, bacteriophages, and any protists from different sources (e.g., from soil, water, plants, animals, or humans).

[0102] The microbiota according to the present invention includes naturally occurring complex microbial communities (e.g., gut microbiota, i.e., the microbial population living in the gut of an animal), and "engineered complex microbial communities", i.e., complex communities generated by a transformation step, wherein the transformation step includes adding isolated beneficial strains to remove potentially harmful microorganisms (e.g., by using rare-cutting endonucleases targeting specific genes of pathogenic symbionts), culturing and amplifying under specific conditions (e.g., co-culture in a suitable medium), etc. The "isolated beneficial strains" in the present invention refer to natural strains known to have beneficial effects under certain conditions (e.g., Akkermansia muciniphila), as well as genetically modified strains, including strains in which potentially harmful genes have been knocked out (e.g., by using rare-cutting endonucleases such as Cas9), and strains into which exogenous genes have been introduced (e.g., by using bacteriophages or CRISPR systems).

[0103] The complex microbial communities and microbiota according to the present invention include "raw" or "native" complex communities or microbiota, i.e., communities obtained directly from a source or donor without post-treatment, and "treated complex microbial communities", including engineered complex communities or microbiota, and any complex microbial communities generated by treating, post-treating, or transforming one or more natural raw complex microbial communities (e.g., complex communities or microbiota that have been filtered, frozen, thawed, and / or lyophilized, and / or complex communities or microbiota picked, isolated, or separated from their initial matrix using techniques well-known to those skilled in the art (e.g., the techniques described in WO 2016 / 170285 and WO 2017 / 103550)).

[0104] The terms "sample", "complex microbial community sample", "CMC sample", and "microbiota sample" are used interchangeably and refer to an initial complex community or microbiota that can be used for processing (including mixing) in the context of the present invention. The terms "product", "mixture", "complex microbial community product", "complex microbial community mixture", "CMC product", and "microbiota mixture" are used interchangeably and refer to an intermediate complex community or microbiota that can be used for further processing (such as co-culturing or additional mixing) in the context of the present invention. The terms "microbiota" and "complex microbial community" refer to either of the above initial or intermediate complex communities and are used interchangeably.

[0105] The term "microbiota ecosystem therapy product" in the present invention refers to any composition containing a complex microbial community (which can be naturally occurring, engineered, native, or processed), but it should be suitable for administration to an individual in need. Microbiota ecosystem therapy (MET) aims to modify the microbiota of an individual to obtain health benefits (such as preventing or alleviating disease symptoms, increasing the probability of an individual's response to treatment, etc.). Generally, for an individual in need, replacing at least part of the dysfunctional and / or damaged ecosystem with a different complex microbial community can complete the microbiota ecosystem therapy. Microbiota ecosystem therapy includes fecal microbiota transplantation (FMT). In the present invention, unless otherwise specified, the term "FMT" is widely used to refer to any type of microbiota ecosystem therapy.

[0106] As Figure 1 shown, the figure shows a complex microbial community processing platform 1, which implements an embodiment of the present invention and can obtain a sample 100 through an initial sample bank or collection bank 10. Although a single collection bank or sample bank is shown in the figure, the samples may be stored in multiple sub-banks, which together constitute the collection bank or sample bank 10.

[0107] The samples of the present invention may comprise or consist of microorganisms from one or more sources and / or from one or more donors 101.

[0108] The samples of the present invention can be from:

[0109] - a single source,

[0110] - at least two sources,

[0111] - a single donor,

[0112] - at least two donors,

[0113] - a single source and a single donor,

[0114] - a single source and at least two donors,

[0115] - at least two sources and a single donor, or

[0116] - at least two sources and at least two donors.

[0117] The term "source" as used in the present invention refers to any environment from which a sample is derived, such as soil, water body, plant part, animal body part or body fluid, or human body part or body fluid. For humans or animals, the source may refer to any part of the body (such as skin, nasal mucosa, etc.) or body fluid (such as intestinal contents (such as fecal samples)).

[0118] The term "donor" as used in the present invention refers to a plant, a physical location (such as a source like soil or water body), an animal or a human, preferably a human.

[0119] The donor can be preselected according to the methods and criteria described in the prior art (such as WO2019 / 171012A1).

[0120] In the illustrated example, certain samples (labeled 100d, 100e, 100f, 100g) are original complex microbial communities or microbiota, i.e., they are obtained directly from the donor without post - treatment.

[0121] Other samples (labeled 100a, 100b, 100c) are "treated samples", i.e., engineered complex microbial communities generated by treating, post - treating or transforming one or more natural original complex communities. As described above, the treatment may include filtering, centrifuging, co - culturing, freezing, lyophilizing the initial complex community, and even mixing the initial complex community, and also includes treatments aimed at separating spores and spore - forming bacteria, such as treatment with ethanol, chloroform or by heating.

[0122] As shown in the figure, the initial complex community can be a sample 100d, 100e, 100f, 100g belonging to the initial sample collection library 10, or an external sample 99.

[0123] The initial sample collection library 10 may include one or more samples from any origin (human, animal, plant, soil, etc.) and any source (feces, skin, nasal cavity, oral cavity, vagina, tumor, etc.), preferably one or more fecal samples from at least one donor (preferably from at least two donors).

[0124] According to a specific embodiment, the samples in the collection library 10 include fecal samples.

[0125] The fecal samples collected from donors can be controlled according to the methods and qualitative criteria described in the prior art (e.g., WO2019 / 171012A1). For example, the qualitative criteria for the samples can include: the consistency of the samples is between types 1 and 6 on the Bristol scale; there is no blood and urine in the samples; and / or there are no specific bacteria, parasites, and / or viruses in the samples, as described in WO2019 / 171012A1.

[0126] The fecal samples can be collected according to any method described in the prior art (e.g., WO2016 / 170285A1, WO2017 / 103550A1, and / or WO2019 / 171012A1). Preferably, the samples can be placed under anaerobic conditions after collection. For example, as described in WO2016 / 170285A1, WO2017 / 103550A1, and / or WO2019 / 171012A1, within 5 minutes after taking the samples, the samples can be placed in an oxygen-impermeable collection device.

[0127] The samples can be prepared according to the methods described in the prior art (e.g., WO2016 / 170285A1, WO2017 / 103550A1, and / or WO2019 / 171012A1).

[0128] All the samples 100a - 100g shown in the figure are actual samples stored in at least one sample library.

[0129] The samples 100y - 100z represented by the dashed lines are theoretical samples that have not been actually collected or produced from donors, and thus are not actually stored in the storage or sample library 10. As described below, these "virtual" samples 100y - 100z are intended to show the theoretical complex community characteristic spectra 110z envisioned by entities (such as computers, operators, researchers, etc.).

[0130] The initial sample collection library 10 can include only the native samples 100d - 100g, or only the processed samples 100a - 100c, or only the virtual samples 100y - 100z, or any combination thereof.

[0131] The samples in the initial sample collection library 10 are input into a mixture production process to produce a mixed resultant product, for example, input into a microbiota-based treatment method process to produce a substance with a mixed resultant characteristic spectrum that is taxonomically and functionally adjusted for the treatment target.

[0132] Figure 2A The first mixture production process of the samples is shown schematically.

[0133] The process begins with selecting (and picking) a set of Γ1 to k microbiota samples from the initial sample collection library 10. The selected samples are labeled as γ1…γ k .γ i,, , and these labels may simply be the identifiers of the samples in the initial sample collection library 10. Each sample in the initial sample collection library has a known microbiota signature.

[0134] A "signature" refers to a description of the relevant complex microbial community or microbiota composition (sample or mixture or product). For example, the signature specifies the relative abundances of the analytical features in the complex community or microbiota composition. "Relative" means that the sum of the abundances equals 1. The relative abundances can be expressed as the mass (or weight) or volume ratios of the individual analytical features in the complex microbial community.

[0135] The types of analytical features can vary according to the relevant application (e.g., in the therapeutic field, according to the disease targeted, in the bioremediation field, according to the pollutant to be eliminated). Analytical features are typically selected from taxa, genes, antibiotic resistance genes, RNA, functions, metabolite features, metabolites, and protein production. The signature can mix different types of analytical features (e.g., taxa and antibiotic resistance genes). A particular embodiment only contemplates using taxa to analyze the complex microbial community.

[0136] A function describes the known role of a protein or protein family (defined phylogenetically, e.g., the KEGG KO database or the NCBI COG database or the Enzyme Commission number database), or can also define the metabolic context (e.g., the BiGG model database at the reaction level, or the KEGG pathway at the metabolic pathway level). Some databases can be specialized databases, such as the CaZy database, which is a catalog of carbohydrate-active enzymes. Any of the above functional categories (or combinations thereof) can be used as features in the matrix model.

[0137] KEGG stands for "Kyoto Encyclopedia of Genes and Genomes", KO stands for "KEGG Ortholog", NCBI stands for "National Center for Biotechnology Information", COG stands for "Cluster of Orthologous Groups", and BiGG stands for "Biochemical Genetics and Genomics".

[0138] Currently, there are a variety of analytical techniques available for obtaining the characteristic spectra of complex communities, including: 16S rRNA gene amplicon (i.e., metagenome) sequencing, NGS shotgun sequencing, amplicon sequencing not based on the 16S rRNA gene, NGS amplicon-based targeted sequencing, 18S / ITS gene sequencing, metagenome sequencing, phylogenetic chip-based analysis, polymerase chain reaction (PCR) identification, mass spectrometry (e.g., LC / MS type, GC / MS type, or MS / MS type), near-infrared (NIR) spectroscopy, nuclear magnetic resonance (NMR) spectroscopy.

[0139] In Figure 2A the process, k samples are mixed or "pooled" together according to their respective mixing ratios M = [α1...α k , where Σα = 1, thus obtaining a mixed result product.

[0140] "Mixing" refers to any actual mixing of samples, thereby generating a new CMC product or a new microbiota composition. The resulting product can be used for the above-mentioned application or transplantation, and thus is also called a mixed result product. For example, the mixed result product can be used as a MET inoculum.

[0141] Figure 2B The production process of the second mixture of samples is shown schematically. This process includes the co-culture operation described below, and thus is a co-culture-based mixture production process.

[0142] The process begins with selecting (and picking) a group of Γ1 to k microbiota samples from the initial sample collection bank 10, as described above.

[0143] The k samples are mixed or "pooled" together according to their respective mixing ratios M = [α1...α k , where Σα = 1, thus obtaining a CMC product. Therefore, the CMC product is an intermediate product, or "pooled product", in the entire mixture production process.

[0144] The CMC product is fed into the co-culture step to obtain j co-culture products or CCMC products. This step involves j bioreactors F1,…,F j (j is an integer greater than 1).

[0145] The term "bioreactor" refers to any device (including containers) that can be used for co - culturing microorganisms, in which culture parameters (such as temperature, pH value, residence time, aeration volume, supply, etc.) can be controlled. In the present invention, the terms "fermenter" and "bioreactor" have the same meaning. Through the bioreactor, the main parameters in different environments can be integrated to reconstruct the collection environment of complex microbial communities. For example, as described in the prior art, such as the research paper published by Cordonnier et al. in 2015 ("Dynamic in vitro human gastrointestinal tract model as a relevant tool for evaluating the survival of probiotic strains and their interaction with the gut microbiota", Microorganisms. December 2015; 3(4):725 - 745), when the complex microbial community comes from the human colonic environment, the bioreactor can integrate the parameters of the human colonic environment, such as pH value, temperature, ileal effluent supply, residence time, and anaerobic state.

[0146] The same starting CMC product is fed into each of the j bioreactors. 'j' can be greater than or equal to 2, greater than or equal to 3, or equal to 4. The bioreactors have at least one different operating parameter. The operating parameters are preferably selected from pH value, temperature, pressure, culture time, residence time, aeration conditions, redox potential, culture medium, light source, and combinations thereof. For example, the j bioreactors have different pH set points. The co - culture step in the bioreactor can last for several hours or even several days, such as 3 days.

[0147] In Figure 2B the process, the CCMC products are mixed or "pooled" together according to their respective controlled ratios B = [β1, β2, …, β j to obtain a mixed resultant product.

[0148] More details of the co - culture step and its subsequent mixing process are provided in Application No. WO2022 / 136694.

[0149] Figure 2C The production process of the third mixture of the sample is shown schematically. Compared with the second process, this process cycles two or more co - culture steps. In this case, the CCMC products are obtained by co - culturing the same starting CCMC product in respective bioreactors (each iteration), and the CCMC products are iteratively mixed to obtain the final mixed resultant product.

[0150] The process starts with selecting (and picking) a group of Γ1 to k microbial community samples from the initial sample collection bank 10, as described above.

[0151] The k samples are mixed according to their respective mixing ratios M = [α1...α kare mixed or "pooled" together, where Σα = 1, thus obtaining the CMC product. Therefore, the CMC product is an intermediate product, or "pooled product", in the entire mixture production process.

[0152] The CMC product is fed into a multi-iteration (or multi-cycle) co-culture step, which involves j bioreactors F1,…,F j (j is an integer greater than 1). Preferably, the same j bioreactors are used throughout the co-culture iteration. However, this method is not mandatory, and different numbers of bioreactors (j1, j2,…, jN, where N is the number of iterations) can also be used. In addition, throughout the iteration, the bioreactors can have different distinguishing operating parameters or different distinguishing parameter settings.

[0153] In the first co-culture iteration, the CMC product is fed into j bioreactors to obtain j co-culture products of the first iteration or CCMC products. The CCMC products of the first iteration are mixed or "pooled" together according to their respective controlled ratios B1 = [β 1,1 , β 1,2 ,…, β 1,j1 to obtain the intermediate mixed result product or "mixed CMCC product" of the first iteration.

[0154] In the i-th co-culture iteration (i = 2…N), the intermediate mixed result product or "mixed CMCC product" obtained from the (i - 1)-th iteration serves as the new starting CMC, which is input into the bioreactors used in the i-th iteration to obtain the CCMC products of the i-th iteration. The CCMC products of the i-th iteration are mixed or "pooled" together according to their respective controlled ratios B i = [β i,1 , β i,2 ,…, β i,ji to obtain the intermediate mixed result product or "mixed CMCC product" of the i-th iteration.

[0155] Preferably, the same j bioreactors and mixing ratios are used throughout the iteration, which means that for any values of (i, m), B i = B m (denoted as 'B'). This is an example of Figure 2C .

[0156] Figure 2B The explanations regarding bioreactors and co-culture steps in

[0157] also apply to each iteration in the third mixture production process. The intermediate mixed result product or "mixed CMCC product" of the N-th iteration is the final mixed result product.

[0158] In certain embodiments of the third mixture production process, during one or more iterations of the process, additional samples from the initial collection library may be added. For example, a new sample is added during the CMCC product mixing process of the i-th iteration. The new sample may have its own ratio β, which is an additional ratio outside of other mixing ratios. New samples may be added in each or some of the multiple iterations, for example, in the third and subsequent iterations, or once every two iterations, or in the last iteration. The same additional sample may be added during the iteration process, or different samples may be added during the iteration process. Of course, more than one additional sample may be added in a given iteration.

[0159] The present invention employs a model of a related mixture production process, such as a learning model. Specifically, the model can be designed and learned in "modules", that is, for mixing operations and co-culture operations respectively, as Figure 3A - 3C shown.

[0160] Figure 3A is shown schematically Figure 2A the first mixture production process modeling shown. The modeling only includes a single learning model 130, called a pool predictor. The pool predictor is used to predict the mixed result product from the samples Γ used in the mixture and their respective proportions M (mixing ratios) in the mixture according to the characteristic spectra. As described below, the advantage of this model is its ability to predict multiple characteristic spectra in one calculation.

[0161] The known characteristic spectra of the samples are shown schematically by reference numeral 110, where the predicted characteristic spectra of the mixed result product (here the CMC product) are labeled as R'.

[0162] This modeling involves predicting the composition of the mixture obtained by mixing the samples 100a - 100z belonging to the initial sample collection library 10. The prediction is divided into two aspects:

[0163] Using a linear method to predict the intermediate CMC characteristic spectra required for the mixture of selected complex microbial community samples Γ;

[0164] Using an interaction model learned from the linearly predicted reference CMC characteristic spectra and the corresponding true reference CMC characteristic spectra to correct the intermediate CMC characteristic spectra to the predicted CMC characteristic spectra. The interaction model preferably uses an interaction square matrix learned from the linearly predicted reference CMC characteristic spectra and the corresponding true reference CMC characteristic spectra.

[0165] Once the interaction model or matrix is learned, this learned interaction model (more specifically, a matrix-based model) can provide accurate prediction results, thus providing relevant hints for the final product without consuming any materials from the initial sample collection library.

[0166] Since the prediction can be implemented by a computer, therefore, even though the number of mixtures to be predicted is large, the number of available samples in the initial sample collection library 10 is large, and the number of features for analyzing complex microbial communities (samples and mixtures) is also large, the predicted CMC feature spectra can still be obtained quickly.

[0167] As Figure 1 shown, it is preferable to provide the feature spectra of the actual samples 100a - 100g, such as 16S sequencing, through an analyzer (or sequencer) 12. The corresponding individual feature spectra thus obtained are labeled 110a - 110g and form the initial feature spectrum collection library or library 11. Of course, 16S rRNA sequencing is not mandatory. As described above, other methods can also be used alone or in combination to provide the feature spectra 110.

[0168] Regardless of the analysis technique used, the individual feature spectra are converted into the same format and stored in a computer memory (not shown) in the form of a matrix or vector a x The coefficient a x (j) of the individual feature spectrum 'x' represents the relative abundance of the analysis feature 'j' in the sample under consideration.

[0169] As mentioned before, for example, by defining the coefficient a x (i) to represent the relative abundance of the analysis feature 'j' in the theoretical sample, some individual feature spectra 110z can be manually constructed by the operator.

[0170] Therefore, the initial feature spectrum collection library 11 can include only the individual feature spectra 110d - 110g corresponding to the native samples 100d - 100g, or only the individual feature spectra 110a - 110c corresponding to the processed samples 100a - 100c, or only the virtual feature spectra 110y - 110z corresponding to the virtual samples 100y - 100z, or any combination thereof.

[0171] Any other feature spectra processed thereafter (such as the so-called intermediate feature spectra or mixture feature spectra) are processed in the same feature spectrum format, such as a vector consisting of the same analysis features 'j' in the same order.

[0172] Preferably, obtaining a bacterial abundance signature means that the signature specifies the relative abundances of analytical features associated with bacteria. More generally, the signature of a complex microbial community can define analytical features associated (preferably associated with bacteria and / or archaea) with one or more microorganisms present in the complex community (bacteria, archaea, viruses, bacteriophages, protozoa, yeasts, algae, and fungi). Of course, the analytical features in the same signature can relate to different microorganisms as described above.

[0173] Preferably, obtaining a genus-based bacterial abundance signature means that the analytical features describe the relative abundances of bacteria at the genus level in a complex microbial community. More generally, the signature of a complex microbial community can define analytical features that specify the relative abundances of microorganisms considered at one or more of the following taxonomic levels: strain, species, genus, family, order, and phylum, preferably genus, family, order, and phylum.

[0174] Under the control of module 14, the predictor module 13 performs a prediction operation to implement the pool predictor 130 in the first mixture production process. Module 14 (referred to as the "testing and decision module" or "decision module") drives platform 1 in order to predict the mixture signature, and / or to determine a set of samples given a target mixture signature, and / or to produce at least one mixture or CCMC result product.

[0175] Modules 13 and 14 are preferably implemented by a computer and have an input / output or user interface (such as a keyboard, mouse, screen) so that an operator can interact with platform 1. The target mixture signature can be defined by the operator interacting with module 14 using the user interface.

[0176] As Figure 3A shown, the pool predictor 130 is a matrix-based predictor that includes two steps for predicting the resulting mixture signature from the initial signatures of the samples mixed together.

[0177] Matrix A defines the individual signatures of all available samples in the acquisition library 10. Matrix A can be formed by the analyzer or sequencer 12, or at least by the individual signatures obtained from the analyzer. In addition, any virtual individual signatures can also be added to the matrix.

[0178] Preferably,

[0179] where j = 1…m, m is the number of analytical features considered, and n is the number of individual signatures 110 in the initial signature acquisition library 11, and thus also the number of samples 100 (including virtual samples) in the initial sample acquisition library 10.

[0180] The square matrix W is the pooling or CMC interaction matrix defined above, which is used to model the interactions between microorganisms. More details about the modeling matrix W, including how to learn the modeling matrix W, will be provided below. The pooling interaction matrix is designed to represent the non-linear interactions between various analytical features when samples are mixed together.

[0181] The prediction operation includes a first step based on a matrix, that is, for at least one mixture of selected samples, using the matrix A to predict the intermediate mixture characteristic spectrum formed by the matrix I: I = P * A, where P is a matrix representing at least one mixture of selected samples from the acquisition library 10.

[0182] The matrix P can define each mixture according to the mass or volume ratio of the samples in the initial sample acquisition library. If necessary, P can define a single mixture.

[0183] For example,

[0184] where {p x (j)} j defines the mixture 'x', and p x (k) represents the proportion of sample k (k ranges from 1 to the number of samples Nsamp in the acquisition library 10 / 11). The sum of the proportions is equal to 1: ∑ k=1…Nsamp (p x (k)) = 1. In other words, for the mixtures involved, p x (k) = α γ k .

[0185] where sample r is not used in mixture x, p x (r) = 0. In other words, for the mixtures involved, for each sample index r not included in Γ, p x (r) = 0.

[0186] The advantage of the matrix-based method is that it can predict different numbers of mixtures simultaneously: each row of the matrix P defines a mixture to be predicted (so 't' mixtures are defined in the above example), and the number 't' varies depending on the prediction.

[0187] For example, the prediction operation I = P * A can be implemented by a computer.

[0188] By defining the linear prediction or "intermediate" CMC characteristic spectrum of the mixture P, the matrix is obtained, where j = 1…m. Since these intermediate CMC characteristic spectra do not take into account the interactions between microorganisms during actual mixing, they are "naive" predictions.

[0189] Thus, according to the present invention, the prediction operation includes a second step of using an interaction model (specifically, the pooling interaction matrix W) to correct the intermediate CMC feature spectrum (i.e., matrix I) into a predicted CMC feature spectrum, represented by the matrix : R = I * W, where j = 1…m.

[0190] Therefore, without consuming any materials from the collection library 10, the predicted CMC feature spectra of different amounts of mixtures can be quickly obtained.

[0191] As expected, the relative abundances r x (j) are non-negative values and together form a complete composition (i.e., the sum of the relative abundances of a given mixture 'x' is equal to 1). However, this may not be the case for matrix multiplication. Therefore, embodiments of the present invention include post-processing the resulting matrix R into R' to meet the biological constraints.

[0192] For example, each negative value in matrix R is truncated, which means that negative abundances are set to 0. Thereafter, the relative abundances r x (j) are normalized, i.e., adjusted (e.g., using linear interpolation) to r' x (j) such that their sum is equal to 1:

[0193] ∑ j=1…m (r′ x (j)) = 1. The final mixed result matrix is as follows where {r′ x (j)} j is a vector representing the predicted CMC feature spectrum of mixture x (defined by {p x (k)} k ). Optionally, before normalization, for analytical features that do not exist in the initial samples mixed together (i.e., for all samples x mixed together, a x (j) are all zero), their non-zero relative abundances (i.e., non-zero values in matrix R) are set to zero.

[0194] The efficiency of the pool prediction stems from modeling the true positive and negative interactions between microorganisms in the mixed samples to form a matrix, namely the so-called pooling interaction matrix W. Then, the true CMC feature spectrum is predicted by efficiently using a two-step matrix-based method.

[0195] For a given set of m analytical features, the pooling interaction matrix W is learned. If the analytical features are re-ordered in the feature spectrum, the coefficients of the pooling interaction matrix W should also be re-ordered accordingly.

[0196] The m analytical features can also evolve over time. For example, due to the discovery of new features, some features become less significant and are thus deleted, and / or some features become more precise as they can be split into more features. The evolution of analytical features can also stem from the enhancement of analytical / sequencing methods and analyzers / sequencers 12 that provide new analytical data, as well as the improvement of bioinformatics methods that combine algorithms with feature reference databases.

[0197] For different diseases or specific treatments, different sets of analytical features can also be considered.

[0198] Both the analytical features themselves and the number of features in the set of analytical features are likely to evolve or change.

[0199] Therefore, each time a new set of analytical features is considered, the pooled interaction matrix W can be recalculated, as well as the matrix A that describes the initial feature spectrum acquisition library 11. The calculated pooled interaction matrix W can be stored in the memory of the pool predictor 130 so that this matrix W can be reused when the corresponding set of analytical features is reused.

[0200] Preferably, the pooled interaction matrix is obtained using machine learning. Machine learning is performed using a set of training data. The training data is constructed from the reference CMC product'ref', where the reference CMC product is generated from multiple mixtures {p ref (k)} of sample k.

[0201] Within 10 minutes to 3 hours, preferably between 30 minutes and 1.5 hours, the actual reference mixture of the sample is homogenized. The homogenization process will be carried out in the temperature range of 0°C to 10°C, preferably in the range of 2°C to 6°C, and more preferably at about 4°C.

[0202] Then, starting from the beginning of mixing, the mixture can remain in a stable state for several hours (at least 16 hours, preferably 24 hours).

[0203] This means that for a stable mixture at 4°C, the pooled interaction matrix represents the interactions that should occur between microorganisms.

[0204] Other pooled interaction matrices representing other mixing conditions can also be generated.

[0205] The individual feature spectrum {a x (j)} j (where j = 1…m) of sample x is known or can be obtained from the sequencer for analyzing sample x. Therefore, using the above linear formula I = P*A, the reference CMC feature spectrum {i ref (j)} j can also be known.

[0206] Referring to the characteristic spectrum of the reference CMC product 'ref', also known as the reference CMC characteristic spectrum {r true (j)} j , which is also known or can be obtained from a sequencer for analyzing the reference CMC product 'ref'.

[0207] The predicted reference CMC characteristic spectrum {r pred (j)} j corresponds to the linearly predicted reference CMC characteristic spectrum {i ref (j)} j and the matrix product with the pooling interaction block matrix W (during the learning process): for a single reference CMC product 'ref', R pred = I ref * W or {r pred (j)} j = {i ref (j)} j * W.

[0208] Machine learning aims to minimize the prediction error of the reference CMC characteristic spectrum. In other words, machine learning aims to obtain the minimum value of a formula based on the true reference CMC characteristic spectrum and the corresponding linearly predicted reference CMC characteristic spectrum where the difference between 'pred - i' and

[0209] 'true - i' refer to the predicted reference CMC characteristic spectrum and the true reference CMC characteristic spectrum corresponding to the same reference CMC product 'i' respectively. 'N' represents the number of reference CMC products considered.

[0210] The training data for machine learning is I ref and R true .

[0211] In some embodiments, the formula for finding the minimum value is the residual vector

[0212] {r true-i (j)} j - {r pred-i (k)} k = {r true-i (j)} j - {i ref (k)} k * W,

[0213] or the residual matrix R true - R pred = R true - I ref * W.

[0214] Any norm can be used: L1, L2, Lp, etc. Preferably, the sum of squared differences (SSD) or its derivative, the mean squared error (MSE), can be used. Alternatively, the minimum chi-squared method can also be used.

[0215] Then, machine learning can aim to solve the following convex optimization problem: where, ‖.‖ 2 is the MSE and N is the CMC product considered in R true and R pred

[0216] In embodiments that avoid overfitting W, this formula adds a regularization term to the difference, preferably a regularization term based on ridge regression (L2). In a variant, a regularization term based on lasso regression (L1) can be used. In another variant, a regularization term based on elastic net regularization (a regularization regression method that linearly combines the L1 penalty term of the lasso regression method and the L2 penalty term of the ridge regression method) can be used. The advantage of the ridge regression method is that it helps to generate a larger number of non-zero coefficients in W, thus enabling more accurate modeling of the interactions between the analyzed features.

[0217] Therefore, machine learning aims to solve the following convex optimization problem: where, ‖.‖2 is the regularization term (preferably ridge regression), I is the identity matrix, and λ is the hyperparameter of the regularization weight.

[0218] In addition, during the machine learning process, constraint conditions can be set such that there are no negative relative abundances in R pred and the sum of the relative abundances of each predicted reference CMC feature spectrum is 1. In other words, preferably, the modified matrix R’ pred is used, that is, first, the negative relative abundances in I ref *W are truncated, and then the sum of the relative abundances of each predicted reference CMC feature spectrum (i.e., each row in I ref *W) is normalized to 1. The modified I ref *W is labeled as Therefore, in an embodiment, machine learning aims to solve the following convex optimization problem:

[0219]

[0220] The training dataset (assuming there are N reference CMC products) is divided into two subsets, one for hyperparameter λ optimization and the other for W optimization.

[0221] ​There are currently known to be a variety of different methods for optimizing λ, including minimizing information criteria (such as minimizing the Akaike information criterion or the Bayesian information criterion) or minimizing cross - validation residuals, which use a first subset of the training data. For such optimization methods, the default setting of W may be different from ID.

[0222] For example, for the training data set and the test data set, when λ varies between 10 -5 and 10 3 , calculate the MSE of the above formula (hyperparameter λ optimization by dividing subsets). The calculated MSE is as Figure 1a shown.

[0223] As shown in the figure, when λ is small, the MSE of the training data set is close to 0, while the MSE of the test data set is very high. In this case, the model is over - fitted. On the other hand, when λ is high, the model is under - fitted. Therefore, λ can be selected to minimize the MSE of the test data set.

[0224] Once λ is known, W is learned based on a second subset of the training data and minimizing cross - validation residuals: perform k - fold cross - validation.

[0225] Divide the training data subset (i.e., {r true-i (j)} j and {i ref-i (j)} j ) into k subsets. Preferably, k is selected from the integers 3 - 20, preferably selected from the integers 4 - 10, and more preferably, k is equal to 5.

[0226] Select each of the k subsets in a round - robin manner (cyclic order) in turn, define the test subset, and the remaining k - 1 subsets will define the training subset.

[0227] In each round of k - fold cross - validation, the model is trained using the training subset, that is, by solving to obtain W. The advantage is that all the CMC feature spectra of the linear predictions of the training subset are input into a single matrix I ref (while inputting the true mixture feature spectrum into R true ) so as to complete the learning of W through a one - time operation.

[0228] Then, use the test subset to check the learned pooled interaction matrix W: substitute the test subset into the matrix - based model R true = I ref *W, and a score based on any norm (such as MSE ) can be obtained.

[0229] Since this operation needs to be repeated for each k test subset, k scores are finally obtained.

[0230] Then, the learning pooling interaction matrix W corresponding to the best score (i.e., the lowest score) can be selected to configure the pool predictor 130.

[0231] Of course, as long as the learning pooling interaction matrix W can be obtained, other machine learning methods can be adopted.

[0232] In some embodiments, the analysis features of the sample 100 (i.e., the features used to form matrix A) are the same as the analysis features of the final mixing result (i.e., the features used to form matrix R). As described above, these analysis features can be taxon, gene, antibiotic resistance gene, function, RNA, metabolite features, and metabolite and protein production.

[0233] In other embodiments, the analysis features of the sample 100 (i.e., the features used to form matrix A) are different (partially or completely) from the analysis features of the final mixing result (i.e., the features used to form matrix R). Any of the above analysis features (taxon, gene, function, etc.) can be adopted.

[0234] For example, when using analysis techniques such as NGS shotgun sequencing, compared with 16S sequencing, more analysis features can be obtained for each sample 100. Therefore, NGS shotgun sequencing can be used to analyze the sample 100 (thereby forming a matrix A containing NGS shotgun analysis features), while the final mixing result can retain a smaller number of analysis features, such as the features obtained by 16S sequencing (thus forming a matrix R containing 16S analysis features). In this case, a matrix I containing NGS shotgun analysis features is formed, and the pooling interaction matrix W is not a square matrix. It can still model the interactions between microorganisms, but in this example, it shows the relationship between NGS shotgun analysis features and 16S analysis features.

[0235] In the specific implementation process, to reduce the massive NGS shotgun analysis features, principal component analysis (PCA) needs to be performed to reduce the large number of features to k principal components (k PCs). In one embodiment, PCA is performed on the sample analysis features, that is, when constructing matrix A. In another embodiment, a matrix I containing massive analysis features is generated, and then PCA is performed on matrix I.

[0236] As described above, when the input mixture is input, the pool predictor 130 outputs the final mixing result matrix

[0237] For more details regarding the pool predictor, reference may be made to co-pending application PCT / EP2022 / 062226. The said application specifically provides experimental results regarding the modeling of the pool predictor, which are hereby incorporated by reference into the present application.

[0238] Figure 3B is shown schematically Figure 2B the modeling process of the second mixture production process of. In this case, the predictor module 13 includes a first pool predictor model 130, a fermenter predictor model 131, and another mixing model (usually another pool predictor model 130). Since the other model reflects the mixing of CCMC products, it may be different from the first pool predictor 130. For example, this model can be learned based on only one set of (predicted and true) CCMC product characteristic spectra, thereby generating different pooling interaction matrices W. However, for ease of implementation (saving memory and learning costs), these two pool predictor models can be the same already learned model.

[0239] For example, the initial mixture {p1(k)} of the initial sample 100 k is input into the pool predictor model 130 to obtain the predicted CMC characteristic spectra {r′1(l)} l , and the CMC characteristic spectra are then input into the fermenter predictor model 131 to obtain j CCMC characteristic spectra, {r2 i (l)} l (i = 1…j), and then these characteristic spectra are input into the second pool predictor 130 to obtain the predicted mixture composition of the final mixed result product.

[0240] The fermenter predictor model 131 is used to predict CCMC products based on characteristic spectra starting from an input or starting complex microbial community (specifically, the CMC products of the initial sample mixture selected from the initial sample collection library).

[0241] As shown in the figure, the fermenter predictor model 131 is a matrix-based model.

[0242] The input matrix Q is used to describe the characteristic spectra of the starting CMC (if there are multiple, one per row) and their corresponding growth conditions, i.e., the operating parameters of the bioreactor.

[0243] The submatrix Q1 is used to describe the characteristic spectra (such as taxonomic abundance values, ranging from 0 to 1): where q1 x ={q1 x (j)} j defines the CMC product 'x', q1 x(k) represents the relative abundance of taxon k (where k ranges from 1 to the number of analytical features). Sub - matrix Q2 is used to describe the growth conditions (such as pH value, temperature, etc. as described above): where q2 x ={q2 x (j)} j defines the growth conditions for the CMC product 'x', and q2 x (k) represents the presence status of condition k. Each column k of sub - matrix Q2 corresponds to a given growth condition (e.g., pH = 5.3, pH = 6.3, pH = 7.3, temperature = t0, temperature = t1, etc.). If the bioreactor used for co - culturing the CMC product 'x' meets condition k, q2 x (k) is set to 1, otherwise it is set to 0. As shown in the figure, Q = Q1|Q2, where '|' is the matrix concatenation operator, indicating that the 'x' - th row of matrix Q (denoted as q x ) is q1 x |q2 x . Thus, each q x defines the starting CMC product (q1 x ) and its growth conditions (q2 x ).

[0244] The CCMC prediction operation includes a matrix - based prediction step, that is, using the co - culture interaction matrix Y, for one or more starting CMC products and their respective growth conditions q1 - q t , predicting the CCMC feature spectrum composed of the matrix : R2 = Q*Y. As shown in the figure, Y has two parts: Y1 describes the interaction between taxa, and Y2 describes the influence of growth conditions (conditions defined in Q2, such as pH value) on taxa.

[0245] In the second mixture production process, the same starting CMC product is fed into j bioreactors with different operating parameters. This means that j rows in matrix Q are used to describe the same starting CMC product (q11 = q12 = … = q1 j ), but corresponding to different operating parameters (for any m, n belonging to 1…j, q2 m ≠q2 n ). Of course, different numbers of rows can be used in one calculation, for example, predicting the CCMC products (and their feature spectra) generated after different starting CMC products pass through the bioreactor.

[0246] Similar to the pool predictor, the expected relative abundance r2 x(j) is non - negative and together forms a complete composition (i.e., the sum of the relative abundances of a given mixture 'x' equals 1). Thus, each negative value in R2 can be truncated. For non - zero relative abundances of analytical features not present in the starting CMC product (i.e., non - zero values in R2), they can be set to zero. The relative abundances r2 x (j) can be normalized.

[0247] The co - culture interaction matrix Y represents the interactions between the various analytical features (taxa) under given operating parameters. Preferably, this interaction matrix is obtained using machine learning. Machine learning is performed using a set of training data. The training data is constructed from the reference CCMC product'ref2' generated from one or more reference complex microbial communities under the following growth conditions (i.e., operating parameters): {q ref}.

[0248] The training of matrix Y can follow the following mixture production process. In fact, co - culture experiments of the starting CMC product can be carried out in at least 2 (preferably 3) bioreactors, where the bioreactors have at least one different parameter selected from the group consisting of pH value, temperature, pressure, culture time, residence time, aeration conditions, redox potential, culture medium, light source, and combinations thereof. For example, a single differential set point can be used. The pH set point of the bioreactor can be selected between 4.5 - 8.0, preferably between 5 - 7.6, and more preferably between 5.2 - 7.4. The other parameters of different bioreactors are set to fixed values. Multiple rounds of co - culture can be carried out simultaneously in parallel or at different times with a delay. The resulting CCMC products are mixed and input into a new round of co - culture and reused as CMC products one or more times. Once multiple rounds of co - culture are completed, the resulting CCMC products need to be sequenced (16S, shotgun sequencing, etc.). The data generated therefrom will be used to train the relevant Y matrix.

[0249] Other co - culture interaction matrices representing other growth conditions can be generated.

[0250] An individual vector q is constructed by simply splicing the known individual feature spectra and growth conditions of the starting CMC product. ref .

[0251] The CCMC feature spectrum of the reference CCMC product'ref2', also known as the true reference CCMC feature spectrum {r2 true (j)} j is also known or can be obtained from the sequencer used for analyzing the reference CCMC product'ref2'.

[0252] The reference predicted CCMC feature spectrum {r2 pred (j)} jcorresponding to the individual vectors {q x (j)} j The matrix product between the co-culture interaction matrix W (during the learning process): for a single reference CCMC product'ref2', R2 pred = Q ref * Y or {r2 pred (j)} j = {q x (j)} j * Y.

[0253] Similar to the pool predictor, machine learning aims to minimize the prediction error (e.g., MSE) of the reference CCMC feature spectrum. For example, the error to be minimized can be or where ‖.‖ 2 is the MSE, N is the number of CCMC products considered in R2 true and R2 pred , ‖.‖2 is the regularization term (preferably ridge regression), and λ is the hyperparameter of the regularization weight. The teachings regarding the pool predictor above still apply.

[0254] The training data for machine learning is Q ref and R2 true , which can be split into subsets according to the aforementioned method, for example, learning Y in a polling manner.

[0255] Then, the learned co-culture interaction matrix Y corresponding to the best score (i.e., the lowest score) can be selected to configure the fermenter predictor 131.

[0256] During use, the output R2 of the fermenter predictor 131 is input into the second pool predictor 130. Therefore, the latter predicts the mixture composition, and the CCMC products are obtained by co-culturing the same complex microbial community (CMC) products in their respective bioreactors separately, and the CCMC products are mixed to obtain the mixture composition, where the method includes:

[0257] (a) Predicting the intermediate mixture feature spectrum of the CCMC product mixture by a linear method;

[0258] (b) Using the mixture interaction model learned from the linearly predicted reference mixture feature spectrum and the corresponding reference true mixture feature spectrum to correct the intermediate mixture feature spectrum into the predicted mixture feature spectrum.

[0259] Previously, the cascaded tandem pool predictor 130, the fermenter predictor 131, and the pool predictor 130 ensured that the overall prediction included the following steps: (by the first pool predictor 130) predicting the CMC characteristic spectrum of the CMC product based on the sample characteristic spectrum of the complex microbial community (CMC) sample selected from the initial sample collection library; (by the fermenter predictor 131) predicting the CCMC characteristic spectrum of the CCMC product based on the CMC characteristic spectrum.

[0260] Given this cascaded tandem, the intermediate mixture characteristic spectrum (generated by the second pool predictor 130) is predicted linearly according to the CCMC characteristic spectrum.

[0261] Therefore, the method can predict the mixture composition obtained from the Figure 2B second mixture production process shown.

[0262] Figure 3C The third mixture production process modeling of Figure 2C is shown schematically. In this case, the predictor module 13 includes the first pool predictor model 130, and a cycle composed of the fermenter predictor model 131 and another mixing model (usually another pool predictor model 130). As described above, these two pool predictors can be different pool predictors. However, for ease of implementation (saving memory and learning costs), these two pool predictor models can be the same already learned model. In a variant, the modeling can involve a paired 'i' cascade composed of the fermenter predictor model 131 i and the pool predictor model 130 i instead of a cycle composed of the fermenter predictor model 131 and the pool predictor model 130. In this way, different pooling (or fermenter) interaction matrices can be learned.

[0263] As shown in the figure, when calculating the predicted mixture composition of the final mixed result product, as long as the Nth (predefined) cycle is not reached, the predicted product characteristic spectrum generated by the second pool predictor 130 will be fed back to the fermenter predictor 131 for the next cycle. In other words, the mixture composition prediction method includes multiple iterations of the above steps (a) and (b), where the predicted mixture characteristic spectrum obtained in the previous iteration step (b) is used to obtain the characteristic spectrum of each CCMC product in the next iteration step (a).

[0264] In some embodiments not shown, when it is expected to add other (e.g., all) initial CMC samples during one or more co-culture iterations, the second pooling 130 optionally adds the characteristic spectra of one or more CMC samples from the initial collection library. These additional samples also have their own exponents γ and mixing ratios β.

[0265] Return Figure 1, as described above, the predictor module 13 is controlled by the module 14. Thus, the module 14 can configure the predictor module 13 according to any one of the Figure 2A - 2C ) mixture production processes implemented Figure 3A - 3C in the model. The user interface of the platform 1 may allow an operator to input the mixture production process to be used.

[0266] The module 14 includes a sample and ratio determination module 140. This module is designed to determine a set of complex microbial community (CMC) samples in the initial sample collection library and their mixing ratios, and generate a mixed result product from the CMC samples using the mixture production process configured according to the mixing ratios. The determination is based on the target mixture characteristic spectrum representing the target mixed result product.

[0267] The module 14 also includes a product generator control module 141. This module is designed to control the product generator module 15, which actually generates a mixture or CCMC result product using the mixture production process according to the determined samples and mixing ratios.

[0268] The sample and ratio determination module 140 is designed to estimate or determine the production parameters of the mixture production process so that the characteristic spectrum of the final CCMC result product is as close as possible to the defined target mixture characteristic spectrum in terms of defined analytical characteristics (such as taxonomic level or functional characteristics).

[0269] As Figure 3A - 3C shown in the example, the production parameters to be determined include: the initial sample set Γ = {γ1…γ k}, the mixing ratios M = [α1...α k at each pooling stage (preferably the same ratio is used for all pooling stages) and the mixing ratios B = [β1...β j at each co-culture stage (if any) (if there are multiple co-culture stages, preferably the same ratio is used for each co-culture stage - Figure 3C ). These data need to satisfy the following constraints: γ i is an integer and satisfies where N is the total number of available initial samples 100; and

[0270] During operation, Σα i and Σβ i equal 1. However, as described below, during candidate calculation, these sums may deviate from the above conditions. In this case, optional constraints can be provided to add a penalty term to the fitness score calculation and reduce the score value. For this purpose, if these sums (Σα i or Σβ i) If any of them is less than 1, the corresponding penalty term needs to be calculated. For example, penalty term = 0.001 * Σα i (or Σβ i ). Similarly, if any of these sums (Σα i or Σβ i ) is strictly greater than 1, the corresponding penalty term needs to be calculated. For example, penalty term = 0.001 * (2 - Σα i )(or 2 - Σβ i ). As a hyperparameter of the production parameter determination model, this penalty term can be adjusted. For example, the change of this hyperparameter relative to the MSE used (for example, a similar method as shown in the previous Figure 1a ) can be studied to adjust its value.

[0271] Hereinafter, a set of possible production parameters is called a "candidate", denoted as c i . Therefore, a candidate is a gene array, where a gene consists of a sample identifier and a mixing ratio: for example, c i = [γ k | [α k | [β k . γ k can be encoded as an integer, while α k and β k can be encoded as floating-point numbers.

[0272] Therefore, the sample and ratio determination module 140 aims at a high-quality candidate (given a target), or in some embodiments, aims at the "best" candidate. Optionally, the module 140 can obtain multiple candidates that meet the requirements and can control the product generator module 15. In fact, according to the quantity of materials (samples 100) available in the initial sample collection library 10, it may be worth trying this solution, that is, driving the product generator module 15 containing multiple groups of different samples to obtain a sufficient amount of final mixture / CCMC result products, and these products all have very similar characteristic spectra (close to the target characteristic spectrum).

[0273] Preferably, an evolutionary algorithm (such as a genetic algorithm) is used to implement the sample and ratio determination module 140. In practical applications, the module 140 first obtains or constructs an initial candidate population, and each candidate represents a set of CMC samples and their mixing ratios in the initial sample collection library. The initial population has a predefined size SIZ (i.e., the number of candidates). Therefore, the module 140 can access the initial sample collection library 10, where each sample has a characteristic spectrum. Each sample is uniquely identified by the exponent γ i . The gene array of each candidate corresponds to the "chromosome" in the theoretical definition of the evolutionary algorithm, and each data (γ i , α i , βi ) can be regarded as a gene.

[0274] In some embodiments, the initial population is constructed randomly, which means randomly selecting SIZ groups of samples (each group having n min -n max samples) within the range of [1…N], and also randomly selecting the relevant mixing ratios α i , β i .

[0275] Then, module 140 iteratively modifies the candidate population using an evolutionary algorithm based on the mixture production process model and the target mixture characteristic spectrum. Each candidate can reproduce with another candidate to form offspring candidates for the next generation population, and these offspring candidates may mutate or adapt. After generating each new population, the fitness of each individual candidate in the new population is evaluated based on a fitness score reflecting the proximity of the target mixture characteristic spectrum.

[0276] Finally, module 140 selects the candidates from the modified population generated by the evolutionary algorithm.

[0277] Figure 4 Shows a genetic algorithm designed specifically for candidates given a target mixture characteristic spectrum.

[0278] The algorithm settings include:

[0279] - The initial population constructed as described above. The population size SIZ can be predefined, such as 100, 200, 500,

[0280] - The individual characteristic spectra 110 of the samples {a i (j)} j ,

[0281] - Optionally, the correspondence matrix can map the analysis feature j to a higher classification level. Hereinafter, when the latter (i.e., the target characteristic spectrum) defines the conditions of a higher classification level, this matrix is used to evaluate the distance between the predicted mixture result characteristic spectrum and the target characteristic spectrum.

[0282] - The mixture production process model, i.e., the CMC interaction matrix or matrix W, and the CCMC interaction matrix or matrix Y (if any),

[0283] - The objective function to be minimized by the algorithm. The mean square error (MSE) can be used,

[0284] - The target characteristic spectrum,

[0285] - Algorithm constraints, such as

[0286] * The minimum and maximum number n of initial samples that make up each candidate min , n max . For example, the minimum and maximum numbers can be 2 and 100 respectively. Of course, other minimum numbers (such as 3, 4, 5, up to 10) and maximum numbers (such as 10, 20, 30, 50 or more) can also be used;

[0287] * The algorithm termination criterion (such as the maximum number of algorithm iterations or the fitness score threshold). The maximum number of iterations can be 100, 500, 1000 or 1500;

[0288] * The mutation probability, used to set the probability of any gene (γ i , α i , β i ) in any candidate of any individual solution undergoing an adaptive change, such as the probability of being replaced by a random value. By default, this probability can be set to 0.3, that is, 30%;

[0289] * The crossover probability, used to set the ratio of the next-generation population generated through crossover operations in the selected candidates, that is, by mixing the genes of two selected candidates. By default, this probability can be set to 0.5, that is, 50%,

[0290] * The parent ratio, used to set the number of candidates selected to participate in reproduction to construct the next-generation population). By default, this ratio can be set to 0.3, that is, 30%,

[0291] * The crossover type, used to set the way of gene exchange. For example, "uniform" crossover can be adopted, that is, randomly select the genes to be exchanged between paired candidates. Specifically, each gene of the new offspring candidate can be selected from either parent candidate based on equal probability,

[0292] * The optional elite ratio, used to set the number of candidates retained in the next-generation population (usually the optimal candidates under the given fitness score below). By default, this ratio can be set to 0.01, that is, 1%.

[0293] As Figure 4 shown, this algorithm includes five steps, which are iteratively repeated based on the current population (initially the initial population) to generate the next-generation population.

[0294] First, based on the mixture production process model and the target mixture characteristic spectrum, for each candidate c in the current population i perform a fitness score SCO(c i ) evaluation (step 400).

[0295] To this end, retrieve candidate c iThe sample characteristic spectrum 110 defined in k}, {β k} is input into the model 13, and the predicted mixture characteristic spectrum PRED(c i ) is obtained.

[0296] For illustrative purposes, for the second mixture production process ( Figure 2B and 3B ), starting from the set of characteristic spectra P corresponding to the selected samples of the variable Γ = {γ k}, the following calculations are performed:

[0297] - First mixing: I = P.A, using the mixing ratio M = {α k}, and then R’ = I.W,

[0298] - Co-culture: R2 = Q1|Q2.Y, where Q1 represents the j-th repeated iteration of R’, and Q2 is used to characterize j bioreactors,

[0299] - Final mixing: I = R2.A, using the mixing ratio B = {β k}, and then I.W is performed, which gives the predicted mixture characteristic spectrum PRED(c i ) of the candidate c i .

[0300] By comparing PRED(c i ) with the target characteristic spectrum TARG, the fitness score SCO(c i ) is calculated. The above penalty value can be added to the score initially evaluated using the evaluation function.

[0301] In some embodiments, PRED(c i ) is fully retained to calculate the fitness score. In other embodiments, if TARG is defined only for certain analysis features (such as certain taxa), only the corresponding subset (corresponding to the taxa) of PRED(c i ) is used to calculate the fitness score.

[0302] For example, MSE can be used as the distance between PRED(c i ) and TARG: ‖PRED(ci)-TARG‖ 2 . Of course, any other distance metric method can also be used, such as the norms L1,…,Lp, SSD, Beta diversity index, or any other known distance metric between analysis features (such as Bray-Curtis, Jaccard similarity, unifrac distance or similarity metric).

[0303] In some embodiments, compared with PRED(ci ) Compared with the taxonomic ranks (such as genus or phylum level) used in [description], TARG may have analytical features defined at higher taxonomic ranks. In this case, PRED(c i ) After post - processing, it can convert the lower - level features of the same high - taxonomic level into a single value corresponding to the higher level. For example, the relative abundances of each taxon related to a higher taxonomic rank are summed up. This operation can be achieved through the correspondence matrix described above. Next, the above - mentioned MSE can be used to obtain the fitness score of candidate c i .

[0304] When TARG defines a specific target abundance for an analytical feature, it is easy to calculate the MSE. However, in some embodiments, the target value of one or more analytical features may be different from the exact value, for example, in the form of an inequality (< or >) or a numerical range (which can be regarded as a double - inequality).

[0305] In this case, the inequality conditions defined by TARG can be first checked against the corresponding analytical features in PRED(c i ). When an analytical feature does not meet its respective inequality constraint, a fitness score is calculated, which includes the distance between the analytical feature and the inequality threshold. On the other hand, when an analytical feature meets its respective inequality constraint, the corresponding distance in the fitness score is set to 0. This operation is implemented as follows: for the analytical feature with an inequality constraint, set the exact value of TARG and its inequality threshold, and then for each analytical feature that meets the inequality constraint, change PRED(c i ) and set its relative abundance to the corresponding inequality threshold (forcing the distance of this feature to be zero). Then, the MSE between the modified PRED(c i ) and TARG can be calculated to obtain the fitness score.

[0306] For example, the inequality can reflect a target diversity criterion, such as a bacterial diversity criterion.

[0307] "Diversity" or "bacterial diversity" refers to the diversity or variability of a complex microbial community (mixture or sample), for example, measured at the genus, species, gene, function, RNA, or metabolite level. Diversity can be expressed by α - diversity parameters to describe a complex community, such as richness (the number of observed species or genera or genes), Shannon index, Simpson index, and its reciprocal Simpson index; and complex communities can be compared by β - diversity parameters, such as Bray - Curtis index, UniFrac index, and Jaccard index.

[0308] Thus, the inequality can represent the minimum or maximum relative abundance of one or more specific analytical features. For example, for a mixed product, it may be necessary to ensure that the proportion (by mass) of a given genus of bacteria is at least 5% of the total compared to the bacteria (this requirement is also specified in other analytical features). The diversity criteria can also define the range within which the relative abundance of one or more specific analytical features falls. Of course, various diversity criteria can be used in combination: a minimum or maximum relative abundance can be set for one analytical feature, a range can be set for another feature, and a maximum relative abundance can be set for a third feature, and so on.

[0309] Once the SCO(c i ) scores of all candidates c in the current population are known, module 140 will select a subset of the current population in step 410 based on the evaluated scores. i ) score, module 140 will select a subset of the current population in step 410 based on the evaluated scores.

[0310] In some embodiments, the number of selected candidates is predefined. For example, this number corresponds to the parental ratio applied to the population. When a parental ratio of 0.3 is applied to a population of 200 candidates, on average 60 candidates will be selected.

[0311] In some embodiments, the candidate with the highest fitness is selected (here, the candidate with the lowest SCO(c i )).

[0312] In some embodiments, candidates are selected based on probability, where the selection probability of each candidate is based on its fitness score. In fact, the candidate with the highest fitness should have a higher selection probability than candidates with poorer fitness scores.

[0313] After selection, a new population of candidates is generated.

[0314] In some embodiments, an elite group is selected from the population based on the elite ratio, and this elite group can include a corresponding number of candidates with the highest fitness. For example, when the elite ratio is 0.01, in a population of 200 candidates, the 2 candidates with the highest fitness will form the elite group, and thus these candidates will necessarily be selected in step 410. The elite group first forms the new generation population.

[0315] Based on the selected candidates, other candidates are obtained by gene crossover between the genes of the selected candidates and / or gene mutations within the gene array, and these candidates form the new generation population. That is, steps 420, 430, 440. The crossover probability defines how many candidates in the new generation (or next generation) population are constructed by crossover between a pair of candidates, while another part of the new population consists of the unchanged selected candidates.

[0316] Step 420 consists of the following operations: pairing the selected candidates, and then possibly performing a partial "chromosome" (i.e., gene) exchange to generate new candidates for the next generation population.

[0317] For example, as long as the new population has not been fully constructed (i.e., the population size has not reached SIZ), new pairing combinations will be generated, which means that a pair of selected candidates will be obtained.

[0318] In some embodiments, each pairing is randomly completed from all the selected candidates. In other embodiments, a decay probability is assigned to each paired selected candidate, that is, the probability that the candidate is randomly selected in the new pairing combination will decrease.

[0319] Then, this pair of candidates can exchange their partial chromosomes based on the crossover probability. If the crossover operation is performed on this pairing, its genes (γ i , α i , β i ) will be mixed, thus generating one or two new candidates to be added to the new generation population. This is the crossover operation performed on the candidates. If the crossover operation is not performed on this pairing, the selected one or two candidates will be retained as new candidates in the new generation population.

[0320] In an embodiment, genes are randomly selected from the gene array and gene exchange is performed between the two candidates in the pairing. In this case, each pairing will reproduce to generate two new candidates. In a variant, each gene is randomly selected from either candidate in the pairing with equal probability for the construction of new candidates. In this case, each pairing will only reproduce to generate one new candidate.

[0321] Then a new set of SIZ candidates is generated to form the new population.

[0322] In step 430, these SIZ candidates will undergo gene mutations within the gene array (chromosome). In the gene array of the candidates, the decision to perform the mutation operation (e.g., performing a bit flip in gene γ i , α i , β i and assigning a new random value to the gene) depends on the mutation probability. When it is determined that a certain candidate needs to undergo gene mutation, the gene involved in the mutation can be randomly selected. The exponent γ i can be modified to a new integer randomly selected from [1…N]. The mixing ratio α i , β i can be modified to a new floating-point value randomly selected from [0…1].

[0323] Mutation helps to maintain the diversity within the population and prevent premature convergence.

[0324] In step 440, a new population is constructed from the SIZ candidates generated by step 430 (regardless of whether these candidates have mutated). By looping back to step 400, the new population can be used for the next iteration of the algorithm or, given the target mixture characteristic spectrum, used as the final population with the aim of selecting (step 460) one or more candidates as the determined solution.

[0325] Therefore, it can be determined in step 450 whether the algorithm termination criterion is met. The aim is to terminate the determination process when the population converges.

[0326] In some embodiments, the algorithm will terminate after reaching a predefined number of iterations (e.g., 1000 iterations).

[0327] In other embodiments, when the fitness of candidate c i in the population, a group of candidates, or a certain proportion of candidates is good enough, the algorithm will terminate. "Good enough fitness" means that the fitness score of the candidate exceeds a predefined fitness score threshold (e.g., when the fitness score is the MSE distance defined above, the fitness score is below a certain threshold).

[0328] Generally, when the number of generations of evolution has reached the maximum value, or the population has reached a satisfactory fitness level, the algorithm will terminate.

[0329] Once the iteration terminates, for example, based on the target mixture characteristic spectrum, candidates are selected (step 460) from the population generated by the evolutionary algorithm. In an embodiment, the optimal candidate (i.e., the candidate with the highest fitness, having the lowest SCO(c i )) in the above example) is selected as the candidate solution for the determination process. Therefore, this candidate provides the samples {γ i} to be used and the mixing ratios {α i}, {β i} required for the mixture production process, generating a mixed result product close to the target mixture characteristic spectrum.

[0330] Alternatively, a candidate can also be randomly selected from the candidates with the highest fitness (e.g., the top 10 candidates with the highest fitness).

[0331] In some embodiments, multiple candidates (e.g., two or three) are selected as the optimal candidates (either by selecting the candidates with the highest fitness or by selecting from a group of candidates with the highest fitness). This method is particularly applicable to scenarios where it is required to generate mixture / CCMC result products from multiple sets of initial samples. Preferably, multiple candidates are selected to ensure that the candidates do not share the same initial samples (i.e., having the same index γ i)。 Its purpose is to ensure that the same materials (initial samples) are not consumed during the production of the mixed result product. In this way, a larger quantity of the product can be produced, and its mixture characteristic spectrum is close to the target mixture characteristic spectrum.

[0332] Therefore, module 141 can utilize the candidate solution c(i) obtained by module 140 to control the actual production process of the CCMC result product 19. For example, module 141 can be used to control the product generator 15 through signaling S1 and optionally through signaling S2, and the generator implements a mixture production pipeline, such as Figure 2A , 2B or as shown in 2C. This means that the method for producing a complex microbial community (CMC) product can first include the following steps: based on the target mixture characteristic spectrum, using the above determination method, obtaining candidates and mixing ratios representing a set of CMC samples in the initial sample collection library.

[0333] For ease of explanation, the following description will focus on a single candidate for producing the target mixed result product. Similar processes can be performed sequentially or in parallel on multiple candidate solutions.

[0334] Once the candidate is known, the production process of the mixed result product 19 can be initiated.

[0335] In an embodiment, module 141 sends a signal to product generator 15 using S1, and at the same time module 141 also carries various samples {γ i} and mixing ratios {α i}, {β i}. In a variant, the signal S1 can be an operator display interface: for example, displaying the samples {γ i} and mixing ratios {α i}, {β i} to the operator on a screen, and the operator manually performs the actual sample picking and mixture production process.

[0336] For ease of explanation, the following will be described based on the second mixture production process, which includes pooling and co-culturing steps. Those skilled in the art will directly adapt this description to the above first and third mixture production processes.

[0337] The product generator 15 can be a machine with mechanical access to the initial sample collection library 10 of the samples (for example, through a controllable articulated robotic arm), including a pooling bioreactor for mixing complex microbial communities, and j bioreactors for co-culturing according to different growth conditions.

[0338] In response to signal S1, given the known amplification factor for each co-culture iteration, the product generator 15 picks (i.e., transfers or obtains) samples {γi}, given the respective mixing ratios {α of the given samples i} and the total volume or mass of the mixed product 19, a certain amount of samples are obtained.

[0339] When the quantity of one of the samples is insufficient, the product generator 15 can send back an error message indicating the sample is missing to the module 141. In this case, the module 141 can restart the mixture production process using another candidate (e.g., the next candidate given a certain fitness score), or can restart Figure 4 the determination process shown, obtain such another candidate, and then trigger the mixture production process again.

[0340] All the obtained samples are injected into the pooled bioreactor for actual mixing.

[0341] Preferably, within 10 minutes to 3 hours, preferably between 30 minutes and 1.5 hours, the samples are homogenized. The homogenization process will be carried out within the temperature range of 0°C to 10°C, preferably within the range of 2°C to 8°C, and more preferably at about 4°C. Then, the resulting CMC product can be maintained for several hours after mixing, can be maintained for at least 16 hours after mixing, and preferably can be maintained for 24 hours after mixing, and is considered stable.

[0342] The CMC product can be cryopreserved before the start of co-culture.

[0343] Then, the product generator 15 sends the CMC product to each of the j bioreactors in equal proportion.

[0344] The CCMC product is obtained through co-culture operation of at least 2 (preferably 3) bioreactors, where the bioreactors have at least one different parameter selected from the group consisting of pH value, temperature, pressure, culture time, residence time, aeration condition, redox potential, culture medium, light source, and combinations thereof.

[0345] Then j CCMC products with different characteristic spectra are obtained.

[0346] Next, given the respective mixing ratios {β of the given CCMC products i} and the total volume or mass of the mixed product 19, the product generator 15 obtains a certain amount of each CCMC product. The obtained CCMC products are put into the pooled bioreactor (optionally adding a certain amount of new samples from the initial collection library), and actual mixing is carried out in the pooled bioreactor using the above parameters. Then, a mixed / CCMC result product is obtained, and its mixture composition is close to the target mixture characteristic spectrum.

[0347] As described above, some of the samples 100y - 100z may be virtual samples. When there is a virtual sample in the candidate solution sample set {γ i}, the sample needs to be actually produced according to its virtual definition (i.e., the corresponding individual characteristic spectrum).

[0348] When the module 141 detects a virtual sample 100y - 100z corresponding to the bacterial composition, the module 141 will send a signal to the sample generator 16 using S2, requesting it to produce the artificial synthetic sample. S2 can identify the target sample and label the required amount of materials.

[0349] The sample generator 16 can mechanically access the isolated strain library 160 (e.g., through a controllable articulated robotic arm) and perform storage access to the file 161, which defines the sample composition in terms of the mixing of individual strains. The sample generator 16 also includes a bioreactor for performing strain mixing.

[0350] In response to the signal S2, given the calibrated material requirements, the sample generator 16 retrieves the definition of the artificial sample (bacterial community) according to the strains, and obtains an appropriate amount of each required strain from the strain library 16. The obtained required strains are put into the bioreactor for actual mixing within 30 minutes at 4°C.

[0351] In an embodiment, the sample generator 16 can access the library 10 and / or an external sample library 99. When the module 141 detects a virtual sample corresponding to an engineered or processed complex community (i.e., a mixture containing samples), the module 141 sends a signal to the sample generator 16 using S2, requesting it to produce the engineered or processed sample. S2 can identify each strain and / or each sample in the library 10 and / or each external sample involved in the mixing operation, and label the required amount of materials.

[0352] In response to the signal S2, the sample generator 16 aspirates or picks up the materials and puts the materials into the bioreactor for actual mixing.

[0353] Once the mixing is completed and stable, i.e., the sample has been generated, it is subsequently stored in the initial sample collection library or the library 10, and the product generator 15 can obtain this sample to actually produce the mixed result product 19.

[0354] Although the above signals S1 and S2 are described as control signals for driving the product generator 15 and the sample generator 16, one or both of these two signals may only be signals shown to the operator for the operator to actually manually perform the mixing operation.

[0355] Samples may disappear over time (either due to the actual production of certain products or due to degradation over time), and new samples can be collected from new donors. The results show that after determining the production target mixing result product for the mixture definition, the collection library 10 may evolve over time (thus causing matrix A to evolve). Thanks to the present invention, the pool predictor 130 can be reconfigured based on the evolved collection library (redefining A, re - learning W) to determine a new candidate solution given the target mixture signature spectrum ( Figure 4 ).

[0356] Figure 5 Figure 500 shows a computer device 500 for managing the production platform 1 in a schematic way. For example, the computer device 500 can implement the predictor module 13 and the test and decision module 14, and can control the sequencer 12, the product generator 15, and the sample generator 16 through adaptation signaling (S1 and S2).

[0357] The computer device 500 can implement at least one embodiment of the present invention. The computer device 500 is preferably a device such as a microcomputer, a workstation, or a lightweight portable device. The computer device 500 includes a communication bus 501, and the following components are preferably connected to this bus:

[0358] - A central processing unit 502, such as a microprocessor, labeled as CPU;

[0359] - A read - only memory 503, labeled as ROM, for storing the computer program for implementing the present invention;

[0360] - A random - access memory 504, labeled as RAM, for storing the executable code of the method according to the embodiments of the present invention, and registers for adjusting and recording the variables and parameters required to implement the method according to the embodiments of the present invention;

[0361] - A communication interface 505, connected to the network 599 for communicating with user or operator devices, and / or for communicating with other devices of the platform 1 (such as the sequencer 12, the product generator 15, and the sample generator 16);

[0362] - A data storage device 506, such as a hard disk or a flash memory, for storing the computer program for implementing the method according to one or more embodiments of the present invention, and any data required for the embodiments of the present invention, especially including individual sample signatures (i.e., the collection library 11).

[0363] Optionally, the computer device 500 may further include a display screen 507 as a graphical interface for the operator, for example, for platform configuration and / or information display (such as candidate solutions) through a keyboard 508 or other pointing input devices (such as setting the target mixture signature spectrum).

[0364] The computer device 500 can optionally be connected to various peripheral devices (not related to the present invention), such as a sequencer 12, etc., and each device is connected to an input / output interface card (not shown).

[0365] Preferably, the communication bus can provide communication and interoperability between various components included in the computer device 500 or peripheral devices connected thereto. The representation of the bus is not restrictive. In particular, the central processor can directly or through another component of the computer device 500 send instructions to any component of the computer device 500.

[0366] The executable code can optionally be stored in any of the following media: read-only memory 503, hard disk 506, or removable digital media (not shown). According to an optional variant, before execution, the executable code of the program can be received via the communication network 599 through the interface 505 and stored in one of the storage media in the computer device 500 storage media (such as the hard disk 506).

[0367] The central processor 502 is preferably adjusted to control and direct the execution of instructions or software code segments of the program according to the present invention, and the instructions are stored in one of the storage media described above. At startup, the program stored in the non-volatile memory (such as the hard disk 506 or read-only memory 503) is transferred to the random access memory 504 and registers for storing variables and parameters required to implement the present invention, and then the random access memory 504 will contain the executable code of the program.

[0368] Experimental results

[0369] Experimental scope

[0370] This experiment aims to study the efficiency of the evolutionary algorithm in the following tasks, namely determining the initial samples and their mixing ratios and obtaining a mixed resultant product with the characteristic spectrum of the target mixture (Experiment 4). Previously, the relevance of the method based on the learning interaction matrix was studied experimentally, and the following experimental modeling and prediction were carried out: the mixture characteristic spectrum of the microbiota sample mixture (Experiments 1 and 2), and the co-culture product characteristic spectrum of the co-culture product obtained in a bioreactor with specific growth conditions (Experiment 3).

[0371] Experiment 1 - Scheme

[0372] Consider the initial sample collection library 10. Using 16S-based microbiota taxonomic analysis, each microbiota sample is sequenced to obtain the corresponding initial characteristic spectrum collection library 11. Thus, 131 taxa (at the genus level) are evaluated as the analysis characteristics.

[0373] Next, sample mixing is achieved. Each mixed product is composed of 3 - 6 samples combined at their respective ratios. The mixing is carried out at a temperature of 4 °C. After completion of the mixing, homogenization is performed within 30 minutes - 1.5 hours. Using the same 16S - based microbiota taxonomic analysis, the mixed products are sequenced under steady - state conditions (i.e., within 16 hours from mixing after homogenization).

[0374] Adopt a k - fold cross - validation strategy (k = 5) to configure the pool predictor 13, that is, to learn λ and the interaction matrix W. The k - fold strategy ensures that no observed data is used as both training data and test set in the same evaluation.

[0375] This modeling method was tested and applied at four different taxonomic ranks (species, genus, family, order). However, due to the very high sparsity of the species - level dataset, it was excluded from the test procedure. Starting from the genus level, the data abundance of the classification assignment table is sufficient for analysis. Therefore, the lower decomposition levels (family, order) are only derived from the genus - level table, and their use is limited to visualization when needed, rather than for the modeling procedure. The main reason is that the composition of higher decomposition levels cannot be derived from the taxonomic levels used in training. From an application perspective, mastering genus - level information is of great value.

[0376] We trained the model for native samples ( Figure 5 ), co - culture samples (Figure 6), and combined samples of both (Figure 7) respectively. MSE was used to quantify the quality when the modeling was applied to the data. At the same time, the MSE of the machine - learning model was systematically compared with that of the linear model (providing naive predictions).

[0377] Experiment 1 - Results

[0378] Figure 5a Shows the initial feature spectrum acquisition library 11 corresponding to the initial sample acquisition library 10, and the initial sample acquisition library 10 only contains native fecal microbiota samples. 27 microbiota samples were considered. Their individual feature spectra are shown as follows.

[0379] Figure 5b Shows the mixture feature spectra of 24 mixed products, which are composed of Figure 5a 3 - 6 microbiota samples out of 27 microbiota samples at their respective ratios or proportions. The saved mixture is defined as {p x (k)} k .

[0380] Figure 5c On the left, it shows that given the mixture definition {p x (k)} k and the individual sample feature spectra {a x (j)} jIn the case of the error caused by the linear prediction of the mixture characteristic spectrum. The linear prediction corresponds to step I = A*P.

[0381] The right side of the figure also shows the error caused by the POOL prediction method (i.e., introducing the product interaction matrix W). W only uses Figure 5a and 5b (native samples) The sample and mixture characteristic spectra are obtained by machine learning, and a k-fold cross-validation strategy is adopted.

[0382] The performance of the model-based method is better than that of the linear method in terms of the native dataset.

[0383] Figure 6a The initial characteristic spectrum acquisition library 11 corresponding to the initial sample acquisition library 10 is shown. The initial sample acquisition library 10 only contains co-cultured fecal microbiota samples. 36 microbiota samples are considered. Their individual characteristic spectra are as shown in the figure.

[0384] Figure 6b The mixture characteristic spectra of 48 mixed products are shown, and these mixed products are composed of Figure 6a 3 - 6 microbiota samples from the 36 microbiota samples mixed at their respective ratios or proportions. The mixture definition {p x (k)} k has been saved.

[0385] Figure 6c The left side shows the error caused by the linear prediction of the mixture characteristic spectrum in the case of a given mixture definition {p x (k)} k and individual sample characteristic spectra {a x (j)} j In the case of, the error caused by the linear prediction of the mixture characteristic spectrum. The linear prediction corresponds to step I = A*P.

[0386] The right side of the figure also shows the error caused by the POOL prediction method (i.e., introducing the mixture interaction matrix W). W only uses Figure 6a and 6b (co-cultured samples) The sample and mixture characteristic spectra are obtained by machine learning, and a k-fold cross-validation strategy is adopted.

[0387] The performance of the model-based method is significantly better than that of the linear method in the co-cultured dataset (the median MSE predicted by the ML model is reduced to 1 / 5 of the original).

[0388] In Figure 7a and 7b the interaction matrix W is obtained by machine learning, and Figure 5a , 5bThe sample characteristic spectra of 6a and 6b (i.e., the native sample and the co-culture sample) and the mixture characteristic spectra are used as training data. The k-fold cross-validation strategy is adopted again.

[0389] Figure 7a shows Figure 5a and 5b (i.e., the native sample and its mixture) when the dataset is applied to the pool predictor 13 configured as such.

[0390] On the left side of the figure, it shows the error caused by the linear prediction of the mixture characteristic spectrum under the given mixture definition {p x (k)} k and the individual sample characteristic spectrum {a x (j)} j .

[0391] On the right side, it shows the error caused by the POOL prediction, that is, involving the interaction matrix W learned as such.

[0392] Compared with the single-dataset model, the combined-dataset model has a slightly improved estimation accuracy when applied to the native dataset.

[0393] Figure 7b shows Figure 6a and 6b (i.e., the co-culture sample and its mixture) when the dataset is applied to the pool predictor 13 configured as such.

[0394] On the left side of the figure, it shows the error caused by the linear prediction of the mixture characteristic spectrum under the given mixture definition {p x (k)} k and the individual sample characteristic spectrum {a x (j)} j .

[0395] On the right side, it shows the error caused by the POOL prediction, that is, involving the interaction matrix W learned as such.

[0396] Compared with the single-dataset model, the combined-dataset model has a significantly improved estimation accuracy when applied to the co-culture dataset (the median MSE predicted by the ML model is reduced to 1 / 4 of the original).

[0397] Experiment 1 - Discussion and Conclusion

[0398] In all cases, model-based prediction improves the accuracy of naive (linear) estimation. This advantage is particularly important in the co-culture dataset, especially for certain taxonomic groups for which the naive method performs poorly in such datasets. The model-based correction method is more efficient, probably because there is more room for improvement with this method. If more data is added for model training, it can be assumed that the overall performance and robustness will be enhanced. The training method allows such models to evolve.

[0399] Experiment 2 - Protocol

[0400] In this experiment, NGS shotgun sequencing was used to analyze Sample 100. Metagenomic sequencing data for 76 pools and 69 donor samples (or individual co-cultures) were obtained.

[0401] Due to the large number of features in NGS shotgun analysis (compared to 16S sequencing, especially when considering species level (rather than genus level), or for certain functions), PCA was used to reduce the dimensionality of each sample feature profile to k PCs.

[0402] Figure 8 The PCA results of the genus relative abundances obtained from NGS shotgun sequencing of native samples (native: sample or mixture) and co-culture samples (co-culture: sample or mixture) are shown. Co-culture samples and native samples show clustering trends respectively.

[0403] Figure 9 The PCA-based strategy of this experiment is summarized. It is clear that in Experiment 1, the "taxon × taxon" interaction matrix W was learned, while in Experiment 2, it was changed to learning the "top k principal components × taxon" interaction matrix W.

[0404] The method for learning the interaction matrix W is the same as that for the 16S analysis in Experiment 1.

[0405] Experiment 2 - Results

[0406]

[0407] Table 1: Comparison of the results of linear prediction, taxon prediction without using PCA, and taxon prediction using PCA

[0408] The interaction matrix W was learned using MSE. In addition, the predicted mixing results (using W) were compared with the true mixing results based on MSE or Bray Curtis distance.

[0409] Both modeling methods (with or without PCA) improved the prediction accuracy of taxonomic feature spectra at the genus and species levels (according to the MSE or BC metrics). The correction based on matrix W had a more significant impact when predicting mixtures from co-culture samples compared to the native samples.

[0410] Reducing the analyzed features by PCA seemed to significantly improve the prediction accuracy for co-culture samples, while only slightly improving the prediction accuracy for native samples.

[0411] Experiment 3 - Protocol

[0412] Three co-culture procedures were adopted in this experiment to generate three datasets, named EXP1, EXP2, and EPX3 respectively.

[0413] For these three procedures, three co-culture bioreactors were used, which could provide three different environmental conditions (only different in pH value) simultaneously:

[0414] Fermenter or bioreactor 1: pH = 5.3,

[0415] Fermenter or bioreactor 2: pH = 6.3,

[0416] Fermenter or bioreactor 3: pH = 7.3.

[0417] Each time, a single microbial community (starting CMC product) was cultured in the three bioreactors, and the specific operations were as follows:

[0418] In EXP1, the starting CMC product grew in the three bioreactors for 14 days, and samples were taken on the 4th day. Finally, three CCMC products were obtained.

[0419] In addition, two other starting CMC products from different donors grew in the same bioreactors for 14 days.

[0420] Therefore, nine CCMC products sampled on the 4th day were obtained through EXP1.

[0421] In EXP2, the starting CMC product grew in the three bioreactors in parallel for 3 days. The resulting CCMC products were frozen and then thawed, and then input into their respective reactors for another 3 days of growth. The resulting CCMC products (frozen and thawed) were input into their respective reactors again for a third 3-day growth. Regular sampling was carried out.

[0422] Therefore, nine different growth processes were carried out in EXP2, resulting in three final CCMC products and six intermediate CCMC products (a total of nine CCMC products).

[0423] In EXP3, the starting CMC products were grown in parallel in three bioreactors for 3 days. The resulting CCMC products were mixed together and then fed into three reactors for another 3 days of growth. The resulting CCMC products were mixed together again and then fed into three reactors for a third 3-day growth. Regular sampling was carried out.

[0424] Therefore, nine different growth processes were carried out in EXP3, resulting in three final CCMC products and six intermediate CCMC products (a total of nine CCMC products).

[0425] Through three experiments (each experiment involving three co-culture bioreactors), a total of 27 CCMC products (grown for 3 or 4 days) were produced.

[0426] Using 16S-based microbiota taxon analysis, the characteristic spectra of the starting CMC products and these 27 CCMC products were obtained by sequencing.

[0427] These characteristic spectra were used to learn the fermenter predictor model. A total of six models were used. The first model is as described above (refer to Figure 3B ): R2 = Q * Y, where Q = Q1|Q2 is used to define the growth conditions (here referring to the pH value), and Y is the co-culture interaction matrix to be learned.

[0428] Other models include:

[0429] Model 2 defined by (Q1 * Y1)+(Q2 * Y2), where Y1 and Y2 must be learned;

[0430] Model 3 defined by (Q2 * Y2)·Q1, where Y2 must be learned, and · represents the Hadamard product, i.e., element-wise multiplication;

[0431] Model 4 defined by (Q1 * Y1)+(J * Y’2), where Y1 and Y’2 must be learned, and J represents an n×1 matrix of all 1s (where n is the number of rows of Q1, i.e., the number of starting CMC products);

[0432] Model 5 defined by ((Q2 * Y2)·Q1)*W), where Y2 must be learned, and W represents the matrix of the pool predictor;

[0433] Model 6 defined by ((Q2 * Y2)·Q1)*Y1), where Y1 and Y2 must be learned.

[0434] Refer to the method of the above-mentioned pool predictor (Experiments 1 and 2), and use three datasets to complete the learning of each model. In particular, a k-fold cross-validation strategy is adopted. The mean squared error (MSE) is used to evaluate the distance between the predicted CCMC feature spectrum and the true CCMC feature spectrum. The calculated distance between the baseline predicted CCMC feature spectrum (naive prediction) and the true CCMC feature spectrum is compared.

[0435] For each taxon output by each bioreactor, the first naive prediction NAIV1 used takes the leave-one-out mean of the relative abundances of all experiments. In fact, after culturing, the classified feature spectrum measured for each sample is compared with the average feature spectrum of all other samples cultured in the same bioreactor. Figure 10a Shows the relative abundances (x-axis) of each sample and each genus (different colors) according to the actual measurement and naive prediction of this leave-one-out mean. Since 3 different bioreactors were tested, there are 3 fitted lines corresponding to each genus of bacteria. When the true relative abundance is more important, the more values we exclude in the leave-one-out mean calculation, so these fitted lines show a slight slope. The simple mean will make the prediction results of each genus / bioreactor combination show a horizontal line.

[0436] The second naive prediction NAIV2 adopted retains the same composition as the starting CMC product, that is, it is assumed that the relative abundances and (more generally) the classified feature spectra are not changed during the culturing process (so this prediction method has 100% absolute naivety). Figure 10b Indicates that this hypothesis does not hold because there are significant deviations between the data points and the y = x baseline (corresponding to perfect prediction).

[0437] Experiment 3 - Results

[0438] Model MSE NAIV2 12.87e-4 NAIV1 5.85e-4 Model 1 5.27e-4 Model 2 5.42e-4 Model 3 10.13e-4 Model 4 5.47e-4 Model 5 9.60e-4 Model 6 6.41e-4

[0439] Table 2: Comparison of prediction results between baseline prediction and model-based prediction

[0440] Experiment 3 - Discussion and Conclusions

[0441] Table 2 confirms that prediction results better than the baseline predictions NAIV1 and NAIV2 can be obtained. The study further shows that Model 1 exhibits the optimal prediction performance on the datasets used.

[0442] It should be noted that the datasets used for model training and evaluation have the disadvantages of small sample size and poor diversity. Compared with the baseline prediction, it is expected that using datasets with more sufficient sample size and better starting CCMC diversity will increase the application benefits of Model 1.

[0443] Experiment 4 - Scheme

[0444] This experiment studied the efficiency of evolutionary algorithms in the following tasks: determining the initial samples and their mixing ratios, and obtaining a mixed product with the characteristic spectrum of the target mixture.

[0445] By differentiating the definition methods of the characteristic spectra of the target mixtures, various sub-experiments were carried out. The initial sample collection library 10 used contained 63 samples analyzed by the 16S sequencing method. Therefore, the relative abundance matrix table of all samples at the genus level has been obtained in this experiment.

[0446] In EXP4.1, at the genus level, a collection library of 100 target characteristic spectra with k = 8 (number of mixed samples) was randomly generated.

[0447] The following steps were repeated 100 times:

[0448] 1 - Randomly generate n α ratios under the following constraints:

[0449] a /

[0450] b / All ratio values satisfy the non-zero constraint condition

[0451] 2 - Use the pool predictor 130 to generate the CMC characteristic spectrum (post-processing sets negative values to 0 and normalizes the sum of relative abundances to 1).

[0452] 3 - Randomly generate a B matrix containing 3 values (3 bioreactors),

[0453] 4 - Obtain the CCMC characteristic spectrum through the fermenter predictor 131, and calculate the final mixed product characteristic spectrum using the pool predictor 130 when the calculated B matrix and CMC characteristic spectrum are input.

[0454] 5 - Store the calculated final mixed product characteristic spectrum, which will be used as the standard target product characteristic spectrum later. Thus, 100 diverse target product characteristic spectra conforming to the actual production process data were generated.

[0455] After that, for each of the above 100 targets as input, the above evolutionary algorithm was executed with the following parameters:

[0456] - Maximum number of iterations = 500

[0457] - Population size = 200

[0458] - Mutation probability = 0.3

[0459] - Elite ratio = 0.01

[0460] - Crossover probability = 0.5

[0461] Then, calculate and store the MSE and BrayCurtis distances between each pair of target queries and the predicted mixture results (computer simulation processing is performed according to the settings shown in Figure 2B ).

[0462] In EXP4.2, the target feature spectrum is not fully defined. The target feature spectrum is defined by seven genera. The sum of their relative abundances is 0.27. The minimization function is only applied to the above seven genera. Other genera are allowed to exist in the final mixture. Their numbers and relative abundances are not limited.

[0463] In EXP4.3, the target feature spectrum is not fully defined. The target feature spectrum is defined by seven genera, three of which are defined as inequality constraints and the remaining four are defined as equality constraints. As described above, the application of the minimization function (MSE) varies depending on the operator of each taxon. Taking the equality constraint as an example, the distance between the target feature spectrum and the relative abundance of this taxon in the predicted feature spectrum needs to be minimized. For the inequality constraint, when the predicted relative abundance does not satisfy the inequality, the above distance needs to be minimized, and when the condition is met, this taxon will no longer be constrained. Other genera are allowed to exist in the final mixture. Their numbers and relative abundances are not limited.

[0464] In EXP4.4, the target feature spectrum is not fully defined, and features at different taxonomic levels are used in this target definition process. Therefore, the target feature spectrum defined in this experiment includes the random complete target feature spectrum (such as the feature spectrum in EXP4.1), and an additional order-level taxonomic constraint is added (the order includes families as taxonomic levels, and the families themselves include genera).

[0465] In EXP4.5, the target feature spectrum is not fully defined, and multi-level taxonomic features and mixed operators are used to construct the target feature spectrum.

[0466] Experiment 4 - Results

[0467] Regarding EXP4.1, Figure 11a shows the distribution of the MSE (left figure) and BrayCurtis (BC) distances (right figure) between each pair of target feature spectra and the corresponding predicted feature spectra after genetic algorithm processing. Therefore, a total of 100 points are co-plotted.

[0468] The MSE and BC distance values are very low. In fact, for the feature spectra obtained from two sequencing of a real sample, the BC distance is rarely lower than 0.20. In EXP4.1, the maximum distance obtained here is lower than 0.09.

[0469] Figure 11bShows the comparison between one target characteristic spectrum (No. 58) among 100 target characteristic spectra and the related predicted characteristic spectra. Each point corresponds to a genus, where the x-value is the expected relative abundance defined in the target, and the y-value is the predicted relative abundance of the optimal candidate obtained through genetic algorithm prediction.

[0470] Since the associated predicted characteristic spectrum of Test 58 has an MSE close to the mean of 100 experiments, Test 58 is selected as an example.

[0471] All points in the figure are densely distributed near the y = x axis, indicating that there is a high similarity between the target characteristic spectrum and the predicted characteristic spectrum for the average result.

[0472] For EXP4.2, Table 3 reports the target relative abundance and predicted relative abundance of each selected genus among the 7 selected genera in the target characteristic spectrum.

[0473]

[0474] Table 3

[0475] The MSE of the optimal candidate obtained in EXP4.2 is 5.84E - 06, and the BC distance is 1.91E - 02. This indicates that even using an incomplete target characteristic spectrum, it is still possible to find relevant candidates.

[0476] For EXP4.3, Table 4 reports the target relative abundance and predicted relative abundance of each selected genus among the 7 selected genera in the target characteristic spectrum.

[0477]

[0478] Table 4

[0479] The MSE of the optimal candidate obtained in EXP4.3 is 5.15E - 08, and the BC distance is 6.02E - 02. This indicates that even using inequality operators and an incomplete target characteristic spectrum, it is still possible to find relevant candidates.

[0480] For EXP4.4, Table 5 reports the target relative abundance and predicted relative abundance defined at multiple taxonomic levels for two taxonomic groups selected from multiple measured taxa.

[0481]

[0482] Table 5

[0483] The MSE of the optimal candidate obtained in EXP4.4 is 5.32E - 06, and the BC distance is 5.95E - 02. This indicates that even using multiple taxonomic group levels and an incomplete target characteristic spectrum, it is still possible to find relevant candidates.

[0484] For EXP4.5, Table 6 reports the target relative abundances and predicted relative abundances for each of the five taxa defined at multiple taxonomic levels.

[0485]

[0486] Table 6

[0487] Although the inequality constraint condition (0.29 ≤ 0.50) is satisfied, the MSE of the optimal candidate obtained in EXP4.5 is 6.8E - 06 and the BC distance is 32.9E - 02 (this value is relatively high). However, compared with the relative abundance value of 0.21 of the predicted feature spectrum, the BC distance is largely affected by the relative abundance value of 0.50 of the target feature spectrum.

[0488] This successful experience shows that it is still possible to find relevant candidates using multiple taxonomic levels and incomplete target feature spectra.

[0489] Experiment 4 - Discussion and Conclusions

[0490] Experiment 4 aimed to evaluate the applicability of a genetic algorithm with established parameters in solving the candidate selection problem. This evaluation used pool and fermenter predictors to generate 100 true feature spectra at the genus level. The results of the main task (EXP4.1) were satisfactory, and all the generated predicted values were highly similar to their respective target values.

[0491] In a real - world scenario, the practice of incompletely defining target values also has practical significance. When features are organized in a hierarchical structure (such as using microbial taxa, or perhaps also functional features such as EC numbers), there may not be enough information to fully list all the necessary features and query through inequalities and different feature hierarchies. Experiments EXP4.2 to EXP4.5 aimed to test each possibility and their combinations, and the results showed that the genetic algorithm can also operate efficiently in the above - mentioned scenario.

[0492] Although Experiment 4 was based on a single computer - simulation evaluation, its successful implementation shows that the genetic algorithm can efficiently find candidates, thus achieving "determining the initial samples and their mixing ratios to obtain a mixed resultant product with the feature spectrum of the target mixture".

[0493] Although the present invention has been described above by reference to specific embodiments, the present invention is not limited to these specific embodiments, and those skilled in the art can make obvious modifications without departing from the scope of the present invention.

[0494] After referring to the above exemplary embodiments, those skilled in the art will envision many further modifications and changes. These embodiments are given only by way of example and are not intended to limit the scope of the invention, which is determined only by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.

[0495] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be advantageously employed.

Claims

1. A computer-aided method for determining a set of complex microbial community (CMC) samples (100) and their mixing ratios (α i , β i ) in an initial sample collection library (10), and using a mixture production process configured with the mixing ratio to produce a mixed resultant product from the CMC samples, the method comprising: Obtain an initial candidate population (c i ), where each candidate represents a set (Γ) of CMC samples and their mixing ratios (M, B) in the initial sample collection library; Based on a model of the mixture production process and a target mixture characteristic spectrum (TARG) representing the target mixed resultant product, an evolutionary algorithm (400 - 450) is applied to iteratively modify a population of candidates. Select (460) a candidate from the modified population generated by the evolutionary algorithm.

2. The method according to claim 1, wherein The mixture production process model includes: Given a mixing ratio, using a linear method to predict the intermediate CMC characteristic spectrum required for the mixture of a selected complex microbial community sample. Using a CMC interaction model learned from the linearly predicted reference CMC characteristic spectrum and the corresponding true reference CMC characteristic spectrum, the intermediate CMC characteristic spectrum is corrected to a predicted CMC characteristic spectrum, which represents the complex microbial community (CMC) product produced by the mixture.

3. The method according to claim 2, wherein, Each candidate includes a mixing ratio representing the respective proportions of the CMC samples to be mixed.

4. The method according to claim 2 or 3, wherein, The mixture production process model further includes one or more cycles of the following steps: Based on the predicted CMC characteristic spectrum or the predicted mixed result characteristic spectrum of the previous cycle, predicting a plurality of CCMC characteristic spectra, which represent co - cultured CCMC products obtained by different co - cultures of the same starting CMC product. Based on the CCMC characteristic spectra, using a linear method to predict an intermediate mixed result characteristic spectrum, which represents a second mixture of the CCMC products at a given mixing ratio. Using a mixture interaction model learned from the linearly predicted reference mixture characteristic spectrum and the corresponding reference true mixed result characteristic spectrum, the intermediate mixed result characteristic spectrum is corrected to a predicted mixed result characteristic spectrum, which represents the mixed CCMC product produced by the second mixture.

5. The method according to claim 4, wherein Each candidate further includes a mixing ratio representing the respective proportions of the CCMC products to be mixed.

6. The method according to claim 4 or 5, wherein Throughout the cycle, the same mixing ratio representing the respective proportions of the CCMC products to be mixed is used.

7. The method according to any one of claims 4-6, wherein The CMC interaction model and the mixture interaction model are the same model.

8. The method according to any one of claims 1-7, wherein, Each candidate is defined by a gene array, which includes a sample identifier and a mixing ratio, and each gene array defines a separate gene.

9. The method according to claim 8, wherein The iteration in the evolutionary algorithm includes: Based on the mixture production process model and the target mixture characteristic spectrum, evaluating the score of each candidate in the current population. Based on the evaluated scores, selecting a portion from the current population. Based on the selected candidates, using gene crossover between the respective genes of the selected candidates and / or gene mutation within the gene array, generating a new population of candidates.

10. The method according to claim 9, wherein, Evaluating the score includes: using the mixture production process model to calculate the distance between the target mixture characteristic spectrum and the predicted mixed result characteristic spectrum according to the candidate.

11. A method for producing a co - cultured complex microbial community (CCMC) resultant product, the method including: Based on the characteristic spectrum of the target mixture, using the determination method according to any one of claims 1-10, a candidate (c i ) is obtained, and the candidate (c i ) represents a set of complex microbial community (CMC) samples (100) and mixing ratios (α i , β i ) in the initial sample collection library (10). Actually picking the CMC samples of the group from an initial sample collection library. Using the mixture production process configured with the obtained mixing ratio to process the picked CMC samples to obtain the CCMC resultant product.

12. The method according to claim 11, wherein The mixture production process includes a first pooling stage, i.e., mixing the selected CMC samples to obtain a CMC product.

13. The method according to claim 12, wherein, The mixing is carried out according to the mixing ratio of the obtained candidates.

14. The method according to claim 12 or 13, wherein The mixture production process further includes a second stage, i.e., one or more iterations of amplifying the starting CMC product, where the iteration includes: (i) co-culturing the CMC product or the mixed result product obtained from the previous iteration in a bioreactor with respective operating parameters to obtain a CCMC product; (ii) mixing the CCMC products to obtain a mixed result product.

15. A computer device, including at least one microprocessor, which can be used to execute the method described in any one of claims 1-14.

16. A non-transitory computer-readable medium for storing a program, when the microprocessor or computer system in the device executes the program, the program can cause the device to execute the method described in any one of claims 1-14.

Citation Information

Patent Citations

  • Method for preparing a fecal microbiota sample

    WO2016170285A1

  • Method for lyophilisation of a sample of fecal microbiota

    WO2017103550A1

  • Stool collection method and sample preparation method for transplanting fecal microbiota

    WO2019171012A1

  • Method of expanding a complex community of microorganisms

    WO2022136694A1