Platform for biotherapeutic drug discovery and method for biotherapeutic drug discovery
Patent Information
- Application Number
- EP2024804477
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-05
- Filing Date
- 2024-11-05
- Publication Date
- 2026-09-09
AI Technical Summary
Current methods for biotherapeutic drug discovery, particularly for microbiome-based therapies, face challenges in isolating and culturing strain-specific microorganisms, optimizing their composition, and scaling up production while maintaining functional integrity.
A comprehensive platform integrating a biobank module, data repository, data engine, biotherapeutic predictor module, and culturomics module to store, process, and predict biotherapeutic drug compositions based on clinical and bioinformatic data, and isolate and culture strain-specific microorganisms using advanced culturomics techniques.
The platform enhances the precision and scalability of biotherapeutic drug development by ensuring strain-specificity, optimizing microbial interactions, and providing a holistic view of microbiome functional capacity, leading to more effective and personalized treatments.
Smart Images

Figure EP2024081247_08052025_PF_FP_ABST
Abstract
Description
[0001] A PLATFORM FOR BIOTHERAPEUTIC DRUG DISCOVERY AND A METHOD FOR BIOTHERAPEUTIC DRUG DISCOVERY
[0002] TECHNICAL FIELD
[0003] The present invention pertains to the development of biotherapeutic drugs, which includes, but is not limited to, proof of concept studies on donor-derived microbiota preparations, biobank sampling, composition prediction, isolation, synthesis, and final composition and formulation of a biotherapeutic drug comprising a plurality of microorganisms.
[0004] BACKGROUND
[0005] Biotherapeutic drugs, often referred to as microbiome-based therapies, comprise a plurality of gut microorganisms. These innovative therapies aim to treat a variety of diseases and conditions by modulating the gut microbiota. Typically, these therapeutic compositions include a diverse consortium of beneficial bacteria naturally found in the human gastrointestinal tract. By restoring or enhancing the balance of these microorganisms, biotherapeutic drugs can significantly contribute to the management and potential cure of conditions such as inflammatory bowel disease (IBD), irritable bowel syndrome (IBS), Clostridium difficile infections, and metabolic disorders like obesity and type 2 diabetes. Emerging research also suggests their potential in treating neurological conditions, such as autism spectrum disorders and depression, by influencing the gut-brain axis. This novel therapeutic strategy leverages the symbiotic relationship between humans and their gut microbiota, offering a promising avenue for precision medicine and personalized healthcare.
[0006] Various proof-of-concept study modalities have been developed, including those on cell cultures, organoids, animal models, humans, and artificial models. However, none have demonstrated as high efficacy and reproducibility as human proof-of-concept studies. In microbiome research, these studies typically involve the use of healthy donor microbiome products, prepared with care to preserve all the microbiota components naturally occurring within the healthy gut, as a full-ecosystem microbiome transfer to the patient’ s gut. This method allows for the observation of whether the medical condition, being a clinical indication, improves or not. Such clinical proof is currently the most efficacious way to infer what is biologically active within the donor-derived microbiome.
[0007] There are various bioinformatic systems available worldwide, which can be categorized into two groups: automated systems for well-established analyses (e g., used in human genomics) and sets of scientist- written scripts prepared for specific analyses, not suitable for large-scale analyses (e.g., used in metagenomic studies). Computational techniques, including bioinformatics, data science, programming, high-throughput computations, mathematics, statistical models, machine learning, artificial intelligence, etc., are used to predict the features responsible for the observed effect. Reliable output data heavily rely on the quality of input data.
[0008] Moreover, isolating, culturing, maintaining, and composing a biosynthetic intestinal microbiome presents a significant challenge. To compose a biotherapeutic drug, it is necessary to employ the most advanced in silico analyses and techniques, such as genomics, culturomics, bioreactor-based cultures, and microfluidics.
[0009] WO2023056341A1 discloses methods and systems for the design, improvement, optimization, or personalization of microbiome therapeutics for various diseases. However, this solution does not provide an effective method for isolating microbial strains that can be used to create a biotherapeutic. Furthermore, it has limited functionalities regarding the selection of microorganisms based on their functional features identified based on multiomics data.
[0010] EP3962653A1 discloses a screening platform utilizing a microfluidic system comprising: at least one droplet inlet for receiving one or more sets of droplets, each set of droplets containing individual droplets, each individual droplet containing a single type of microorganism and / or compound or chemical mixture; and a micro-well array, where each micro-well can accommodate a single droplet. In the description, document EP3962653A1 reveals a method for preparing a defined microbiome composition using high-precision droplet methods. However, this solution is not suitable for isolating single strains of microorganisms from fecal samples and does not enable their further culturing for use in biotherapeutics.
[0011] The publication Huang, Y., Sheth, R.U., Zhao, S. et al. High-throughput microbial culturomics using automation and machine learning. Nat Biotechnol 41, 1424-1433 (2023) https: / / doi.org / 10.1038 / s41587-023-01674-2. discloses a high-throughput, open-source, robotic platform for bacterial strain isolation, enabling rapid generation of isolates on demand. The disclosed approach utilizes machine learning, which leverages colony morphology and genomic data to maximize the diversity of isolated microorganisms and allows for targeted selection of specific types of microorganisms.
[0012] Previously known solutions do not allow for optimal design of biotherapeutic drugs and obtaining strains for their production. This is particularly true for biotherapeutic drugs containing a plurality of strains that can interact with each other. Therefore, it is necessary to develop a platform that allows for the optimization of the process of selecting the composition of a biotherapeutic drug containing multiple microorganisms based on reliable experimental data that also takes into account the functional characteristics of the microbiome. Such a platform should provide the means for selecting the composition of the biotherapeutic drug and for isolating and culturing the strains of microorganisms that have been selected for it.
[0013] SUMMARY OF THE INVENTION
[0014] The object of the invention is a platform for biotherapeutic drug discovery, wherein a biotherapeutic drug comprises plurality of microorganisms, at least part of which are strainspecific, wherein the platform comprises: a biobank module for storing biological samples of a subject and a donor in conditions appropriate to the type of a biological sample; a computer system comprising: a data repository module for storing a clinical data and a bioinformatic data associated with gut microbiota of the subject and the donor, wherein the bioinformatic data comprises data obtainable by nucleic acid sequencing, metabolomic or proteomic, wherein the clinical data and the bioinformatic data are obtainable by a faecal microbiota transplantation proof-of-concept interventions or any interventions using live microorganisms to modulate the gut microbiota, on the subject; a data engine module for processing the clinical data and the bioinformatic data, wherein processing the clinical data and the bioinformatic data comprises: transforming the bioinformatic data to obtain useful form of the bioinformatic data; annotating of gut microbiota features based on the bioinformatic data; linking gut microbiota features with the clinical data to identify gut microbiota features that are associated with desired clinical outcome; a biotherapeutic predictor module for predicting of the biotherapeutic drug composition based on data obtained from the data engine module, wherein strain-specific (but not exclusively) microorganisms selected for the biotherapeutic drug composition are specific strains selected from specific donor samples; and a culturomics module for isolating and culturing at least strain-specific (but not exclusively) microorganisms selected for the biotherapeutic drug composition from specific donor samples stored in the biobank module.
[0015] Within the context of the present disclosure, strain-specificity means that particular strain from the proof-of-concept study is finally in the biotherapeutic drug ensuring that the same strain composes the final drug) microorganisms
[0016] The platform integrates multiple advanced technologies, ensuring a comprehensive and precise approach to biotherapeutic drug discovery, thereby enhancing the likelihood of identifying effective microorganisms for therapeutic use (meeting the strain-specificity phenomenon). The described platform for biotherapeutic drug discovery offers several significant advantages, particularly in the context of developing microbiota-based drugs that utilize strain-specific microorganisms. These advantages address critical challenges in the field and provide innovative solutions that enhance the efficacy, precision, and scalability of biotherapeutic drug development.
[0017] Obtaining biotherapeutic drugs with strain-specific microorganisms is crucial for several reasons. Precision and efficacy are significantly improved as different strains of the same microbial species can have vastly different genetic and functional capacities. By selecting specific strains known to produce desired therapeutic effects, the efficacy of the biotherapeutic drug can be maximized, ensuring that the drug can target specific health conditions more effectively. Additionally, strain-specific microorganisms allow for the development of personalized treatments tailored to the unique microbiome composition of individual patients, leading to better clinical outcomes and reduced side effects. Furthermore, using well- characterized strains ensures that the biotherapeutic drug can be produced consistently, with predictable and reproducible effects, which is essential for regulatory approval and clinical use.
[0018] Developing biotherapeutic drugs with strain-specific microorganisms involves several challenges. Isolation and cultivation of many beneficial microorganisms are difficult using traditional microbiological methods, compounded by the need to maintain the viability and functional integrity of these strains during the drug development process. The complexity of microbiome interactions, particularly in the human gut microbiome, presents another challenge due to the intricate interactions between different microbial species, making it difficult to understand and replicate these interactions in a therapeutic context. Additionally, the development of strain-specific biotherapeutics requires the integration and analysis of vast amounts of clinical, genomic, transcriptomic, proteomic, and metabolomic data, which must be processed to identify the most effective strains and their interactions. Once effective strains are identified, scaling up their production while maintaining their functional properties is a significant challenge, requiring optimized cultivation conditions.
[0019] The described platform addresses these challenges through several innovative solutions. The comprehensive biobank module stores biological samples from both subjects and donors under optimal conditions, ensuring the availability of high-quality samples for downstream analysis and drug development. The advanced data repository and the engine module store and process clinical and bioinformatic data, integrating data from various sources, including nucleic acid sequencing, metabolomics, and proteomics, to provide a comprehensive understanding of the gut microbiota. The data engine module employs advanced machine learning and deep learning algorithms to process and analyse the data, including quality filtering, genome assembly, taxonomic classification, gene prediction, and functional annotation. The data engine module can also annotate gut microbiota features like protein-protein interactions, microbial interactions, and metabolic potential, linking these features to clinical outcomes. The biotherapeutic predictor module predicts the optimal composition of the biotherapeutic drug based on the processed data, selecting specific strains from donor samples that are most likely to produce the desired therapeutic effects. The culturomics module isolates and cultures the selected strain-specific microorganisms under optimal conditions, utilizing high-throughput microfluidic systems, robotic technologies, and advanced bioreactor systems to ensure scalable and reproducible cultivation. By incorporating genomic, transcriptomic, proteomic, and metabolomic data, the platform provides a holistic view of the microbiome's functional capacity, allowing for more accurate predictions and the development of more effective biotherapeutics. The platform's use of high-throughput technologies and automation enhances the efficiency and scalability of the drug development process, including automated systems for microbial isolation, cultivation, and data analysis. The platform's ability to analyse individual patient data and predict personalized biotherapeutic compositions ensures that the developed drugs are tailored to the specific needs of each patient, improving clinical outcomes.
[0020] The platform aims to restore the full ecosystem of the microbiota, providing a comprehensive therapeutic effect. By preserving the "strain specificity / dependence" phenomenon, the platform ensures that the final product closely mimics the natural microbiome, enhancing its efficacy and stability in the host.
[0021] Additional advantages include enhanced stability and engraftment, as the platform considers specific engraftment-facilitating genes and co-occurrence optimization to ensure that the selected strains can effectively engraft and function within the patient's gut microbiome The platform aims to restore the full ecosystem of the gut microbiome, providing a stable and functional microbial community that supports overall health. By ensuring reproducibility, consistency, and high-quality data integration, the platform facilitates compliance with regulatory requirements for biotherapeutic drug development. The use of advanced culturomics and bioreactor systems allows for the cultivation of previously unculturable microorganisms, expanding the range of potential therapeutic strains. It must be stressed that not every strain in the biotherapeutic must meet the strain-specificity phenomenon, but at least most of them.
[0022] Preferably, the biotherapeutic drug comprises at least 2, preferably 10, 20, 30, 50, 100, 200, 300 or other number higher than 2, strain-specific (but not exclusively) microorganisms encoding specific functions being a main part of the biotherapeutic “capacity”, as “functional capacity”. Including at least 2, but generally at least 10 or more (50, 100, 200, 300) donordependent microorganisms in the biotherapeutic drug ensures a rich and diverse microbial composition, potentially increasing the therapeutic efficacy and stability of the drug. Most important in the biotherapeutic design are the functions (meaning genes content) encoded within particular microorganisms composing the biotherapeutic, and all the functionality needed to treat some disease or state could be encoded in as less aas two bacteria, but most probably it will be more bacteria, especially taking into account that functions redundancy is very desired.
[0023] Preferably, transforming the bioinformatic data comprises operations selected from group comprising: data quality filtering, genome assembling, metagenome assembling, metagenome assembled genome binning, transcriptome assembling, taxonomic classification, gene predicting, functional annotating, proteome profiling, metabolomic profiling.
[0024] Preferably, the gut microbiota features are selected from group comprising: taxa abundance, functional features abundance, protein-protein interactions, metabolic potential, microorganisms interactions, transcriptomic profile, proteomic profile, metabolomic profile.
[0025] Preferably, the biological samples of a subject and a donor are collected in the form of triad of biological samples comprising three sets of biological samples, wherein a first set of biological samples is collected from the subject before the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of-concept intervention or standard procedures, on the subject, a second set of biological samples is collected from the subject after the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of- concept intervention or standard procedures, and a third set of biological samples is collected from the donor.
[0026] Collecting biological samples in a triad format allows for a detailed comparison of microbiota before and after intervention, providing robust data to identify effective microbial strains and especially their functions, and their impact on clinical outcomes.
[0027] Preferably, the biotherapeutic drug further comprises at least one additive from a group consisting of: metabolites, proteins, bacteriophages, comprises nucleic-acids or its products, and intestinal lumen / wall components.
[0028] The inclusion of additives such as metabolites, proteins, bacteriophages, nucleic acids, and intestinal lumen / wall components can enhance the therapeutic potential and functionality of the biotherapeutic drug, addressing specific health conditions more effectively.
[0029] Preferably, the biotherapeutic predictor module is configurate for predicting of the biotherapeutic drug composition by: identifying optimal biotherapeutic drug composition for an individual, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine biotherapeutic drug composition for the individual, considering specific physiological and therapeutic requirements of the individual; generating stable biotherapeutic drug composition for specific indication, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication or patient profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome; generating stable biotherapeutic drug composition for a wide populations, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication in a wide patients / people profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome.
[0030] Utilizing a generative Al-based model for predicting biotherapeutic drug composition ensures a personalized and stable formulation, tailored to individual, and wide populations physiological and therapeutic needs, thereby increasing the likelihood of successful treatment outcomes.
[0031] Preferably, the culturomics module comprises: an untargeted culturomics module for isolating and culturing a diverse spectrum of microorganisms from the donor sample stored in the biobank module comprising: a classical culturomics submodule for manual culturing microorganisms from specific donor samples in fluid environments or on solid substrates in plurality combinations of culturing media paired with enriching substances; a machine learning guided robotic culturomics submodule for automated culturing plurality of microorganisms on solid substrates and identifying by image analysis potentially the most valuable microbial colonies based on their morphology and growth; a droplet microfluidic culturomics submodule for ultra-high-throughput culturing of cultures from individual bacterial cells and consortia or heterogenic samples by generating plurality of droplet microreactors, wherein volume of each droplet microreactors ranges from 10 picolitres to 10 nanolitres, wherein each droplet microreactors initially contain single cells or micro-consortia, which are next incubated to propagate bacterial cell divisions; a targeted culturomics module for isolating and culturing strain-specific microorganisms selected for the biotherapeutic drug composition comprising: a single-cell genome sequencing submodule for sequencing the genomes of single bacterial / microorganism cells; a reverse-genomic submodule for isolating a target microorganism using fluorescence cell sorting, wherein the target microorganisms expressing specific membrane protein bearing extracellular epitope unique to the target microorganism, wherein specific membrane protein bearing extracellular epitope unique to the target microorganism is identified by genome analysis, wherein fluorescently labelled antibodies binds to extracellular epitope unique to the target microorganism and allows the target microorganism to be recognized by the sorter; a bulk long reads sequencing submodule for predicting specific membrane protein bearing extracellular epitope unique to the target microorganism for use as a input for the reverse-genomic submodule; a culturing optimization module (330) for optimization of culturing conditions for isolated strain-specific microorganisms comprising: a high-throughput culture optimization submodule for determining of culture conditions, including growth media concentrations and additives for the target microorganism by using microdroplet reactors comprising varying concentrations of various chemical enrichment substances in growth media to facilitate the growth of the target microorganism; a high-throughput co-culturing submodule 332 for merging of microdroplets containing different strains of bacteria and co-culturing thereof; a bioreactors and scaling-up submodule 333 for culturing and scaling-up culturing of biomass from the combined strains selected for the biotherapeutic drug.
[0032] The combination of untargeted and targeted culturomics provides comprehensive microbial diversity. The untargeted culturomics module is designed to isolate and culture a wide range of microorganisms from donor samples, employing classical culturomics, machine learning-guided robotic culturomics, and droplet microfluidic culturomics. This multi-faceted approach ensures that a broad spectrum of microbial species, including rare and slow-growing bacteria, are captured. The classical culturomics submodule allows for manual culturing in various media, the robotic submodule automates the process and uses image analysis to identify valuable colonies, and the droplet microfluidic submodule enables ultra-high-throughput culturing at the single-cell level. The targeted culturomics module focuses on isolating specific strains of microorganisms that are predicted to be beneficial for biotherapeutic applications. It includes single-cell genome sequencing, reverse-genomics for targeted isolation using fluorescence cell sorting, and bulk long-read sequencing for antigen prediction. This targeted approach ensures that specific, functionally important strains are isolated and cultured under optimal conditions.
[0033] By combining untargeted and targeted approaches, the platform maximizes the likelihood of isolating both a wide array of microbial species and specific strains with desired therapeutic properties. Untargeted culturomics casts a wide net to capture microbial diversity, while targeted culturomics hones in on specific strains identified through genomic and proteomic analyses. This dual approach ensures that no potentially beneficial microorganisms are overlooked. The targeted culturomics module includes a high-throughput media / culture optimization submodule that determines the optimal culture conditions for specific strains. This ensures that once a target microorganism is isolated, it can be cultured efficiently and effectively, maximizing its viability and functional potential. This is particularly important for strains that may have unique or stringent growth requirements.
[0034] The integration of single-cell genome sequencing and reverse-genomics in the targeted culturomics module provides detailed genomic and functional insights into the isolated strains. This allows for the identification of specific genes and metabolic pathways that are crucial for the therapeutic efficacy of the strains. Such detailed information is invaluable for the rational design of microbial consortia for biotherapeutic drug. The use of machine learning-guided robotic culturomics and droplet microfluidic culturomics in the untargeted module significantly enhances the throughput and efficiency of the isolation process. Automation reduces the time and labour required for culturing and identifying microorganisms, allowing for the rapid screening of large numbers of samples.
[0035] The platform is integrated with biobanking modules, which collect and store biological samples from donor-dependent microbiota preparations and proof-of-concept studies. This integration ensures a continuous supply of high-quality samples and allows for the longitudinal study of microbiota changes in response to therapeutic interventions.
[0036] In another aspect, the invention relates to a method for biotherapeutic drug discovery, wherein a biotherapeutic drug comprises plurality of strain-specific (but not exclusively) microorganisms, the method comprising: storing in a biobank biological samples of a subject and a donor in conditions appropriate to the type of a biological sample; performing specific instructions in computer system comprising: storing in a data repository a clinical data and a bioinformatic data associated with gut microbiota of the subject and the donor, wherein the bioinformatic data comprises data obtainable by nucleic acid sequencing, metabolomic or proteomic, wherein the clinical data and the bioinformatic data are obtainable by a faecal microbiota transplantation or other microbiota-oriented interventions in proof-of-concept intervention or standard procedures on the subject; processing the clinical data and the bioinformatic data stored in the data repository, comprising: transforming the bioinformatic data to obtain useful form of the bioinformatic data; annotating of gut microbiota features based on the bioinformatic data; linking gut microbiota features with the clinical data to identify gut microbiota features that are associated with desired clinical outcome; predicting of the biotherapeutic drug composition based on data obtained by processing the clinical data and the bioinformatic data, wherein strain-specific microorganisms selected fo the biotherapeutic drug composition are specific strains selected from specific donor samples; isolating and culturing at least strain-specific microorganisms selected for the biotherapeutic drug composition from specific donor samples stored in the biobank.
[0037] The method ensures a systematic and thorough approach to biotherapeutic drug discovery, from sample storage to data processing and strain isolation, increasing the reliability and effectiveness of the resulting biotherapeutic drug.
[0038] Preferably, the biological samples of the subject and the donor are collected in the form of triad of biological samples comprising three sets of biological samples, first set of biological samples is collected from the subject before the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of-concept intervention or standard procedures, on the subject, a second set of biological samples is collected from the subject after the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of-concept intervention or standard procedures, and a third set of biological samples is collected from the donor.
[0039] Collecting biological samples in a triad format provides a comprehensive dataset for analysing the impact of faecal microbiota transplantation, aiding in the identification of effective microbial strains and their therapeutic potential. Preferably, predicting of the biotherapeutic drug composition comprises: identifying optimal biotherapeutic drug composition for an individual, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine biotherapeutic drug composition for the individual, considering specific physiological and therapeutic requirements of the individual; generating stable biotherapeutic drug composition for specific indication, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication or patient profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome; generating stable biotherapeutic drug composition for a wide populations, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication in a wide patients / people profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome.
[0040] The use of a generative Al-based model for predicting biotherapeutic drug composition ensures a personalized and stable formulation, tailored to individual physiological and therapeutic needs, thereby increasing the likelihood of successful treatment outcomes. Preferably, isolating at least strain-specific microorganisms selected for the biotherapeutic drug composition comprises: isolating and culturing a diverse spectrum of microorganisms from the donor sample stored in the biobank using untargeted culturomics methods, comprising: manual culturing microorganisms from specific donor samples in fluid environments or on solid substrates in plurality combinations of culturing media paired with enriching substances; automated culturing plurality of microorganisms on solid substrates and identifying by image analysis potentially the most valuable microbial colonies based on their morphology and growth; ultra-high-throughput culturing of cultures from individual bacterial cells and consortia or heterogenic samples by generating plurality of droplet microreactors, wherein volume of each droplet microreactors ranges from 10 picolitres to 10 nanolitres, wherein each droplet microreactors initially contain single cells or micro-consortia, which are next incubated to propagate bacterial cell divisions; isolating and culturing strain-specific microorganisms selected for the biotherapeutic drug composition, comprising: sequencing the genomes of single bacterial / microorganism cells using a single-cell sequencing methods; isolating a target microorganism using fluorescence cell sorting, wherein the target microorganism expressing specific membrane protein bearing extracellular epitope unique to the target microorganism, wherein specific membrane protein bearing extracellular epitope unique to the target microorganism is identified by genome analysis, wherein fluorescently labelled antibodies binds to extracellular epitope unique to the target microorganism and allows the target microorganism to be recognized by the sorter; predicting specific membrane protein bearing extracellular epitope unique to the target microorganism based on bulk long reads sequencing data; optimization of culturing conditions for isolated strain-specific microorganisms comprising: determining of culture conditions, including growth media concentrations and additives for the target microorganism by using microdroplet reactors comprising varying concentrations of various chemical enrichment substances in growth media to facilitate the growth of the target microorganism; merging of microdroplets containing different strains of bacteria and co-culturing thereof; culturing and scaling-up culturing of biomass from the combined strains selected for the biotherapeutic drug.
[0041] The method's high-throughput culturing and targeted isolation techniques ensure the efficient and precise identification and cultivation of strain-specific microorganisms, enhancing the development of effective biotherapeutic drugs.
[0042] Preferably, transforming the bioinformatic data comprises operations selected from group comprising: data quality filtering, genome assembling, metagenome assembling, metagenome assembled genome binning, transcriptome assembling, taxonomic classification, gene predicting, functional annotating, proteome profiling, metabolomic profiling.
[0043] Preferably, the gut microbiota features are selected from group comprising: taxa abundance, functional features abundance, protein-protein interactions, metabolic potential, microorganisms interactions, transcriptomic profile, proteomic profile, metabolomic profile.
[0044] BRIEF DESCRIPTION OF DRAWINGS
[0045] The invention is shown by means of example embodiment in a drawing, wherein: Fig. 1 shows an overall structure of an embodiment of the platform according to the invention; Fig. 2 shows an overall structure of an embodiment of the culturomics module;
[0046] Fig. 3 shows schematically steps of the method for biotherapeutic drug discovery according to the invention.
[0047] DETAILED DESCRIPTION
[0048] In certain aspects, the invention pertains to a platform for biotherapeutic drug discovery. Within the context of this invention, a biotherapeutic drug is a biotechnological medication composed of a plurality of strain-specific microorganisms, naturally found in the human or animal gastrointestinal tract. This biotherapeutic drug is designed to restore or enhance the balance of microorganisms in the gastrointestinal tract and may be beneficial in treating conditions such as antibiotic-resistant bacteria colonization, immunomodulation-needed conditions, inflammatory bowel disease (IBD), irritable bowel syndrome (IBS), Clostridium difficile infections, and metabolic disorders like obesity and type 2 diabetes, and many others.
[0049] The platform includes several modules: a biobank module 101 for storing biological samples of a subject and a donor under conditions suitable for the type of a biological sample; a data repository module 201 for storing a clinical data and a bioinformatic data associated with gut microbiota of the subject and the donor; a data engine module 202 for processing the clinical data and the bioinformatic data; a biotherapeutic predictor module 203 for predicting the biotherapeutic drug composition based on data obtained from the data engine module; and a culturomics module 300 for isolating at least strain-specific (but not exclusively) microorganisms selected for the biotherapeutic drug composition from specific donor samples stored in the biobank module 101. The data repository module 201, the data engine module 202 and the biotherapeutic predictor module 203 are components of a computer system 200.
[0050] The biological samples and the clinical data of the subject and the donor stored in the biobank module 101 can be collected through a faecal microbiota transplantation proof-of- concept intervention on the subject for specific medical indications of interest. The faecal microbiota transplantation proof-of-concept intervention can be performed through a dedicated clinical platform 401, which includes a clinical trial / experiment design module and a clinical inference strategy planning module with a clinical / biological sampling planning module. Any clinical indication, status, or condition enrolled into the clinical platform 401 allows for strategic assessment of the clinical inference strategy and sampling planning. The biological samples can take various forms such as stool, blood, saliva, urine, and others. Stool samples are essential for identifying gut microbiome characteristics of the donor and the subject. Other types of biological samples can provide additional clinical and bioinformatics data useful in identifying the biotherapeutic drug composition.
[0051] Biological samples collected in the biobank module 101 can be used to obtain the bioinformatic data. This can be performed by a multiomics laboratory 402 configured to generate bioinformatic data through nucleic acid sequencing, metabolomics, or proteomics. However, bioinformatic data can also be obtained from external repositories 501, such as public databases (e.g., NCBI, ENA) or medical databases.
[0052] In some embodiments, the data repository module 201 is a data storage system that can comprise a plurality of submodules.
[0053] A primary data storage submodule serves as a centralized repository where all inputs and outputs of each analytical pipeline are securely stored. This includes the raw data and final processed data generated during the execution of bioinformatic / computational pipelines from - omics studies, such as genomic, metabolomic, proteomic, transcriptomic, or single cell studies, or other data generated by basic molecular, microbiological, or other studies. The data storage also contains clinical data, within the metadata submodule. Such storage allows for reproducibility and traceability of each step of analysis, which is crucial during medical research. It also allows for reanalysis in case improvements and modifications are made to the analytical pipelines, while preserving previous versions of results. This facilitates easy comparison and combination of results obtained with different analytical tools.
[0054] A metadata submodule stores all available technical, biological, and medical metadata for each dataset. This submodule is designed to integrate with other components, facilitating data management, analysis, and interpretation of acquired results.
[0055] The data engine module 202 is configured to process the clinical data and the bioinformatic data by: transforming the bioinformatic data to obtain useful form of the bioinformatic data; annotating gut microbiota features based on the bioinformatic data; and linking gut microbiota features with the clinical data to identify gut microbiota features associated with a desired clinical outcome. Processing the clinical data and the bioinformatic data may also include additional operations, but these are not crucial for the platform.
[0056] In some embodiments, the data engine module 202 is designed to extract and process various types of information from metagenomic, genomic, transcriptomic, proteomic, and / or metabolomic bioinformatic data combined with clinical data. The extracted information is subsequently stored in a highly efficient data repository module 201.
[0057] In certain embodiments, the data engine module 202 comprises multiple containerized pipelines that execute all stages of data analysis. For example, metagenomic data can be quality checked, assembled into metagenomes, sorted into metagenome assembled genomes (MAGs), and deduplicated. At this point, it is feasible to merge the metagenomic analytical pathway with the genomic analytical pathway for subsequent steps: taxonomic classification, gene prediction, and functional annotation. The final steps of estimating abundance and detecting interactions with other microorganisms present in the original metagenomic sample are conducted independently of genomic data. This method allows for the streamlined development of individual analytical pipelines (including machine learning-based solutions) and their combination into flexible analytical pathways. It is possible to merge different types of data (e g., genomic and transcriptomic data) and numerous datasets (e.g., all metagenomic data available on the platform) to extract as much useful information as possible in each case. Incorporating diverse gut microbiome data modalities, such as transcriptomics and metabolomics, enables a comprehensive analysis of the consortium's full potential. Solely relying on genomic data overlooks the fact that some genes, despite being present, may not be actively expressed. Transcriptomics, which examines RNA transcripts, helps us understand which genes are being actively transcribed and provides insight into the regulatory mechanisms governing gene expression. Metabolomics, on the other hand, profiles the small molecules and metabolites produced by microbial communities, offering a snapshot of their metabolic activities and pathways. By integrating transcriptomics and metabolomics with genomic data, we gain a more holistic understanding of the microbiome's functional capacity. This approach reveals the actual biochemical activities within the microbiome, allowing us to identify active metabolic pathways and the interactions between different microbial species. Such detailed insights are crucial for developing and optimizing biotherapeutics, as they provide a clearer picture of how microbial consortia function and interact within the gut environment. Moreover, with this integrated data, we can construct detailed metabolic models of bacterial consortia. These can model how microbial communities respond to different environmental conditions or therapeutic interventions, thereby informing the design of more effective microbiome-based therapies. Ultimately, leveraging transcriptomics and metabolomics alongside genomics enables a more precise and functional analysis of microbial consortia, paving the way for advanced applications in microbiome research and biotherapeutics development.
[0058] In some embodiments, the data engine module 202 encompasses machine learning and deep learning algorithms that allow for: categorization of provided data, prediction of proteinprotein interactions, prediction of interactions between microorganisms constituting the microbiome, and modelling the metabolic potential of the microbiome. Machine learning (e.g., Random Forest, Support Vector Machines, k-Nearest Neighbors, XGBoost) and deep learning (e g., Neural Networks, Convolutional Neural Networks, Transformers) methods assist scientists, microbiologists, and statisticians in classifying and identifying important gut microbiota features (specific bacterial strains, metabolites, etc ), which are included in biotherapeutic drugs to be effective medicine for particular diseases. Predictions of proteinprotein interactions allow for the identification of peptides with antimicrobial properties with strain-detailed resolution. Interactions and metabolism modelling help to identify stable, yet diverse bacterial consortia, predict their growth conditions in bioreactors, and their impact on the patient's gut microbiome. At the consortium cultivation level in bioreactors, the composition of the consortium will be determined for the given culture medium. The consortium will be metabolically optimized to prevent the accumulation of intermediate metabolites during biosynthesis, which could disrupt the growth of some strains within the consortium. The analyses and culturing will utilize the phenomenon of cross-feeding among bacterial strains, making it possible to culture strains within the consortium that would not proliferate independently.
[0059] In some embodiments, the data engine module is configured to link gut microbiota features with clinical data to identify gut microbiota features that are associated with a desired clinical outcome.
[0060] The biotherapeutic predictor module 203 is configured to select strain-specific microorganisms, which are specific strains selected from specific donor samples. A key feature of the invention is that the strains identified and selected for use in the biotherapeutic drug are directly linked to the sample. This invention preserves a "strain dependence" phenomenon, maximizing the effectiveness of the final product. In certain embodiments, the biotherapeutic predictor module 203 amalgamates all data and algorithms to forecast the optimal theoretical composition or compositions of a biotherapeutic drug for a specific condition. This composition may be a singular entity or divided into subconsortia for separate testing or merging at the end of the synthesis process. The biotherapeutic predictor module 203 processes each medical dataset assigned to a condition, utilizing data about gut microbiota features from the data engine module, such as metabolites, proteins, bacteriophages, metabolic pathways, nucleic acids or their products and subtypes, and bacteria. These components, either individually or in combination, maximize the chances of curing the disease or condition for each individual patient. Subsequently, the biotherapeutic predictor module 203 integrates the results for all analysed patients, their samples, and gathered data, and adjusts them based on the microbiome composition in the general population. The final outcome is a comprehensive list of components with the highest theoretical probability of curing the disease in as many patients as possible. The biotherapeutic predictor module 203 can also simulate the efficacy of the composed biotherapeutic drug consortia in in silico models and / or compare it to Faecal Microbiota Transplantation (FMT) or other microbiota therapeutics.
[0061] In some embodiments, the platform according to the invention is designed to discover a biotherapeutic drug comprising a number of strain-specific microorganisms, with a value selected from a group consisting of: at least 20, 30, 50, 60, 100, 120, 150, 200, 300. Preferably, the biotherapeutic drug comprises at least 50 strain-specific microorganisms.
[0062] In certain embodiments, the transformation of bioinformatic data involves operations selected from a group consisting of: data quality filtering, genome assembling, metagenome assembling, metagenome assembled genome binning, transcriptome assembling, taxonomic classification, gene predicting, functional annotating, proteome profiling, metabolomic profiling. These operations are standard bioinformatics procedures aimed at obtaining high- quality data and facilitating more in-depth analyses.
[0063] In some embodiments, the gut microbiota features are selected from a group consisting of: taxa abundance, functional features abundance, protein-protein interactions, metabolic potential, microorganisms interactions, transcriptomic profile. According to conducted experiments, these gut microbiota features are crucial in identifying microorganisms that will be a significant component of the biotherapeutic drug. In certain embodiments, the biological samples of the subject and the donor are collected in the form of a triad of biological samples, comprising three sets of biological samples. The first set is collected from the subject before the faecal microbiota transplantation proof-of- concept intervention, the second set is collected from the subject after the intervention, and the third set is collected from the donor. This triad of biological samples enables clinical and experimental inference towards biological mechanisms that warrant the observed effect and, for instance, prediction of what is particularly effective within the complex therapy, to extract it and compose a biosynthetic drug.
[0064] In some embodiments, the biotherapeutic drug further comprises at least one additive from a group consisting of: metabolites, proteins, bacteriophages, nucleic acids or their products, and intestinal lumen / wall components.
[0065] In certain embodiments, the biotherapeutic predictor module 203 is configured to predict the biotherapeutic drug composition by performing specific procedures. One such procedure involves identifying the optimal biotherapeutic drug composition for an individual, wherein a generative Al-based model analyses data obtained from the data engine module 202 to determine the biotherapeutic drug composition for the individual, considering specific physiological and therapeutic requirements of the individual. This personalized approach takes into account the unique characteristics and health needs of each patient, ensuring that the selected microbial consortia can effectively support their specific physiological and therapeutic requirements. Another procedure involves generating a stable biotherapeutic drug composition for a specific indication, wherein a generative Al-based model analyses data obtained from the data engine module 202 to determine a stable biotherapeutic drug composition specific to a particular health indication or patient profile, ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome.
[0066] In certain embodiments, the culturomics module 300 comprises an untargeted culturomics module 310, a targeted culturomics module 320, and a culturing optimization module 330. The untargeted culturomics module 310 is configured to isolate and culture a diverse spectrum of microorganisms from the donor sample stored in the biobank module 101. In some embodiments, the untargeted culturomics module 310 comprises a plurality of submodules. One submodule of the untargeted culturomics module 310 may be a classical culturomics submodule 311, which manually cultures microorganisms from specific donor samples in fluid environments or on solid substrates using a variety of culturing media and enriching substances. This submodule represents a methodological advancement in microbiology, designed to facilitate the isolation, identification, and cultivation of a diverse range of microbial species from biological specimens. The technique is based on the principle of exploiting a wide combination of culture conditions to promote the growth of select, often difficult-to-culture species, while concurrently suppressing the proliferation of dominant bacterial communities. This method requires a rigorous and meticulous process where bacteria are subjected to culturing on various media types and in different volumes. Whether cultured in fluid environments within flasks or on solid substrates such as Petri dishes, this procedure is characterized by its empirical reliance on the combination of various cultivation media and enriching substances, along with various bacterial cultivation conditions. The identification of isolated microorganisms within classical culturomics often relies on techniques such as MALDI-TOF, complemented by 16S rRNA amplification and sequencing, especially for strains with previously unknown spectra or entirely new bacterial species.
[0067] Another submodule of the untargeted culturomics module 310 may be a machine learning guided robotic culturomics submodule 312. This submodule automates the culturing of a variety of microorganisms on solid substrates and identifies potentially valuable microbial colonies based on their morphology and growth through image analysis. The machine learning guided robotic culturomics submodule 312 can be coupled with microfluidics and microorganism identification means to automate the recognition of new colonies, subculturing of cultures undergoing division, and evaluation, using algorithms, to determine if the colony differs from those already in the database. This submodule selects potentially valuable microbial colonies based on their morphology and growth, then transfers them to multi-well plates for identification and biobanking. In some embodiments of the invention, droplets with microcolonies of bacteria can be deposited instead of bacterial suspension. This submodule enables the growth and identification of a higher number of strains, as cultivation in microdroplets has been proven to provide a better representation of rare taxa and slow-growing bacteria compared to plate-based methods.
[0068] Another submodule of the untargeted culturomics module 310 may be a droplet microfluidic culturomics submodule 313, which is used for ultra-high-throughput culturing of cultures from individual bacterial cells and consortia or heterogeneous samples by generating a plurality of droplet microreactors. This submodule serves as a tool for ultra-high-throughput culturing of cultures from individual bacterial cells and consortia or heterogeneous samples (stool, urine, blood, saliva or others) preferably at the level of selection of >500 strains / consortia from at least 20 million microcultures (but can be less if needed) in a single experiment (i.e., assessing and culturing sample from 1 donor or patient). Preferably, the droplet microfluidic culturomics submodule 313 generates >20 million droplet microreactors (but can be less), of approximately 100 picolitres (but this volume can range from 10 picolitres to 10 nanolitres). These microreactors initially contain single cells or micro-consortia, which are then incubated to propagate bacterial cell divisions. In total, the droplet microfluidic culturomics submodule 313 can conduct cultures from samples of faeces or other materials, e g. from 200 sample’s donors, resulting in the production of >8 billion droplet microreactors (as an example). From these, a fraction of cultures (e.g. 100 000), based on predefined conditions, can be selected and deposited on multi-well plates (a single droplet is sorted into 50 microliters (range 1 microliter to 1 millilitre) of medium in each well of a plate. During the assay, the droplets can also be formed in the format of double emulsions, water / oil / water (w / o / w) - meaning aqueous droplets containing nutrients surrounded by a thin layer of oil and suspended in the aqueous phase. Liquid microdroplets are selected based on the growth and morphological features of microcolonies and sorted at ultra-high throughput (thousands of droplets per minute) using a flow cytometer sorter (FACS) into individual wells on multi-well plates for further analysis and identification. This approach allows for the isolation of the maximum number of microcolonies in liquid culture and is synergistic with the robotic culturomics module (point b), in which strains from isolates can grow on solid substrates. This technique allows culturing significantly more unique isolates than classical culturomics and may be more efficient than the machine learning guided robotic culturomics submodule 312, and combined together can yield a higher percentage of microorganisms (e.g. bacteria, especially difficult to grow ones) than classical microbiology methods and / or all the above-mentioned methods separately.
[0069] On the other hand, the targeted culturomics module 320 is configured for isolating and culturing strain-specific microorganisms selected for the biotherapeutic drug composition. The targeted culturomics module 320 is designed for the specific labelling, annotation, isolation, and maintenance of selected microorganisms deemed significant for the biotherapeutic drug composition. Simultaneously with the untargeted culturomics module 310 or if the untargeted culturomics module 310 experiments would not provide the cultures of microorganisms being of special interests to compose the final biotherapeutic drug version, the targeted culturomics module 320 provides particular strains or cells in a targeted mode. The selection of desirable strains in targeted culturomics is based on molecular analyses, such as genome or / and proteome studies of bacteria. In some embodiments, the targeted culturomics module 320 comprises a plurality of submodules.
[0070] One submodule of the targeted culturomics module 320 can be a single-cell genome sequencing submodule 321, designed for sequencing the genomes of individual bacterial / microorganism cells. This submodule utilizes short- and / or long-read methods of nucleic acid sequencing, notably Oxford Nanopore Sequencing, but also potentially PacBio or other methods, for sequencing the genomes of single bacterial / microorganism cells. Molecular barcoding in a microfluidic format, combined with dedicated protocols for library preparation and particularly long-read sequencing methods such as ONT, are essential for obtaining more complete genomes of target bacteria / microorganisms. This includes all DNA content, such as the genome, mobile genetic elements like transposons, plasmids, and others. This approach is particularly useful for microorganisms that have not been successfully cultured by non-targeted culturomics methods and culturomics supported with metabolic modelling. The single-cell genome sequencing submodule 321 comprises a microfluidic system that enables the encapsulation of individual bacterial / microorganism cells in picolitre-sized droplets, followed by cell lysis and the labelling of DNA fragments of their complete genome with molecular markers. These markers are released from polyacrylamide beads co-encapsulated with the bacteria / microorganisms. After sequencing, it is possible to assemble more complete genomes of target bacteria / microorganisms, along with extrachromosomal genetic elements such as DNA plasmids, and assign them to a specific taxon from which they originate. This process enables the prediction of metabolic pathways encoded in bacterial genomes, thereby enriching the culture medium, or co-culture with other bacteria, and selecting specific surface antigens for other submodules.
[0071] Another submodule of the targeted culturomics module 320 can be a reverse-genomic submodule 322, designed for isolating a target microorganism using fluorescence cell sorting. The reverse-genomic submodule 322 employs a sophisticated approach rooted in reverse genomics guided isolation. By leveraging genomic data harvested from Single-cell Amplified Genomes (SAGs) and Metagenome-Assembled Genomes (MAGs), this method meticulously searches for genes encoding membrane proteins bearing extracellular epitopes unique to a target microorganism. Once identified, these antigenic peptides undergo a synthesis process, subsequently serving as foundational elements for producing specific antibodies. An enhancement step follows, wherein the resultant antibodies are conjugated with fluorescent markers, ensuring their detectability under predetermined conditions. Upon entering the incubation phase, these fluorescently-conjugated antibodies are exposed to composite microbiota samples, setting the stage for targeted interaction. Employing Fluorescence Activated Cell Sorting (FACS) technology, cells exhibiting fluorescence — indicative of successful antibody interaction — are systematically isolated from the overarching community. Notably, a salient feature of this protocol is the unwavering preservation of cellular viability, ensuring that cells remain receptive to subsequent cultivation. However, the cultivation's success remains contingent on accurately determining the microbe-specific cultivation parameters. Further enriching this module, data derived from the single-cell genome sequencing submodule 321 is instrumental in ascertaining the amino acid sequence of antigens situated on the bacterial cell's surface. This invaluable insight paves the way for the synthesis of bespoke antibodies tailored to the earmarked antigen. The final phase encompasses the use of a state-of- the-art flow cytometer equipped with a sorter, facilitating the meticulous selection of fluorescently labelled strains. Post-sorting, these identified cells are allocated into designated wells on a multi-well plate, thereby priming them for an optimized cultivation procedure in the previously described or subsequent modules.
[0072] Another submodule of the targeted culturomics module 320 can be a bulk long reads sequencing submodule 323, designed for predicting specific membrane protein bearing extracellular epitope unique to the target microorganism for use as input for the reverse- genomic submodule 322. The bulk long reads sequencing submodule 323 allows for subsequent bacteria / microorganism antigens prediction without performing single-cell genomics, but sequencing the bulk of isolated DNA, coming from the sample / consortium, with adequately high-resolution covering >80% of the genome and having long protein sequences to design an antigen and compatible antibody to perform targeted isolation of desired bacteria / microorgani sm .
[0073] The culturing optimization module 330 is configured for optimization of culturing conditions for isolated microorganism strains. In some embodiments, the culturing optimization module 330 comprises a plurality of submodules.
[0074] One submodule of the culturing optimization module 330 can be a high-throughput culture optimization submodule 331, designed for determining culture conditions, including growth media concentrations and additives for the target microorganism by using microdroplet reactors comprising varying concentrations of various chemical enrichment substances in growth media to facilitate the growth of the target microorganism. Microdroplet reactors comprise varying concentrations of various chemical enrichment substances in growth media to facilitate the growth of microorganisms identified by the untargeted or targeted culturomics modules (e.g. the reverse-genomics submodule). The high-throughput culture optimization submodule 331 has the capability to label, for example, >= 10,000 combinations of enrichment substances using a mixture of fluorescent dyes in, for example, >= 20 million droplets, or fewer droplets if needed (each with an optimum volume of 100 picolitres, but also different volumes from the range 10 picolitres to 10 nanolitres can be used) in a single experiment. Multiplexed fluorescence measurement is conducted in droplets where colony growth is observed to read the combinations of specified concentrations of media and chemical additives. Higher volumes (200 microliters per well) of liquid with a determined nutrient and additives composition are added to each well. In a single experiment, e.g., 96 droplets will be deposited in a multi-well plate (which means gathering information on e.g., 96 possible combinations of media where bacterial growth occurs).
[0075] Another submodule within the culturing optimization module 330 could be a high- throughput co-culturing submodule 332. This submodule facilitates the combination of microdroplets containing bacteria (both healthy and pathogenic, the latter to challenge those intended to compose the biotherapeutic drug) with liquids containing similar bacteria in hundreds of thousands of combinations in at least 20 million droplet bioreactors per experiment, or fewer if required. Microdroplets where bacterial or microorganism growth is detected, indicating that co-culturing these specific organisms stimulates their growth, are sorted using microfluidic devices and deposited into multi-well plates containing 200 microliters of growth medium per well. Additionally, the high-throughput co-culturing submodule 332 can inject fluorescently labelled pathogens, including bacteria against which the biotherapeutic drug is designed to act (e.g., C. difficile, antibiotic-resistant bacteria, etc ), into droplets along with bacteria / microorganisms / components selected by the biotherapeutic predictor module 203. This is done to confirm lytic or lytic-enhancing properties of predicted or newly identified strains and their potential use in the biotherapeutic drug. In total, at least 100 co-cultivation assays can be conducted in a single experiment, depositing droplets in two plates with, for example, 96 wells each, totalling approximately 192 bacterial consortia with the target strain, including 96 consortia that inhibit the growth of at least one out of 10 antibiotic-resistant pathogenic strains.
[0076] Another submodule of the culturing optimization module 330 could be a bioreactors and scaling-up submodule 333. This tool is designed for the culturing and scaling-up of biomass from the combined strains that constitute the biotherapeutic drug. The bioreactors and scaling- up submodule 333 uses data collected earlier in the platform to optimize cultivation conditions for efficient bacterial biomass production. This data includes the composition of consortia, type of medium, type and concentration of enriching substances, and cultivation conditions such as the gas mixture, temperature, and cultivation duration. This submodule operates a system of scalable bioreactors under anaerobic conditions. The bioreactors in the system have been chosen to ensure high throughput initial cultures and also allow for the cultivation of strains separately or in smaller consortia. The subsequent bioreactor system enables scaling, where smaller volume bioreactors can serve as pre-seed and seed fermenters for the production bioreactor, depending on the need. The system adheres to the inoculum dilution ratio of 1 :20- 1:100. The bioreactors are equipped with specially selected probe systems, including growth sensors for biomass, pH, DO, redox, and fluorescence, including measurements of riboflavin B2 and NAD(P)H, as well as external oxygen sensors. It is beneficial if the system utilizes components from bioreactors designed for eukaryotic cell cultures to maintain fully anaerobic conditions and avoid the use of silicone elements. The bioreactor system represents state-of- the-art equipment suitable for culturing from microliters through millilitres to litters. During culturing, monitoring of metabolites excreted by the bacterial consortium is conducted as one of the indicators of the biological activity of the system in real-time mode. The composition of the bacterial consortium can be determined by metagenomic analysis, quantitative assessments such as qPCR, microbiological techniques such as colony counting, or combinations of these methods.
[0077] In another aspect, the present invention relates to a method for biotherapeutic drug discovery, wherein the biotherapeutic drug comprises a plurality of strain-specific microorganisms. The method is particularly advantageous for identifying and developing biotherapeutic drugs that leverage the unique properties of gut microbiota to achieve desired clinical outcomes. The method comprises several steps, which are described in detail below.
[0078] The first step in the method involves storing 601 biological samples from both a subject and a donor in a biobank. These samples may include, but are not limited to, fecal samples, blood samples, and tissue biopsies. The storage conditions are tailored to the type of biological sample to ensure the preservation of the microbiota and other biological materials. For instance, fecal samples may be stored at -80°C to maintain microbial viability, while blood samples might be stored in cryogenic vials with appropriate cryoprotectants.
[0079] The next step involves storing 602 clinical data and bioinformatic data associated with the gut microbiota of the subject and the donor in a data repository. The clinical data may include patient history, treatment outcomes, and other relevant medical records. The bioinformatic data comprises data obtainable by nucleic acid sequencing, metabolomic, or proteomic analyses. This data is collected from faecal microbiota transplantation (FMT) proof- of-concept interventions performed on the subject. The data repository is designed to securely store and manage large datasets, ensuring data integrity and accessibility for subsequent processing.
[0080] The clinical and bioinformatic data stored in the data repository is then processed 603 to extract meaningful insights. This processing 603 step involves several sub-steps:
[0081] The bioinformatic data is transformed 604 into a form that can be readily analysed. This may involve normalizing sequencing data, converting raw metabolomic spectra into quantifiable metabolites, or translating proteomic data into protein expression levels. Advanced bioinformatic tools and algorithms are utilized to ensure high-quality data transformation.
[0082] The transformed bioinformatic data is used to annotate 605 specific gut microbiota features. This annotation process involves identifying microbial species, strains, and their functional capabilities. Machine learning algorithms and reference databases are utilized to achieve accurate and comprehensive annotations.
[0083] The annotated gut microbiota features are then linked 606 with the clinical data to identify gut microbiota features that are associated with desired clinical outcomes. Statistical analyses and correlation studies are performed to establish these associations. For example, specific microbial strains may be linked to improved gastrointestinal health or enhanced immune response.
[0084] Based on the data obtained from processing 603 the clinical and bioinformatic data, the composition of the biotherapeutic drug is predicted 607. This involves selecting specific strains of microorganisms from the donor samples that are associated with the desired clinical outcomes. Computational models and predictive analytics are employed to optimize the selection of strain-specific microorganisms, ensuring the efficacy and safety of the biotherapeutic drug.
[0085] The final step involves isolating and culturing 608 the selected strain-specific microorganisms from the donor samples stored in the biobank. Advanced microbiological techniques are used to isolate pure cultures of the desired strains. These strains are then cultured under controlled conditions to produce sufficient quantities for therapeutic use. Quality control measures are implemented to ensure the purity, viability, and consistency of the cultured microorganisms.
[0086] In summary, the method for biotherapeutic drug discovery described herein provides a systematic approach to identify and develop biotherapeutic drugs based on the unique properties of gut microbiota. By integrating clinical data, bioinformatic analyses, and advanced microbiological techniques, the method enables the discovery of effective and personalized biotherapeutic treatments. In some embodiments, the biological samples of the subject and the donor are collected in the form of a triad of biological samples comprising three sets of biological samples. A first set of biological samples is collected from the subject before the faecal microbiota transplantation proof-of-concept intervention on the subject, a second set of biological samples is collected from the subject after the faecal microbiota transplantation proof-of-concept intervention, and a third set of biological samples is collected from the donor.
[0087] In some embodiments, predicting the biotherapeutic drug composition comprises:
[0088] - Identifying the optimal biotherapeutic drug composition for an individual, wherein a generative Al-based model analyses the clinical data and the bioinformatic data stored in the data repository to determine the biotherapeutic drug composition for the individual, considering specific physiological and therapeutic requirements of the individual;
[0089] - Generating a stable biotherapeutic drug composition for a specific indication, wherein a generative Al-based model analyses the clinical data and the bioinformatic data stored in the data repository to determine a biotherapeutic drug composition specific to a particular health indication or patient profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome.
[0090] In some embodiments, isolating at least strain-specific microorganisms selected for the biotherapeutic drug composition comprises:
[0091] - Isolating and culturing a diverse spectrum of microorganisms from the donor sample stored in the biobank using untargeted culturomics methods, which include manual culturing of microorganisms from specific donor samples in fluid environments or on solid substrates in a plurality of combinations of culturing media paired with enriching substances;
[0092] - Automated culturing of a plurality of microorganisms on solid substrates and identifying potentially the most valuable microbial colonies based on their morphology and growth through image analysis;
[0093] - Ultra-high-throughput culturing of cultures from individual bacterial cells and consortia or heterogeneous samples by generating a plurality of droplet microreactors, wherein the volume of each droplet microreactor ranges from 10 picolitres to 10 nanolitres, and each droplet microreactor initially contains single cells or micro-consortia, which are then incubated to propagate bacterial cell divisions; - Isolating and culturing strain-specific microorganisms selected for the biotherapeutic drug composition, which includes sequencing the genomes of single bacterial / microorganism cells using single-cell sequencing methods;
[0094] - Isolating a target microorganism using fluorescence cell sorting, wherein the target microorganism expresses a specific membrane protein bearing an extracellular epitope unique to the target microorganism, which is identified by genome analysis, and fluorescently labelled antibodies bind to the extracellular epitope unique to the target microorganism and allow the target microorganism to be recognized by the sorter;
[0095] - Predicting a specific membrane protein bearing an extracellular epitope unique to the target microorganism based on bulk long reads sequencing data;
[0096] - Optimizing the culturing conditions for isolated strain-specific microorganisms, which includes determining the culture conditions, including growth media concentrations and additives for the target microorganism by using microdroplet reactors comprising varying concentrations of various chemical enrichment substances in growth media to facilitate the growth of the target microorganism;
[0097] - Merging of microdroplets containing different strains of bacteria and co-culturing thereof;
[0098] - Culturing and scaling-up the culturing of biomass from the combined strains selected for the biotherapeutic drug.
[0099] In some embodiments, transforming the bioinformatic data comprises operations selected from a group comprising: data quality filtering, genome assembling, metagenome assembling, metagenome assembled genome binning, transcriptome assembling, taxonomic classification, gene predicting, functional annotating, proteome profiling, metabolomic profiling.
[0100] In some embodiments, the gut microbiota features are selected from a group comprising: taxa abundance, functional features abundance, protein-protein interactions, metabolic potential, microorganisms interactions, transcriptomic profile.
Claims
CLAIMS1. A platform for biotherapeutic drug discovery, wherein a biotherapeutic drug comprises plurality of microorganisms, at least part of which are strain-specific, wherein the platform comprises: a biobank module (101) for storing biological samples of a subject and a donor in conditions appropriate to the type of a biological sample; a computer system (200) comprising: a data repository module (201) for storing a clinical data and a bioinformatic data associated with gut microbiota of the subject and the donor, wherein the bioinformatic data comprises data obtainable by nucleic acid sequencing, metabolomic or proteomic, wherein the clinical data and the bioinformatic data are obtainable by a faecal microbiota transplantation proof-of-concept interventions or any interventions using live microorganisms to modulate the gut microbiota, on the subject; a data engine module (202) for processing the clinical data and the bioinformatic data, wherein processing the clinical data and the bioinformatic data comprises: transforming the bioinformatic data to obtain useful form of the bioinformatic data; annotating of gut microbiota features based on the bioinformatic data; linking gut microbiota features with the clinical data to identify gut microbiota features that are associated with desired clinical outcome; a biotherapeutic predictor module (203) for predicting of the biotherapeutic drug composition based on data obtained from the data engine module (202), wherein strainspecific (but not exclusively) microorganisms selected for the biotherapeutic drug composition are specific strains selected from specific donor samples; and a culturomics module (300) for isolating and culturing at least strain-specific (but not exclusively) microorganisms selected for the biotherapeutic drug composition from specific donor samples stored in the biobank module (101).
2. The platform for biotherapeutic drug discovery according to claim 1, characterised by that the biotherapeutic drug comprises at least 2, preferably 10, 20, 30, 50, 100, 200, 300 or other number higher than 2, strain-specific (but not exclusively) microorganisms encoding specific functions being a main part of the biotherapeutic “capacity”, as “functional capacity”.
3. The platform for biotherapeutic drug discovery according to any one of proceeding claims, characterised by that transforming the bioinformatic data comprises operations selected from group comprising: data quality filtering, genome assembling, metagenome assembling, metagenome assembled genome binning, transcriptome assembling, taxonomic classification, gene predicting, functional annotating, proteome profiling, metabolomic profiling.
4. The platform for biotherapeutic drug discovery according to any one of proceeding claims, characterised by that the gut microbiota features are selected from group comprising: taxa abundance, functional features abundance, protein-protein interactions, metabolic potential, microorganisms interactions, transcriptomic profile, proteomic profile, metabolomic profile.
5. The platform for biotherapeutic drug discovery according to any one of proceeding claims, characterised by that the biological samples of a subject and a donor are collected in the form of triad of biological samples comprising three sets of biological samples, wherein a first set of biological samples is collected from the subject before the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of-concept intervention or standard procedures, on the subject, a second set of biological samples is collected from the subject after the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of- concept intervention or standard procedures, and a third set of biological samples is collected from the donor.
6. The platform for biotherapeutic drug discovery according to any one of proceeding claims, characterised by that the biotherapeutic drug further comprises at least one additive from a group consisting of: metabolites, proteins, bacteriophages, comprises nucleic-acids or its products, and intestinal lumen / wall components.
7. The platform for biotherapeutic drug discovery according to any one of proceeding claims, characterised by that the biotherapeutic predictor module (203) is configurate for predicting of the biotherapeutic drug composition by: identifying optimal biotherapeutic drug composition for an individual, wherein generative Al-based model analysing data obtained from the data engine module 202 todetermine biotherapeutic drug composition for the individual, considering specific physiological and therapeutic requirements of the individual; generating stable biotherapeutic drug composition for specific indication, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication or patient profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome; generating stable biotherapeutic drug composition for a wide populations, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication in a wide patients / people profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome.
8. The platform for biotherapeutic drug discovery according to any one of proceeding claims, characterised by that the culturomics module (300) comprises: an untargeted culturomics module (310) for isolating and culturing a diverse spectrum of microorganisms from the donor sample stored in the biobank module (101) comprising: a classical culturomics submodule (311) for manual culturing microorganisms from specific donor samples in fluid environments or on solid substrates in plurality combinations of culturing media paired with enriching substances; a machine learning guided robotic culturomics submodule (312) for automated culturing plurality of microorganisms on solid substrates and identifying by image analysis potentially the most valuable microbial colonies based on their morphology and growth; a droplet microfluidic culturomics submodule (313) for ultra-high-throughput culturing of cultures from individual bacterial cells and consortia or heterogenic samples by generating plurality of droplet microreactors, wherein volume of each droplet microreactors ranges from 10 picolitres to 10 nanolitres, wherein each droplet microreactors initially contain single cells or micro-consortia, which are next incubated to propagate bacterial cell divisions; a targeted culturomics module (320) for isolating and culturing strain-specific microorganisms selected for the biotherapeutic drug composition comprising: a single-cell genome sequencing submodule (321) for sequencing the genomes of single bacterial / microorganism cells;a reverse-genomic submodule (322) for isolating a target microorganism using fluorescence cell sorting, wherein the target microorganisms expressing specific membrane protein bearing extracellular epitope unique to the target microorganism, wherein specific membrane protein bearing extracellular epitope unique to the target microorganism is identified by genome analysis, wherein fluorescently labelled antibodies binds to extracellular epitope unique to the target microorganism and allows the target microorganism to be recognized by the sorter; a bulk long reads sequencing submodule (323) for predicting specific membrane protein bearing extracellular epitope unique to the target microorganism for use as a input for the reverse-genomic submodule; a culturing optimization module (330) for optimization of culturing conditions for isolated strain-specific microorganisms comprising: a high-throughput culture optimization submodule (331) for determining of culture conditions, including growth media concentrations and additives for the target microorganism by using microdroplet reactors comprising varying concentrations of various chemical enrichment substances in growth media to facilitate the growth of the target microorganism; a high-throughput co-culturing submodule 332 for merging of microdroplets containing different strains of bacteria and co-culturing thereof; a bioreactors and scaling-up submodule 333 for culturing and scaling-up culturing of biomass from the combined strains selected for the biotherapeutic drug.
9. A method for biotherapeutic drug discovery, wherein a biotherapeutic drug comprises plurality of strain-specific (but not exclusively) microorganisms, the method comprising: storing (601) in a biobank biological samples of a subject and a donor in conditions appropriate to the type of a biological sample; performing specific instructions in computer system comprising: storing (602) in a data repository a clinical data and a bioinformatic data associated with gut microbiota of the subject and the donor, wherein the bioinformatic data comprises data obtainable by nucleic acid sequencing, metabolomic or proteomic, wherein the clinical data and the bioinformatic data are obtainable by a faecal microbiota transplantation or other microbiota-oriented interventions in proof-of- concept intervention or standard procedures on the subject;processing (603) the clinical data and the bioinformatic data stored in the data repository, comprising: transforming (604) the bioinformatic data to obtain useful form of the bioinformatic data; annotating (605) of gut microbiota features based on the bioinformatic data; linking (606) gut microbiota features with the clinical data to identify gut microbiota features that are associated with desired clinical outcome; predicting (607) of the biotherapeutic drug composition based on data obtained by processing the clinical data and the bioinformatic data, wherein strain-specific microorganisms selected fo the biotherapeutic drug composition are specific strains selected from specific donor samples; isolating and culturing (608) at least strain-specific microorganisms selected for the biotherapeutic drug composition from specific donor samples stored in the biobank.
10. The method for biotherapeutic drug discovery according to claim 9, characterised by that, the biological samples of the subject and the donor are collected in the form of triad of biological samples comprising three sets of biological samples, first set of biological samples is collected from the subject before the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of-concept intervention or standard procedures, on the subject, a second set of biological samples is collected from the subject after the faecal microbiota transplantation or other microbiota-oriented interventions in proof-of-concept intervention or standard procedures, and a third set of biological samples is collected from the donor.
11. The method for biotherapeutic drug discovery according to claim 9 or 10, wherein predicting of the biotherapeutic drug composition comprises: identifying optimal biotherapeutic drug composition for an individual, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine biotherapeutic drug composition for the individual, considering specific physiological and therapeutic requirements of the individual; generating stable biotherapeutic drug composition for specific indication, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication orpatient profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome; generating stable biotherapeutic drug composition for a wide populations, wherein generative Al-based model analysing data obtained from the data engine module 202 to determine stable biotherapeutic drug composition specific to particular health indication in a wide patients / people profile by ensuring the presence of beneficial bacterial strains and functions that are crucial for the desired therapeutic outcome.
12. The method for biotherapeutic drug discovery according to any one of claim 9-11, wherein isolating at least strain-specific microorganisms selected for the biotherapeutic drug composition comprises: isolating and culturing a diverse spectrum of microorganisms from the donor sample stored in the biobank using untargeted culturomics methods, comprising: manual culturing microorganisms from specific donor samples in fluid environments or on solid substrates in plurality combinations of culturing media paired with enriching substances; automated culturing plurality of microorganisms on solid substrates and identifying by image analysis potentially the most valuable microbial colonies based on their morphology and growth; ultra-high-throughput culturing of cultures from individual bacterial cells and consortia or heterogenic samples by generating plurality of droplet microreactors, wherein volume of each droplet microreactors ranges from 10 picolitres to 10 nanolitres, wherein each droplet microreactors initially contain single cells or micro-consortia, which are next incubated to propagate bacterial cell divisions; isolating and culturing strain-specific microorganisms selected for the biotherapeutic drug composition, comprising: sequencing the genomes of single bacterial / microorganism cells using a singlecell sequencing methods; isolating a target microorganism using fluorescence cell sorting, wherein the target microorganism expressing specific membrane protein bearing extracellular epitope unique to the target microorganism, wherein specific membrane protein bearing extracellular epitope unique to the target microorganism is identified by genome analysis, wherein fluorescently labelled antibodies binds to extracellular epitope uniqueto the target microorganism and allows the target microorganism to be recognized by the sorter; predicting specific membrane protein bearing extracellular epitope unique to the target microorganism based on bulk long reads sequencing data; optimization of culturing conditions for isolated strain-specific microorganisms comprising: determining of culture conditions, including growth media concentrations and additives for the target microorganism by using microdroplet reactors comprising varying concentrations of various chemical enrichment substances in growth media to facilitate the growth of the target microorganism; merging of microdroplets containing different strains of bacteria and coculturing thereof; culturing and scaling-up culturing of biomass from the combined strains selected for the biotherapeutic drug.
13. The method for biotherapeutic drug discovery according to any one of claim 9-12, characterised by that transforming the bioinformatic data comprises operations selected from group comprising: data quality filtering, genome assembling, metagenome assembling, metagenome assembled genome binning, transcriptome assembling, taxonomic classification, gene predicting, functional annotating, proteome profiling, metabolomic profiling.
14. The method for biotherapeutic drug discovery according to any one of claim 9-13, characterised by that the gut microbiota features are selected from group comprising: taxa abundance, functional features abundance, protein-protein interactions, metabolic potential, microorganisms interactions, transcriptomic profile, proteomic profile, metabolomic profile.