Process for predicting a condition or pathology

A deep neural network-based method analyzing the microbiome and clinical data effectively predicts NEC in premature newborns, addressing the current diagnostic challenges with high accuracy and potential for early intervention.

FR3155837A1Pending Publication Date: 2025-05-30UNIVERSITE CLERMONT AUVERGNE +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2023013206
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current methods for diagnosing necrotizing enterocolitis (NEC) in premature newborns are unreliable and lack effective early detection strategies, primarily due to the complex and dynamic nature of the neonatal gut microbiota.

Method used

A deep neural network-based method that analyzes the intestinal or fecal microbiome of premature newborns, combined with relevant clinical data, to predict the presence or absence of NEC, allowing for early therapeutic interventions.

Benefits of technology

The method achieves high accuracy in predicting NEC, with a 94.6% success rate and an AUROC of 0.987 ± 0.01, significantly improving early detection and potentially reducing morbidity and mortality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to an in vitro method for predicting a condition or pathology from a biological sample taken from the digestive system and / or from the excretions of a subject. More particularly, the invention relates to an in vitro method for predicting a pathology of the digestive system or an extra-digestive pathology of a subject from the analysis of the microbiota present in a biological sample taken from the digestive system and / or from the stools of a subject. Even more particularly, the invention relates to a method for predictive diagnosis of necrotizing ulcerative enterocolitis (NEC) in premature newborns. Figure for abstract: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method for predicting a condition or pathology Technical field of the invention

[0001] The present invention relates to an in vitro method for predicting a condition or pathology from a biological sample taken from the digestive system and / or from the excretions of a subject. More particularly, the invention relates to an in vitro method for predicting a pathology of the digestive system or an extradigestive pathology of a subject from the analysis of the microbiota present in a biological sample taken from the digestive system and / or from the stools of a subject. Even more particularly, the invention relates to a method for predictive diagnosis of necrotizing ulcerative enterocolitis (NEC) in premature newborns. The present method is therefore in the field of in vitro diagnosis and personalized medicine. State of the art

[0002] Necrotizing enterocolitis (NEC) is the most common life-threatening gastrointestinal emergency encountered by preterm infants in neonatal intensive care units. It is defined as ulcerative inflammation of the intestinal wall. Current clinical practice for diagnosing NEC is based on clinical, radiological, and hematological findings constituting the Bell criteria, according to a recent review (D'Angelo et al., 2018). Clinical signs of early NEC are often very subtle and may initially manifest as feeding intolerance and nonspecific symptoms before gastrointestinal symptoms become evident.These include increased gastric residuals, bloody stools, and abdominal distension; these may progress to generalized hypotonia, lethargy, and cardiorespiratory failure, which may also be features of other neonatal conditions, including sepsis and viral intestinal infections. If the disease is not diagnosed and treated early, it can lead to severe sepsis, intestinal perforation, and significant morbidity and mortality.

[0003] To date, clinicians have no reliable diagnostic tool for predicting NEC. The pathophysiology of NEC remains poorly understood and effective methods for its early detection have yet to be established. Therefore, current efforts to understand and predict NEC are focused on the study of its risk factors. Preterm birth represents the most important risk factor for the development of NEC. In newborns with very low birth weight (<1.5 kg at birth), the incidence of NEC ranges from 5% to 13%. In addition, Prolonged administration of antibiotics during the first week of life and substitution of breast milk with formula or infant formula are frequently associated with the later onset of NEC.

[0004] Colonization of the gut microbiota has been widely considered to play a role in the development of NEC in preterm neonates, despite the consistent failure to identify a single opportunistic pathogen or pathogenic microbial community as the cause of NEC. This failure is primarily due to the early and highly dynamic establishment of the neonatal gut microbiota, influenced by many factors, including environment, sex, gestational age, mode of delivery, feeding method, and antibiotic treatments.

[0005] The association between intestinal colonization and the development of necrotizing enterocolitis (NEC) in preterm infants has been of long-standing interest. However, despite extensive research, there has been a lack of consistent identification of a specific opportunistic pathogen or pathogenic microbial community that can be directly linked to the onset of NEC.

[0006] Therefore, there is a very significant need for a reliable and reproducible method for predictive diagnosis of pathologies affecting newborns, particularly premature newborns. Such a predictive diagnosis would make it possible to identify newborns at risk and anticipate the management of pathologies likely to seriously affect their lives. Description of the invention

[0007] The inventors have now developed and optimized a deep neural network and a method allowing a predictive diagnosis of a condition or pathology in a subject, using a precise analysis of the microbiome of said subject. According to a particular embodiment, the invention relates to a method for predictive diagnosis of a digestive pathology or an extra-digestive pathology, based on a precise analysis of the intestinal, or fecal, microbiome of said subject, in combination with at least one relevant clinical data. A method according to the invention potentially makes it possible to carry out early therapeutic interventions in order to prevent the development or the worst complications of a pathology.

[0008] A method according to the invention has the advantage of being reliable, reproducible and relatively quick to implement; it meets a previously unmet clinical need.

[0009] In a particular aspect of a method according to the invention, said deep neural network based on the intestinal microbiota is combined with at least one clinical data of the subject, to predict the presence, or on the contrary the absence, of a condition which may lead to pathology or a full-fledged pathology in a subject.

[0010] A method according to the invention comprises a step of reconstructing at least a part, preferably the major part, of the sequence of the gene expressing the 16S rRNA and / or the sequence of the gene expressing the 18S rRNA of the microorganisms present in the biological sample, followed by a step of taxonomic classification of said reconstructed genes. The analysis of the classification of the reconstructed genes expressing the 16S rRNA and / or the 18S rRNA of said microorganisms allows the precise identification, at the genus and possibly species level, and the measurements of relative abundances of a plurality of microorganisms present in the biological sample.

[0011] A method according to the invention uses all of the metagenomic data of the microbiota which then allows the complete reconstruction of the 16S and / or 18S rDNAs and a precise affiliation of the microorganisms of the microbial community at the genus or species level, or even the identification of new microorganisms.

[0012] The study of the prior art shows that several methods for analyzing the microbiota have been developed and tested. However, a method comprising the amplification of fragments of a size less than 300 base pairs of the gene expressing the 16S rRNA has several limitations: the biases likely to be generated during the amplification step carried out by PCR can alter the vision of the real diversity of the microbiota, the short length of the fragment provides only a low taxonomic resolution.

[0013] On the other hand, a method according to the invention makes it possible to transmit simple predictive information to a clinician. A method according to the invention therefore has the advantages of high accuracy in prediction, combined with high speed and high throughput.

[0014] Furthermore, the optional application of a reclassification strategy on acquired longitudinal samples makes it possible to further improve the performance of a method according to the invention. With regard to a pathology as serious as infant ECUN for example, although the effectiveness of a method according to the invention is superior to any known diagnostic method for this pathology, the possibility of having an additional gain in efficiency in predictive diagnosis represents a valuable asset for patients and for clinicians.

[0015] The inventors have in particular developed a predictive diagnostic approach for ECUN based on a learning approach using deep neural networks. The processing of metagenomic sequencing data from the stools of premature babies collected before the onset of the pathology by RiboTaxa bioinformatics chaining of 2 studies made it possible to generate intestinal microbiota input data associated with metadata to perform learning by deep neural networks making it possible to distinguish premature babies who have developed ECUN and premature babies who have not developed this digestive pathology.

[0016] An artificial intelligence-based ECUN prediction model is therefore very useful for identifying premature newborns at risk, strengthening monitoring and enabling rapid therapeutic responses that avoid potential serious health problems. Indeed, a method according to the invention makes it possible to very effectively predict ECUN early and to distinguish unaffected infants just as effectively. Indeed, a method according to the invention, developed on 1355 unique microbial species and 5 clinical data (sex, gestational age, day of life / day of sampling, mode of birth and birth weight), has demonstrated its ability to distinguish complex inter-individual microbial interactions in non-ECUN and ECUN samples with an accuracy of 94.6% and an AUROC of 0.987 ± 0.01. Such a degree of reliability is unmatched to date among predictive diagnostic methods for ECUN, particularly in premature newborns.

[0017] The inventors clearly confirmed the effectiveness of the model's prediction on a new cohort of infants which was established by the inventors after the development of the method, and which had not been integrated during the training of the deep neural network.

[0018] Furthermore, a method according to the invention allowed the inventors to observe the association of several species of microorganisms with the presence of the pathology, on the one hand, and to observe the association of several species of microorganisms with the absence of the pathology, on the other hand. Indeed, the inventors observed that several species of Lactobacillus were associated with non-ECUN cases, while several other bacterial species such as: unclassified Enterobacter, unclassified Enterobacteriaceae, Enterococcus faecalis, unclassified Klebsiella, Haemophilus parainfluenzae, Enterococcus durons and Enterobacter cancerogenus were associated with ECUN cases. These results suggest that the prediction of the phenotype is a function of both dominant, subdominant and even rare taxa, highlighting that no individual species or taxonomic group of species is exclusively responsible for an increased risk of ECUN.Instead, it is likely that various microbial consortia can cause inflammatory cascades leading to the onset of ECUN. A method according to the invention therefore further allows for a precise mapping of microorganisms associated with the presence of a condition that can lead to pathology or of a pathology, and of microorganisms associated with the absence of a condition that leads to pathology or of a pathology, on the other hand.

[0019] A method according to the invention has the advantage, from a simple sample of microbiota, for example in the stools of a subject, and its direct metagenomic sequencing, of determining with high certainty a predictive diagnosis of a digestive system disease or extra-digestive disease. This approach can be used in the context of personalized medicine to assess the relevance of more precise clinical monitoring and / or the use of therapeutic treatment, monitoring and treatment methods are well known.

[0020] The search for microbial signatures for the predictive diagnosis of pathologies is particularly complex due to the very strong inter-individual variations. To date, several microbiota analysis techniques exist. However, current approaches do not allow for precise characterization of microbiotas. This is a major challenge in revealing links between these communities of microorganisms and pathologies. The establishment of predictive diagnoses would make it possible to anticipate care or even carry out preventive treatments. The approach of learning metagenomic sequencing data using deep neural networks following pre-processing of the data with software such as RiboTaxa makes it possible to obtain levels of prediction that are unparalleled until now.

[0021] An approach as described in the present patent application is useful for determining predictive diagnoses in many conditions or pathologies, or even in the context of personalized medicine. This approach can not only be used for determining predictive diagnoses but also for evaluating the response of a subject to a therapeutic treatment and / or for monitoring the evolution over time of said condition, the effectiveness of said treatment or said pathology.

[0022] Furthermore, a method according to the invention can make it possible to identify microorganisms which would be key players in various healthy or pathological situations. Microorganisms could thus be identified as being able to play the role of probiotics or for the development of new treatments or even new diagnostics. Detailed description of the invention

[0023] The present invention has as its first object an in vitro method for predictive diagnosis of a state or pathology, from a biological sample taken from the digestive system and / or from the excretions of a subject, said method comprising the following steps: (a) providing nucleic acid from a plurality of microorganisms present in said biological sample, b) determination by sequencing of the nucleotide sequence of at least one nucleic acid chosen from: a gene fragment expressing 16S rRNA, a gene fragment expressing 18S rRNA, a 16S rRNA fragment and an 18S rRNA fragment, of a plurality of microorganisms, to generate a plurality of nucleotide sequences, c) organizing a plurality of nucleotide sequences determined during step b) to reconstruct the nucleotide sequence of at least one gene fragment expressing 16S rRNA, one gene fragment expressing 18S rRNA, one 16S rRNA fragment and one 18S rRNA fragment of a plurality of microorganisms, d) from the results of step c), determining the identity, taxonomic classification, and relative abundance of a plurality of microorganisms present in said sample, and e) from the characteristics determined during step d), determination, by a previously trained classification model, of the predictive diagnosis of a condition or pathology.

[0024] According to a particular aspect, an in vitro method for predictive diagnosis of a condition or pathology, from a biological sample taken from the digestive system and / or from the excretions of a subject, said method comprising an additional preliminary step of isolating the nucleic acid from a plurality of microorganisms present in said biological sample.

[0025] The term "nucleic acid" means the nucleic acid molecules present in the biological sample, in particular deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), in particular ribosomal RNA (rRNA) and even more particularly 16S rRNA and 18S rRNA. The term "gene expressing 16S rRNA" means the DNA nucleotide sequence comprising the nucleotide sequence expressing 16S rRNA. A gene expressing 16S rRNA is also called "16S rDNA". The term "gene expressing 18S rRNA" means the DNA nucleotide sequence comprising the DNA nucleotide sequence encoding 18S rRNA. A gene expressing 18S rRNA is also called "18S rDNA". According to a particular aspect, a method according to the invention comprises determining by sequencing the nucleotide sequence of at least one gene fragment expressing 16S ribosomal RNA (rRNA) and / or a gene fragment expressing 18S rRNA.According to another particular aspect, a method according to the invention comprises the determination by sequencing of the nucleotide sequence of at least one 16S rRNA fragment and / or one 18S rRNA fragment. The isolation of the nucleic acid molecules from the biological sample is carried out according to any technique well known to the person skilled in the art.

[0026] Genes expressing the small subunit of rRNA, i.e. genes called "16S rDNA" for prokaryotic microorganisms, such as bacteria and archaea, and "18S rDNA" for eukaryotes, including yeasts, are used to enable the description of the structure of the microbiota (Chakoory et al., 2022).

[0027] According to a particular aspect of the method of the invention, the nucleic acid of at least 30% of the microbiota present in the biological sample is analyzed in a method according to the invention. More particularly, the nucleic acid of at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% of the microbiota present in the biological sample is analyzed.

[0028] According to another particular aspect, the nucleotide sequence of the gene expressing the 16S rRNA or of the gene expressing the 18S rRNA of at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% of the microbiota present in the biological sample is analyzed.

[0029] "Sequencing" means any known method for determining the nucleotide sequence of a nucleic acid. Among these methods, direct metagenomic sequencing known as "shotgun" is preferred. The use of sequencing data from gene capture approaches by hybridization is also preferred.

[0030] The nucleotide sequences of the genes expressing the 16S rRNA of microorganisms and of the genes expressing the 18S rRNA of microorganisms are at least partially known and are listed in specialized databases that are publicly accessible. Among these databases, the SILVA database is accessible on the website "arb-silva.de". Another example of databases is the "Greengenes" database. The person skilled in the art can easily determine whether a given nucleotide sequence comes from a known or unknown microorganism, or from the human host.

[0031] By "identity determination" is meant the determination of the genus of a microorganism present in the sample and optionally the determination of the species of a microorganism present in the sample. The determination of the identity of a microorganism is carried out from the nucleotide sequence of the 16S rDNA and / or the 18S rDNA.

[0032] By "determination of the taxonomic classification" is meant the organization of a plurality of microorganisms, based on knowledge of the identity of said microorganisms, into hierarchical categories, in other words into taxonomic ranks, these categories consisting of belonging to the domain of life (least precise rank) to the definition of the species (most precise rank). The taxonomic classification is carried out by comparing the 16S rDNA sequences and / or the 18S rDNA sequences determined with, respectively, 16S rDNA sequences and / or the 18S rDNA sequences contained in databases. Among the databases that can be used, the SILVA database may be mentioned in particular.

[0033] By “determination of the relative abundance”, we mean the determination for each of the microorganisms considered for the method according to the invention, of the abundance of the microorganism relative to the total abundance of microorganisms considered for the method according to the invention.

[0034] More particularly, a method according to the invention comprises determining the nucleotide sequence of at least 20% of the 16S rDNA and / or the 18S rDNA of a plurality of microorganisms derived from the biological sample. Even more particularly, a method according to the invention comprises determining the nucleotide sequence of at least 20% of the 16S rDNA or the 16S rRNA of a plurality of microorganisms derived from the biological sample and the nucleotide sequence of at least 20% of the 18S rDNA or the 18S rRNA of a plurality of eukaryotic cells derived from the biological sample.

[0035] By "a fragment of at least 20%" is meant a fragment of at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, less than 98%, at least 99% or 100% of the nucleotide sequence considered.

[0036] More particularly, in a method according to the present invention, during step c) of reconstructing at least one nucleotide sequence, at least 70% of the length of the gene expressing the 16S rRNA and / or at least 70% of the length of the 16S rRNA is reconstructed. The length of a 16S rDNA gene being approximately 1500 base pairs on average, a nucleotide sequence of at least 70% of the length of the gene comprises approximately 1050 base pairs, on average.

[0037] According to another particular aspect, a method according to the invention comprises an additional step consisting of the reclassification of a biological sample, following a first classification of this sample. This method is used when longitudinal samples are taken from a subject before the appearance of the pathology and at least 3 samples taken are analyzed to predict the phenotype of the subject. The subjects presenting an inconsistent first classification between the samples and for which at least 3 samples are available are identified. Then, the number of samples in each of the phenotypic groups (pathology or absence of pathology) is calculated and the sample (or samples) of the minority group is automatically reclassified in the phenotypic group having the largest number of samples attached to this phenotypic group.The predicted phenotype of the subject will then correspond to the phenotypic group having the largest number of samples attached to this phenotype. For example, if among three samples of a subject, two samples No. 1 and No. 2 are considered as predictive of the pathology and one sample No. 3 is considered as predictive of the absence of pathology, sample No. 3 is reclassified as predictive of the pathology, the subject is therefore reclassified among the subjects for which the pathology is predicted.

[0038] By "reclassification" of a sample is meant the re-examination of the prediction of pathology or non-pathology for a sample among the samples taken longitudinally from a subject and used to predict the subject's phenotype.

[0039] By "subject" is meant an individual, in particular a human being or an animal, preferably a mammal, and preferably a human subject. According to a particular embodiment of a method according to the invention, the subject is chosen from: adults, children, children in the perinatal period, newborns and premature newborns. According to a particular aspect of the method of the invention, the subject is a human newborn.

[0040] A biological sample used in a method according to the invention is chosen from: a sample taken from the digestive system and a sample taken from the excretions of the subject, in particular a sample taken from the stools of the subject. Said sample is taken in a conventional and well-known manner by a specialist. A given biological sample comprises a community of microorganisms referred to as "microbiota".

[0041] Among the microbiota hosted by a human subject, we can distinguish the cutaneous microbiota, the mucosal microbiota, the pulmonary microbiota, the oral microbiota, the vaginal microbiota and the microbiotas of the digestive system (oral or salivary microbiota, stomach microbiota, small intestine microbiota, colonic microbiota, anal microbiota). The microbiota present in the stools, or fecal microbiota, corresponds to all the microorganisms found in the stools following transit through the digestive system of a subject, which may reflect the intestinal microbiota in the broad sense with a closer proximity to the colonic microbiota. Transient microorganisms can also be found in this microbiota. The term "microbiome" refers to all the genomes carrying the genes hosted by the microorganisms constituting the microbiota.The microbiome can also be considered as the set of microorganisms including their genomes in a particular biological environment such as the colon. A "subject" is understood to mean an individual, in particular a human being or an animal, preferably a mammal.

[0042] The term "digestive system" refers to all the organs of multicellular animals that take in food, digest it to extract nutrients, and excrete waste in the form of fecal matter. The organs of the human digestive system include the mouth, salivary glands, pharynx, esophagus, stomach, pancreas, liver, gallbladder, bile duct, small intestine, and large intestine. The large intestine includes the ascending colon, transverse colon, sigmoid colon, and rectum. The term "excretion" refers to unusable or toxic waste that is expelled by the subject, such as urine, feces, or stools, or secretory products such as bile or saliva.

[0043] “Pathology” means a disease, a risk of an adverse event, a biological imbalance or discomfort. By "condition or pathology of the digestive system" is meant a condition or pathology affecting at least one organ chosen from: the mouth, salivary glands, pharynx, esophagus, stomach, pancreas, liver, gallbladder, bile duct, small intestine and large intestine. The large intestine includes the ascending colon, transverse colon, sigmoid colon and rectum. According to a particular aspect of the method of the invention, said pathology is an intestinal pathology.

[0044] According to a particular embodiment, the invention relates to an in vitro method for predictive diagnosis of a digestive condition or pathology, from a biological sample taken from the digestive system and / or from the excretions of a subject. Even more particularly, among said digestive conditions and pathologies, mention may be made of: digestive cancers, i.e. affecting at least one of the organs of the digestive system, chronic inflammatory diseases, such as in particular Crohn's disease, ulcerative colitis, irritable bowel syndrome and celiac disease.

[0045] According to a particular embodiment, the invention relates to an in vitro method for predictive diagnosis of an extra-digestive condition or pathology, from a biological sample taken from the digestive system and / or from the excretions of a subject.

[0046] By "extra-digestive condition or pathology" is meant a condition or pathology not directly affecting an organ of the digestive system but one of the consequences of which is likely to directly or indirectly affect the microbiota of the digestive system and vice versa. Among the extra-digestive, or non-digestive, conditions and pathologies for which a predictive diagnosis can be carried out by a method according to the invention, mention may be made of: diabetes, obesity, cardiovascular diseases, metabolic diseases, liver diseases, kidney diseases, urogenital diseases, pulmonary diseases, joint diseases, muscle diseases, inflammatory diseases, asthma, allergies, arthritis, neurodegenerative diseases (Parkinson's, Alzheimer's, etc.), psychiatric diseases, behavioral diseases, all types of cancer for all types of organs.

[0047] According to a particular aspect of the method of the invention, said pathology is a digestive pathology of a subject chosen from: children, infants (children beyond their first month of life and up to the age of 24 or 30 months) and newborns (children less than 28 days old according to the definition of the World Health Organization), said newborns being born at term, i.e. between the 37th week and the end of the 40th week of amenorrhea, or premature, i.e. born before the 37th week of amenorrhea.

[0048] According to a particular aspect of the method of the invention, said pathology is an intestinal pathology chosen from enterocolitis, said pathology is more particu particularly necrotizing enterocolitis (NCE). "Necrotizing enterocolitis" means a disease characterized by inflammation and necrosis of the intestinal mucosa. More particularly, in a method according to the invention, said pathology is necrotizing enterocolitis in infants. Even more particularly, in a method according to the invention, said pathology is necrotizing enterocolitis in premature infants.

[0049] According to a particular aspect of the method of the invention, step e) is carried out from the characteristics determined during step d) and at least one clinical data characteristic of the subject. By "at least one clinical data" is meant one, two, three, four, five, six, seven, eight, nine, ten or more than ten clinical data characteristic of the subject.

[0050] Among the clinical data of the subject that can be used for implementing a method according to the invention for determining a predictive diagnosis of a pathology, in particular a pathology of the newborn, we can cite clinical data of the subject himself and, in the case of a newborn, clinical data relating to his mother, among which: - the actual age of the subject at whom the sample was taken, in number of days of life (DOL) - birth weight, - gestational age at birth, - the method of birth (vaginal or caesarean), - the gender of the subject (masculine, feminine), - the dosage of blood components or markers of the subject or the mother, - the dosage of fecal components or markers of the subject or the mother, - the presence of at least one other pathology, - the administration of medical treatment, - the ethnicity of the subject's mother, - the diet of the subject's mother, - the lifestyle of the subject's mother (physical activity, consumption of alcohol, tobacco, drugs), - the presence of at least one pathology in the subject's mother, - the administration of at least one medical treatment to the subject's mother, - a combination of at least two of these clinical data.

[0051] In a method according to the invention, the age of the mother can be defined in number of years or by her belonging to an age group. More particularly, the age of the mother can be attributed to one of the following two groups: "less than 35 years" and "equal to or greater than 35 years".

[0052] By "ethnicity" we mean a group of people who are brought together by a certain number of characters. In a method according to the invention, the characteristic “ethnicity” is notably chosen from the group consisting of: “African-American”, “American-Indian”, “Black”, “White”, “Caucasian”, “Hispanic”, “Asian”, “Multi-ethnicity”.

[0053] Clinical data are encoded as follows: categorical data (such as gender and mode of birth in the case of newborns) are converted into vectors using "one hot encoding", i.e. all elements of the vector are converted to 0 except the categorical variable which is converted to 1. Continuous value data (actual age, birth weight and gestational age in the case of newborns) are transformed into a discrete variable by creating a set of contiguous intervals ("bins" in English) that cover the range of values ​​of the variable. The clinical data "day of life" is discretized into intervals with an increasing step of 9 (from 0 to 99 days) and 99 (100 to 499 days). A time step of 1 could also be considered over the first 3 weeks of life where the pathology most frequently appears.The clinical data "weight" is discretized into intervals with an increasing step of 99 (from 500 to 2899 grams). The weight of the children can also be monitored if necessary by intervals of 9 throughout the first 3 weeks of life until the possible appearance of the pathology. The gestational age at birth is converted into factors due to the limited number of values.

[0054] The present invention more particularly relates to an in vitro method for predictive diagnosis of necrotizing ulcerative enterocolitis in a newborn, from a biological sample taken from the stools of said newborn, said method comprising the following steps: (a) isolation of nucleic acid from a plurality of microorganisms present in said biological sample, b) determination by sequencing, more particularly metagenomics, even more particularly by direct metagenomic sequencing, of the nucleotide sequence of at least one fragment of genes expressing 16S rRNA and / or a fragment of genes expressing 18S rRNA of a plurality of microorganisms to generate a plurality of nucleotide sequences, c) organization of a plurality of nucleotide sequences determined during step b) to reconstruct the nucleotide sequence of at least one fragment of the genes expressing the 16S rRNA and / or a fragment of the genes expressing the 18S rRNA of a plurality of microorganisms, d) from the results of step c), determining the identity of a plurality of microorganisms present in the sample, the taxonomic classification, the and the relative abundance of a plurality of microorganisms present in said sample, and determination of at least one clinical data characteristic of said newborn, and (e) from the characteristics determined during step (d), determination, by a previously trained classification model, of a predictive diagnosis of ECUN or absence of ECUN, said at least one clinical data being chosen from: age in number of days since birth, birth weight, gestational age, mode of birth, gender, ethnicity of the mother, age of the mother, diet of the mother, lifestyle, pathologies and treatments, the result of the dosage of blood or even fecal components or markers of the child or even the mother, the administration of medical treatment to the child or even the mother, the presence of at least one other pathology in the child or a combination of several of these characteristics.

[0055] According to a particular aspect of a method according to the invention, step e) of determining a predictive diagnosis of a state or pathology is carried out by means of a previously trained classification model.

[0056] By “classification model” we mean in particular a classification model comprising: - a previously trained neural network, in particular during supervised learning, in particular a deep neural network, - a machine learning algorithm and - a training dataset. The classification model may consist of a computer program. The computer program may be written in any computer language, such as for example: C, C++, JAVA, Python, etc. The said classification model therefore potentially executes a technical function, consisting here of steps of the classification process. The execution of the said program by a computer produces a digital object having technical characteristics.

[0057] By "previously trained" is meant a process allowing the neural network to learn to associate in a weighted manner a microorganism or a group of microorganisms and at least one relevant clinical data item with a phenotype (pathology or non-pathology). In the case of a method relating to a digestive pathology, said at least one relevant clinical data item is in particular related to the pathology considered. In the case of a method relating to a newborn subject or a premature newborn, said at least one relevant clinical data item is linked to the clinical data of the newborn and possibly of its mother, or with prematurity.

[0058] For the implementation of a method according to the invention, step e) of determining a predictive diagnosis of a state which may lead to a pathology or a pa Theology is performed using a pre-trained classification model comprising an algorithm, said algorithm being chosen from the group consisting of: a neural network (NN), a decision tree, K-nearest neighbors (KNN), a random forest (RF), a naive Bayesian classification (NF), an extreme gradient boosting (XGBoost) algorithm, a logistic regression and a support vector machine (SVM).

[0059] In one embodiment of the method according to the invention for determining a predictive diagnosis of a condition or pathology, the method further comprises a step of determining at least one first profile, also referred to as a "signature", of a plurality of microorganisms, said at least one first profile being characteristic of a prediction of the presence of said condition or pathology. By "prediction of the presence of said condition or pathology", is meant a prediction of the presence of said condition or pathology greater than 50%, preferably a probability greater than or equal to 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or equal to 100%.

[0060] According to this particular embodiment, from the results of step d) of a method according to the inventions, at least one first profile of a plurality of microorganisms statistically associated with a prediction of the presence of said state or said pathology is determined. In other words, according to this particular embodiment a method according to the invention makes it possible to establish at least one “signature” of a plurality of microorganisms statistically associated with a prediction of the presence of said state or said pathology.

[0061] A first profile of microorganisms statistically associated with a prediction of the appearance and development of ECUN, obtained by a method according to the invention, is characterized in particular by the presence of microorganisms of the genus: - Unclassified Enterobacter, - Unclassified Enterobacteriaceae, - Enterococcus faecalis, - Unclassified Klebsiella, - Haemophilus parainfluenzae, - Enterococcus durans and - Enterobacter carcinogenus. Indeed, these microorganisms are present or present in greater quantity in biological samples statistically associated with the prediction of the presence of ECUN.

[0062] In one embodiment of the method according to the invention for determining a predictive diagnosis of a condition or pathology, said method further comprises a step of determining at least one second profile of a plurality of microorganisms characteristic of a predictive diagnosis of absence of a condition or pathology. By "prediction of absence of said condition or pathology" is meant a prediction of absence of said condition or pathology greater than 50%, preferably a probability greater than or equal to 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or equal to 100%.

[0063] According to this particular embodiment, from the results of step d) of a method according to the invention, at least one second profile of a plurality of microorganisms statistically associated with a prediction of absence of ECUN is determined.

[0064] Said at least one second profile of a plurality of microorganisms statistically associated with a prediction of absence of ECUN, obtained by a method according to the invention, is characterized in particular by the presence of microorganisms of several species of Lactobacillus associated with non-ECUN cases. Indeed, these microorganisms are present or present in greater quantity in the biological samples statistically associated with a prediction of absence of ECUN. Other microorganisms may be cited, such as: Bifidobacterium, Bifidobacterium longum, Bacteroides, Bacteroides fragilis, Lactobacillus casei.

[0065] The present invention also relates to a classification model previously trained on a training data set to determine, according to a method according to the invention, a predictive diagnosis of a condition or pathology.

[0066] More particularly, the present invention also relates to a classification model previously trained on a training data set to determine, according to a method according to the invention, a predictive diagnosis of the presence of ECUN or the absence of ECUN.

[0067] According to a third aspect, the invention relates to the use of a previously trained classification means, according to the invention, for determining a state or pathology.

[0068] According to a particular aspect of this third aspect, the invention relates to the use of a classification means according to the invention for determining a predictive diagnosis of ECUN or absence of ECUN.

[0069] According to a fourth aspect, the invention relates to the use of a computer tool capable of ensuring the precise determination of the structures of microbial communities. Said computer tool makes it possible in particular to determine the identity of the microorganisms and their relative abundance, for determining a predictive diagnosis of a condition or pathology. An example of such a computer tool is RiboTaxa (Chakoory et al., 2022).

[0070] More particularly, according to this fourth aspect, the invention relates to the use of a computer tool capable of ensuring the precise determination of the structures of microbial communities for the determination of a predictive diagnosis of ECUN or absence of ECUN.

[0071] The invention therefore also relates to an in vitro method for predictive diagnosis of a condition or pathology, from a biological sample taken from the digestive system and / or from the excretions of a subject, said method comprising the following steps: (a) isolation of nucleic acid from a plurality of microorganisms present in said biological sample, b) determination by sequencing of the nucleotide sequence of at least one nucleic acid chosen from: a gene fragment expressing 16S rRNA, a gene fragment expressing 18S rRNA, a 16S rRNA fragment and an 18S rRNA fragment, of a plurality of microorganisms, to generate a plurality of nucleotide sequences, c) organizing a plurality of nucleotide sequences determined during step b) to reconstruct the nucleotide sequence of at least one gene fragment expressing 16S rRNA, one gene fragment expressing 18S rRNA, one 16S rRNA fragment and one 18S rRNA fragment of a plurality of microorganisms, d) from the results of step c), determination of the identity, taxonomic classification, and relative abundance of a plurality of microorganisms present in said sample using a computer tool capable of ensuring the precise determination of the structures of microbial communities, such as for example RiboTaxa, and e) from the characteristics determined during step d), determination, by a previously trained classification model, of the predictive diagnosis of a condition or pathology.

[0072] The present invention is further explained by the following figures and examples.

[0073] [Fig.l] shows an overview of the steps followed for the general method of predicting a condition or pathology, starting with training the neural network from metagenomic sequencing data of a mixed population comprising healthy subjects and subjects with pathology, followed by the method of predicting the risk of a pathological condition using the trained and optimized model.

[0074] [Fig.2] represents the final structure of the model optimized to predict ECUN.

[0075] [Fig.3] represents the rate of true positives (on the ordinate) as a function of the rate of false positives (on the abscissa), the AUC is equal to 0.987.

[0076] [Fig.4] represents the precision (on the ordinate) as a function of the sensitivity (on the abscissa), the AUC is equal to 0.992.

[0077] [Fig.5] illustrates the reclassification of samples following an initial classification. The unlabeled circle on the left represents the actual phenotype of the infant. Samples from infants without pathology are shown in dark gray and samples from ECUN infants in light gray. Each labeled circle represents a sample collected from each of the infants and the numbers inside the circles correspond to the day of sampling (in days of life). The color of these circles represents the phenotype predicted by the neural network according to the same color code as the unlabeled circles. The single square represents the samples that were reclassified into the “control” group and the double square represents the samples that were reclassified into the “ECUN” group.

[0078] [Fig.6] represents the 20 input features of the neural network contributing most to the prediction of ECUN or non-ECUN phenotypes. Examples

[0079] Example 1: Predictive diagnosis of ECUN using deep neural networks Materials and methods

[0080] For data collection, a meta-analysis was performed using the keywords (in English): “premature infants” AND (“stool microbiome” OR “intestinal mi-crobiome”) AND “shotgun metagenomics” AND “necrotizing enterocolitis”. This search led to a total of eleven articles that were carefully selected to verify that the sequencing method used was the “shotgun” method implemented on the Illumina platform and to exclude studies applying only the metabarcoding method. In addition, the studies had to be prospective studies that had carried out longitudinal stool samples in a population of preterm newborns housed in a neonatal intensive care unit.The studies were to include longitudinal stool samples from newborns who had developed NEC, before the onset of the pathology, designated as "pre-NEC samples", and longitudinal stool samples from premature newborns who had not developed NEC (control). The inventors also selected studies that included clinical information such as: mode of birth (vaginal or cesarean), gender (male-female), gestational age (in weeks), actual age (in days of life) and birth weight (in grams). Finally, two studies were retained.

[0081] Raw data and metadata from Masi et al. (2021) were downloaded from from the European Nucleotide Archive (ENA) under BioProject PRJEB39610 (n = 524; 974.51 GB). In addition to their own cohort, Olm et al. (2019a) also used sequencing data from different previously published datasets. All raw data and metadata used on the Olm cohort were downloaded from the Sequence Read Archive (SRA) under the BioProjects: PRJNA294605 (n = 141; 596.53 GB), PRJNA417343 (n = 184; 152.21 GB) PRJNA396794 (n = 295; 1.35). Tb), PRJNA376566 (n = 358; 905.22 Gb) and SRA study SRP052967 (n = 60; 114.21 Gb).

[0082] The pre-existing cohorts used for deep neural network training represent a total of 524 and 1,038 metagenomic data extracted from Masi et al. (2021) and Olm et al. (2019a), representing 1,305 control samples (from 160 infants) and 257 ECUN samples (from 48 ECUN infants), respectively. No samples collected after the onset of ECUN were analyzed. Five clinical metadata features were collected and reported in both studies, such as phenotypes (control, ECUN), mode of birth (vaginal, cesarean), gender (boy, girl), gestational age at birth (in weeks), day of life (in days), and newborn birth weight (in grams), and infant identification. Only Masi et al. (2021) reported that their entire cohort received probiotics (Lactobacillus acidophilus, Bifido-bacterium infantis, and B. bifidum).

[0083] Data preprocessing with RiboTaxa (Chakoory et al., 2022) was used to obtain taxonomic profiles from raw metagenomic datasets through 16S and 18S rRNA sequence reconstruction using the SILVA (Quast et al., 2013) SSU 138.1 NR99 database. For each cohort, raw reads were provided as input to RiboTaxa. Reads were processed to remove Illumina adapters, known Illumina artifacts, and to trim both ends of reads when the base quality score was below Q20. Resulting reads containing more than one “N,” or with quality scores below 20 averaged across the read, or a length below 60 bp, were discarded.

[0084] For the reconstruction of the gene expressing 16S / 18S rRNA, the default parameters were used, except for the following parameters (Masi cohort: --max_read_length = 151, — insert_mean = 144, —insert_stddev = 100; Olm cohort: --max_read_length = 301, — insert_mean = 120, —insert_stddev = 100), which depend exclusively on the sequencing length of the input data. For each cohort, the parameters — max_read_length represented the size of the longest read of the input data and — insert_mean and —insert_stddev were estimated using the script Mean_size.py (https: / / gist.github.com / timoast / af73c0e9fac00187ee49). The Reconstructed 16S / 18S rDNA sequences were then classified at different taxonomic levels, from domain to species and after discarding human eukaryotic sequences, relative abundances were calculated by RiboTaxa.

[0085] All output taxonomy tables were grouped into a single table containing all species-level profiles using RiboTaxa's RiboTaxa_group_taxonomy.sh script.

[0086] For each sample, the taxonomic relative abundances provided by RiboTaxa were normalized to avoid the influence of highly abundant taxa by using a standard scale to normalize each abundance value via the transformation below, called min-max normalization:

[0087] [Math.l] X — (x — Xmin) / (Xmax ~

[0088] where: x is the original data, x' is the normalized data xmin and xmax are respectively the minimum and maximum values ​​of the original value (abundance). The above equation is a linear transformation that preserves all the abundance ratios of the original data after normalization.

[0089] Data discretization and vectorization were performed as follows: For model training, the normalized species abundance profiles. In addition, clinical information on phenotype (control, ECUN), mode of delivery (vaginal, cesarean), sex (boy, girl), gestational age at birth (in weeks), day of life (in number of days, "Day of Life" or "DOL") and birth weight (in grams) of the newborn were the main inputs to the deep neural network. Apart from phenotype, mode of birth and sex which were discrete variables, other clinical characteristics were continuous values.To better handle these high-dimensional data, they were transformed by discretization, which is the process of transforming a continuous-valued variable into a discrete variable by creating a set of contiguous intervals (or bins) that span the range of the variable's values. Grouping numeric features into interval-based groups is beneficial for classification and can significantly improve model performance. The "DOL" metadata was discretized into bins from 0 to 499 with an increasing step size of 9 (0-99) and 99 (100-499). Infant weight was grouped into bins from 500 to 2,899 with an increasing step size of 99. Gestational age was converted to factors due to the limited number of values.

[0090] The next step was to apply a one-hot encoding technique to the categorical data using LabelEncoder from the scikit-learn library (Pedregosa et al., 2011). Thus, the discrete values ​​were vectorized, that is, all elements of the vector are converted to 0 except the categorical variable, which is converted to 1. After this step, a data set comprising 47 categorical values ​​and 1,282 numerical values ​​(normalized microbial abundances) for each sample was obtained.

[0091] A fully connected deep neural network model was implemented using the Python programming language and dedicated libraries such as scikit-learn, Tensorflow (Abadi et al., 2016) and Keras (Keras, 2023).

[0092] For training and optimization of the neural network, the Rectified Linear Unit (ReLU) activation function was used for all hidden layers. Activation functions play an important role in training neural networks. They provide the necessary nonlinearity for the model to learn complex representations. The neuron dropout technique on each hidden layer (0.1-0.5) was also employed to mitigate overfitting in neural networks. Neuron dropout is a learning method that involves randomly removing neurons during model training, with the removed nodes being excluded from subsequent steps.

[0093] For training and optimizing the neural network, the activation function of the output layer uses the Softmax function to assign a value based on a probability between 0 and 1 to each class (ECUN, control). The loss function the cross-entropy between the target value and the predicted value was used, and this loss was optimized over epochs (number of times the full dataset will be propagated through the neural network) using the Adam optimizer (Kingma and Ba, 2017) with variable learning rates, ranging from 0.0001 to 0.01. The number of hidden layers was varied from 1 to 3, the number of epochs from 1 to 40, and the number of neurons in the first hidden layer from 32 to 512 with an increasing step size of 32. The number of neurons in the hidden layers was set to half that of the previous layer, except for the first hidden layer. These optimizations were implemented using Keras (https: / / github.com / keras-team / keras-tuner). A total of 584,482 hyperparameters were tested. The optimal model was selected based on a low loss in the validation value, which allowed identifying the best hyperparameters of the deep neural network ([Fig.2]). In addition, a stopping criterion was used whereby training was stopped if the validation value did not improve after 10 epochs.

[0094] The model training was carried out as follows: a total of 1,562 samples were divided into training and test sets in a ratio of 8:2 to obtain 80% training data and 20% test data. For training, K-Fold cross-validation was applied with the training data. These are divided into K subsets of almost equal size; Kl subsets being used for model training and the remaining subset for validation of the produced model. The objective was to build and adjust a trained model on different subsets of training data. In this way, K models are built, each time with a redistribution of the K subsets and the definition of new hyperparameters. The best combination of hyperparameters for each model is selected by averaging the accuracy metric of the K models. This approach not only allowed more efficient use of the available training data, but also increased the reliability of the prediction model. Once the best combination of hyperparameters was determined, a final classification model was trained using the training dataset, which was then tested on the test dataset.

[0095] The following table 1 brings together the characteristics of the deep neural network thus obtained.

[0096] [Tables 1] Hyperparameters Range Optimized value Training rate 0.0001-0.01 0.01 Number of hidden layers 1-3 3 Number of neurons in the first hidden layer 35-512 416 Epoch 1-40 18 Dropout rate 0.1-0.5 0.4

[0097] The model performance was estimated on the test dataset by comparing the true phenotype of the sample with the predicted phenotype. If the model correctly classifies an ECUN sample, it is considered a true positive (TP for True Positive), otherwise it is a false negative (FN for False Negative). On the other hand, if the model correctly classifies a non-ECUN sample (Control), it is considered a true negative (TN for True Negative), otherwise it is a false positive (FP for False Positive).A multi-metric evaluation was applied due to class imbalance to measure the model performance using accuracy, true positive rate (TPR) and true negative rate (TNR), receiver operating characteristic (ROC) / AUROC area under the curve (AUC), and precision-recall AUC (PR-AUC).

[0098] Accuracy is the ratio of the number of correct predictions to the total number of predictions, it is calculated as follows:

[0099] [Math.3] Accuracy = (TP+TN) / (TP+FP+TN+FN)

[0100] The TPR, also called sensitivity, is the ratio of correctly classified ECUN samples (TVN-ECUN), it is calculated as follows:

[0101] [Math.4] TPR-ECN =TP / (TP+FN)

[0102] The TNR, also called specificity, is the ratio of correctly classified control / non-ECUN samples (TNR-non-ECUN) and is calculated as follows:

[0103] [Math.5] TNR-non-ECN =TN / (TN+FP)

[0104] Finally, AUROC measures the false positive (FP) rate on the TPR and PR-AUC measures the sensitivity over accuracy (ratio of TPs to the total number of TPs and FPs). AUCs were calculated in python using the scikit-learn package (Pedregosa et al., 2011) and plotted using matplotlib (Hunter, 2007) (v3.1). The 95% confidence intervals (CIs) of the AUCs were estimated using the bootstrap method (Efron and Tibshirani, 1994) with 1000 iterations. ROC curves and Sankey plot were generated using matplotlib and plotly (v5.15.0) respectively to compare the model and evaluate the final model performance on the test dataset.The Mann-Whitney U test was used to determine significant changes in AUROC and PR-AUC between the model trained from metadata and microbiota microorganism relative abundance data, and without metadata (microbiota only); a p value was considered significant when <0.05.

[0105] Sample reclassification based on longitudinal infant sampling was performed as follows: intensive longitudinal data can be used to explore the evolution of the microbiota over time. In the present study, the inventors took advantage of the longitudinal newborn sampling provided by the Masi and Olm cohorts to reclassify samples misclassified by the deep neural network. Infants for whom phenotype prediction was inconsistent across samples and who had at least 3 stool samples in the test dataset were identified. The number of samples in each phenotypic group was calculated and the sample(s) in the minor phenotypic group were automatically reclassified to the phenotypic group with the largest number of samples.To ensure that samples were correctly reclassified, the true phenotype of newborns and TPR-ECUN and TNR-non-ECUN were recalculated. Finally, a lollipop plot was generated to visualize the reclassification of longitudinal samples ([Fig. 5]) using the ggpubr package (v0.4.0).

[0106] SHAP: Models can be interpreted by calculating the importance of the data input data related to the model's classification performance. In this study, the importance of input elements (metadata, microorganisms) was calculated using SHAP. The DeepLIFT function (Lundberg and Lee, 2017; Shrikumar et al., 2019), also known as SHAP's DeepExplainer, is a method for decomposing the output of a neural network (prediction) by assigning contribution values ​​to each data item in the neural network's input. This function highlights the input data (metadata, microorganisms) with the most weight in predicting a phenotype.

[0107] To further evaluate the performance of the optimized model, a cohort of 19 preterm infants was constructed, including 9 newborns who developed ECUN, a total of 56 stool samples were collected longitudinally.

[0108] This study was approved by the Ethics Committee of CPP-Sud-Est VI (protocol code 2021 / CE 26, approval date is May 4, 2021). The CORTECs cohort aims to address prenatal and postnatal risk factors for ECUN. All prematurely born children hospitalized in the Neonatal Intensive Care Unit (NICU) of the Clermont-Ferrand University Hospital were proposed to enter the cohort. Written informed consent was obtained from the families of study participants prior to enrollment. Infant stools were collected daily during their NICU stay, between May 2021 and June 2022. Stools were collected in a diaper using a sterile loop and then dispensed into eNAT buffer (Copan) before being briefly held at 4°C. Samples were stored at -80°C until DNA extraction.

[0109] Cases of NEC were identified by physicians based on systemic and abdominal findings and radiographic features. They were stratified according to disease severity according to Bell stages. Cases of NEC were matched to a control neonate (two to one case) who did not develop NEC. Case-control matching was based on gestational age at delivery, mode of delivery, sex, birth weight, and pre- and postnatal antibiotics. For each NEC infant, available samples were selected within a 1-week window before the onset of NEC, and samples from corresponding control cases were matched according to the age of the NEC patient.

[0110] Genomic DNA was extracted using the standard operating protocol for fecal samples (Protocol H) recommended by the International Human Microbiome Standards (IHMS SOP 07 VI). DNA quality was assessed using the Nanodrop 2000 Fluorometer (Thermo Scientific) and the Agilent 4150 Ta-peStation system with Genomic DNA ScreenTapes (Agilent). DNA quantity was assessed using the Qubit 3 Fluorometer (Invitrogen) with the Qubit dsDNA High Sensitivity Assay Kit (Invitrogen). Hybridization capture of the gene expressing 16S rRNA and sequencing data processing: Capture probes were designed to target the 16S rRNA-expressing gene (Gasc et al., 2016). Sequencing libraries were produced for each sample using the Nextera XT Library Preparation Kit. The gene capture experiment was performed according to the protocol described by Ribière et al. (2016) and Comtet-Marre et al. (2023). Briefly, biotinylated RNA capture probes were obtained by in vitro transcription. 500 ng of libraries were mixed with 2.5 pg of salmon sperm DNA and incubated with 500 ng of biotinylated probes in hybridization buffer for 24 h at 65 °C. Probe / target heteroduplexes were captured using 500 μg of streptavidin-coated paramagnetic beads (Dynabeads M-280 Streptavidin, Invitrogen).Beads were collected using a magnetic stand (Ambion), washed once with 500 pL of 1 x SSC / 0.1% SDS buffer, and then three times with 500 pL of 0.1 x SSC / 0.1% SDS buffer preheated to 65°C. Captured DNA fragments were eluted with 50 pL of 0.1 M NaOH and transferred to a sterile tube containing 70 pL of 1 M Tris-HCl buffer pH 7.5. Captured DNA was amplified by PCR with 25 cycles using primers complementary to Illumina adapters. To increase enrichment efficiency, a second round of capture was performed. Captured DNA was then sequenced on the Illumina MiSeq 2 x 300 bp platform. Raw sequencing data were processed using the RiboTaxa pipeline to reconstruct 16S rRNA-expressing genes, obtain the taxonomic affiliation of the reconstructed sequences (down to the species level), and determine relative species abundances, as described above.

[0111] Prediction using the optimized model: Before prediction, species abundances were normalized by a min-max scale. Continuous variables in the metadata were discretized, and all variables were encoded using one-hot encoding. Since the hybridization gene capture approach allows for a more refined description of the microbial community compared to shotgun metagenomics, all species that were not used during training were excluded. The optimized model was then invoked. The relative abundance table of microorganisms concatenated with relevant clinical data was input into the trained model, and its predicted value was recorded. Finally, each prediction was compared to the true value (control or ECUN) to calculate the number of correct predictions.SHAP plots for each prediction were also generated using matplotlib to explain the different predictions. This cohort also allowed for reclassification of longitudinal samples from the same infant using the same approach as described above. Results

[0112] Direct (“shotgun”) sequencing was selected to assess the microbiota at high resolution (species level) to improve the identification of specific microbial signatures associated with specific disease states. 1,562 metagenomic stool samples were analyzed using RiboTaxa (Chakoory et al., 2022). This uniformity allows a single model to accommodate data from diverse study protocols. The controlled, high-quality species-level relative abundance profiles and five clinical data were used to train a deep neural network to predict the risk of ECUN.

[0113] To facilitate comparison of metadata between the “Masi” and “Olm” cohorts, only metadata at the intersection of the two studies was used. Clinical and demographic data of the patients are presented in Table 2.

[0114] [T ables 2] Masi Cohort Olm Cohort Control (n = 34) Z 't QH LU c. Control (n = 126) ECUN (n = 34) Gender Male (%) 36.9 91.1 50.8 41.6 Female (%) 63.1 8.9 49.2 58.4 Mode of birth Vaginal (%) 75.1 51.9 30.5 &7.6 Cesarean 24.9 48.1 69.5 82.4 Mean gestational age at birth [SD] 25.0 [1.85] 24.8 [n 45] 28.4 [2.23] 28.9 [2.47] Mean number of days alive [SD] 59

[74] 21

[14] 23

[16] 17 [9] Mean birth weight (g) 748 704 1194 970

[0115] To standardize microbial characterization between the two independent cohorts, all metagenomic sequencing data were reanalyzed using RiboTaxa. After discarding human contaminants, we were left with only bacterial and archaeal 16S rDNA sequences representing the fecal microbiota of 208 preterm infants. Taxonomic profiling of the two cohorts led to the identification of a total of 1,282 unique species: Masi: 1,237 species; Olm: 1,343 species.

[0116] Microbial diversity analysis revealed higher species diversity in non-ECUN samples with a total of 1,238 species identified compared to ECUN samples (539 species). RiboTaxa identified 495 common species between the two groups and, interestingly, 743 species were unique in non-ECUN samples, compared to ECUN samples with 44 unique species in ECUN. Significant differences (p < 0.05) in α and [3 diversity further confirm a distinct microbial composition between preterm control and ECUN infants.

[0117] Unclassified Enterobacteriaceae, Unclassified Escherichia-Shigella, Ente- Unclassified Enterobacter and unclassified Klebsiella were the most abundant species overall. Although the abundance of unclassified Klebsiella and unclassified Escherichia-Shigella did not differ significantly (p > 0.05) between the two groups, the mean relative abundances of unclassified Enterobacter and unclassified Enterobacteriaceae were significantly higher in ECUN samples than in preterm controls (p < 0.05).

[0118] Furthermore, the mean abundance of Bifidobacterium breve, unclassified Bifidobacterium and unclassified Citrobacter were significantly higher in control samples while unclassified Staphylococcus were significantly more abundant in ECUN samples (p < 0.05). All these highlighted microbial species, although informative for a better understanding of ECUN risk factors, were not statistically strong enough to establish reliable and universal characteristics that could represent a prediction of ECUN risk.

[0119] ECUN is often distinguished by a complex alteration in microbiome composition and diversity, encompassing altered interactions between dominant and rare microorganisms. To better differentiate complex microbial interactions in ECUN and non-ECUN newborns, a deep neural network was developed to predict ECUN risk before disease onset. The deep neural network was trained using 1,402 different features (1,355 species and 47 clinical data: 10 gestational age groups, 18 weight groups, 15 DOL, 2 birth modes, and 2 sex groups). It was decided to retain all species detected in all samples, instead of applying feature selection before training to retain inter-individual variations between infants. The objective was to obtain a high-performance model while maintaining inter-variability between individuals.A comprehensive performance evaluation system that separates the test set (20%) and the training set (80%) was performed during the hyperparameter optimization phase using 10-fold cross-validation. A total of 584,482 trainable parameters were tested and the optimal hyperparameter setting for the final model had 448 units (neurons) in the 1st hidden layer and a total of 3 hidden layers with the number of units set to half that of the previous layer. The optimal learning rate was 0.01 and the model was best fitted at epoch 27. Due to the relatively small number of input features, the model training was performed on: i861inux32, 4.0 GB RAM x 8 cores (32.8 GB total), without GPU and completed in 4 min 51s.

[0120] The evaluation of the final model was carried out on the test set composed of 313 samples (140 infants) associated with the five metadata (mode of birth, sex, gestational age, birth weight, and DOL). The model identified a total of 297 correctly predicted samples (125 infants), yielding an accuracy of 94.9%. Among the 260 preterm control samples (112 infants), the model predicted 249 non-ECUN samples (101 infants) and misclassified 11 samples as ECUN, producing a TNR-non-ECUN score of 95.8%. On the other hand, among 53 ECUN samples (28 infants), the model correctly identified 48 ECUN samples (24 infants) and misclassified 5 samples as non-ECUN, yielding a TPR-ECUN of 90.6%. In repeated testing of the deep neural network, we demonstrated an AUROC of 0.987 ± 0.01 ([Fig.3]), suggesting a good balance between sensitivity and specificity and a PR-AUC value of 0.992 ([Fig.4]).

[0121] In this study, the longitudinal sampling of newborns provided by Masi et al. (2021) and Olm et al. (2019a) was used to apply a unique sample reclassification approach to increase the phenotypic correct prediction rate. In case the sample prediction was inconsistent among serial samples (in the test set) belonging to the same infant, the number of samples in each phenotypic group was calculated and the sample(s) in the minor phenotypic group were automatically reclassified to the phenotypic group with the largest number of samples. In this study, 16 samples (from 16 infants) out of 313 tested samples were misclassified by the deep neural network, among which 6 samples belonged to 6 infants (2 controls and 4 ECUN) for whom at least three serial samples were available.To reclassify the 6 samples, 22 longitudinal samples belonging to the 6 infants were retrieved and sorted by sampling day. Using this technique, 2 samples were reclassified into the ECUN group and 4 samples were reclassified into the control group. To ensure that all samples were correctly predicted, we compared them to the true phenotype of each infant and confirmed the correct classification of all samples. After this process, 131 out of 140 infants were correctly classified, giving a TPR-ECUN and TNR-non-ECUN of 94.3% and 97.3%, respectively.

[0122] In the dataset, the latest control sample was obtained at DOL 447 while the last ECUN sample before the onset of pathology was collected at DOL 67. Therefore, to minimize microbial variations in the control samples beyond DOL 67, all control samples collected after this time point were excluded. The final dataset contained 1,172 control samples (160 infants) and 257 ECUN samples (42 infants), encompassing 1,394 input features (1,355 species and 39 clinical data: 10 gestational age groups, 18 weight groups, 7 DOLs, 2 birth modes, and 2 sex groups). The newly developed model demonstrated a high level of accuracy, reaching 92.3% and an AUROC of 0.960 ± 0.02. The model accurately predicted 227 control samples (102 infants) with a TNR-non-ECUN of 97%, while 7 samples (7 infants) were misclassified as ECUN. On the other hand, the model correctly identified 37 ECUN samples belonging to 19 infants, yielding a TPRECUN of 71.15%, but misclassified 15 samples (9 infants) as controls. Moreover, when serially reclassifying the samples, 3 infants (1 control and 2 ECUN) were accurately reclassified, leading to a TPR-ECUN and TNR-non-ECUN of 74% and 97.4%, respectively. Nevertheless, these results suggest that limiting the number of control samples had a negative impact on ECUN prediction, while using all samples (regardless of DOL) during training allowed the model to better differentiate control and ECUN subjects.The initial model (accuracy = 94.9%) was therefore retained.

[0123] The deep neural network is also interesting because it allows the identification of key species contributing to the model prediction. The top 20 features contributing to the model validation were summarized by the SHAP explainer and presented in [Fig.6]. The four most important contributors were Lactobacillus spp. and their high SHAP values ​​were associated with preterm control samples. Although the growth of Lactobacilli is generally promoted by breastfeeding, all preterm infants received probiotics that directly influence the abundance of Lactobacilli. In addition, RiboTaxa detected two unclassified bacteria, bacterium_129 and bacterium_ARb03, which contributed to the prediction of non-ECUN infants. Phylogenetic analysis revealed that bacterium_129 was potentially a new Lactobacillus species sharing 96.32% identity with L.casei, while bacterium_ARb03 shared 99.71% identity with Bacillus cereus strain ADY07. In contrast, the affiliated species unclassified Enterobacteriaceae, uncultured Syntrophomonas, Streptomyces vanillaeus, unclassified Enterobacteriaceae, and Enterococcus faecalis contributed the most to the ECUN classification. Interestingly, low-abundance species such as Ruminococcus sp., uncultured Staphylococcus, Streptococcus parasanguinis, and Proteus spp. were also observed to contribute to the ECUN classification, while unclassified Bifidobacterium were associated with non-ECUN samples.

[0124] The predictive performance of the model comes mainly from microbiome data. Excluding clinical and demographic characteristics (Table 1) during training resulted in performance values ​​(AUROC = 0.931, PR-AUC = 0.956) that were not significantly different from those of the model including metadata (p > 0.05 for the Mann-Whitney U test between ROC and precision-specificity curves with and without metadata).

[0125] To evaluate the model's performance on data outside of that used for model training, the inventors constructed their own cohort with a broader range of clinical and demographic data. 56 fecal samples from 9 preterm infants who developed ECUN and 10 preterm infants without ECUN included in the CORTECS cohort were analyzed. Cases and controls were matched based on sex, birth weight, mode of delivery, gestational age, and antibiotic administration. Detailed patient metadata is provided in the following Table 3.

[0126] [Tables3] Cohort Control ECUN Gender Male 58.4 64.3 Female (%) 41.4 35.7 Mode of birth Vaginal (vaginalX%) 20.7 21.4 Caesarean 79.3 78.6 Mean gestational age at birth [standard deviation] 30.4 [2.01] 29.7 [2.58] Mean number of days alive [standard deviation] 13 15 Mean birth weight (g) 1431 1372

[0127] Taxonomic profiling was performed using RiboTaxa and individual species-level profiles were retained. This resulted in the identification of 509 unique species in both groups with a reduced diversity of 349 species in ECUN samples compared to 420 species in control samples. RiboTaxa highlighted 260 common species between the two groups while 160 and 89 unique species were detected in the control and ECUN group, respectively. In this particular cohort, the gut microbiome composition of each preterm infant undergoes dynamic changes, regardless of their phenotypic classification. This can be attributed to clinical events that occur during intensive care unit treatment, as well as inter-individual variations that shape the microbiota.Species of the family Enterobacteriacea (unclassified Klebsiella, unclassified Escherichia-Shigella, and Enterobacter spp.) that are commonly associated in the pathogenesis of NEC were present in both groups and varied in relative abundance among infants.

[0128] Due to an insufficient number of samples (56), it was not possible to perform model training on this dataset. Therefore, the species relative abundance profiles and their associated metadata (sex, birth weight, gestational age, DOL and birth weight) were used to predict each sample's phenotype, using the neural network previously optimized deep neural network. Overall, the deep neural network yielded a TPR-ECUN of 82% (22 out of 27 ECUN samples) and a TNR-non-ECUN of 76% (22 out of 29 controls).

[0129] Among the features contributing to the different predictions, the prediction of control samples was mainly driven by the presence of a higher abundance of Lactobacillus spp. including L. rhamnosus, L. casei and Lactobacillus sp. In contrast, an ECUN prediction was made when the sample had a higher abundance of Enterococcus faecalis, Veillonella ratti, unclassified Klebsiella, Enterococcus durans, Enterobacter cancerogenus, Clostridium neonatale or C. per-fringens. Low abundant species such as uncultured Staphylococcus, Haemophilus parainfluenzae and Staphylococcus epidermidis contributed to the ECUN prediction in some samples, confirming that the phenotype prediction was influenced by the dominant and low abundant taxa, which showed a tendency to co-variate. This observation suggests the existence of a complex network of ecological interactions between abundant and rare taxa.Unlike microbial features explaining model learning, sample prediction was based on both microbial species and clinical data. Birth weight <800 g and gestational age <30 weeks were both factors often associated with NEC, while vaginal delivery and gestational age >31 weeks were characteristics of non-NEC infants.

[0130] After phenotype prediction, sample reclassification was performed based on at least three serial samples from the same infants. This led to the reclassification of 7 samples into 3 controls and 4 ECUN cases, among which 5 samples matched the true phenotype, resulting in improved TPR-ECUN (93%) and TNR-non-ECUN (86%). However, the two misclassified samples (one of each phenotype belonging to DI015control and HI125ECUN) implied that all remaining infant samples had also been misclassified by the model. The features explaining the prediction of DI015 control as ECUN were the high abundance of E. faecalis at 50% and 61% in samples collected on days 10 and 13 respectively. Conclusions

[0131] The results of this study are based on a systematic and comprehensive search with specific criteria of the gut microbiome based on shotgun metagenomics, merging two independent studies totaling 1,562 serially collected stool samples from 160 preterm controls (n = 1,305) and 48 ECUN infants (before ECUN onset, n = 257). All sequencing data were analyzed with the RiboTaxa pipeline (Chakoory et al., 2022), allowing the reconstruction of nearly complete 16S rDNA genes to provide a Accurate description of the gut microbiota down to the species level, including identification of dominant (>1%), subdominant (<1%), and rare (<0.1%) microorganisms. This approach led to a higher resolution of 1,282 unique bacterial species associated with preterm controls and ECUN infants.

[0132] The significant differences in α and [3] diversity further confirmed a distinct and dynamic microbial composition within and between the preterm and ECUN control groups. The gut microbiota of two infants may follow very different pathways as they progress from an initial state rich in aerobic microorganisms to a more diverse and anaerobic mature state. It is likely to be influenced by factors such as sex, gestational age, birth weight, day of life, mode of delivery, viral infections, dietary changes, and antibiotic administration. Many studies have examined neonatal stools to discern bacterial colonization patterns related to ECUN. These investigations have highlighted the correlations between microbiota characteristics and the occurrence of ECUN, however, without complete congruity in their results.

[0133] The results of the present study are consistent with previous studies, as the inventors also demonstrated the presence of unclassified Enterobacteriaceae, unclassified Escherichia Shigella, unclassified Enterobacter and unclassified Klebsiella at varying abundances in the control and ECUN samples. Although the abundance of unclassified Klebsiella and unclassified Escherichia Shigella did not differ significantly between the two groups, the average relative abundance of unclassified Enterobacter and unclassified Enterobacteriaceae was significantly higher in the ECUN samples than in the preterm controls.These highlighted microbial species, although providing valuable information on the identification of risk factors for ECUN, have not demonstrated sufficient consistency to establish definitive characteristics that could serve as a necessary or sufficient cause for the development of ECUN. Instead, ECUN is described as a microbial pattern of non-linear nature, characterized by complex changes in the composition and diversity of the microbiome, including altered interactions between dominant and rare microorganisms.

[0134] Olm et al. applied gradient-enhanced classification to distinguish ECUN infants from controls using taxonomic data and achieved 64% accuracy (Olm et al., 2019a).

[0135] In this study, the inventors developed and optimized a deep neural network to predict the risk of ECUN before the onset of the disease using gut microbiome profiles coupled with five metadata (sex, weight at birth, gestational age, DOL and mode of birth), potentially enabling early interventions to prevent the worst complications of the disease. The 1,355 species from the 1,562 samples were used as input features to obtain a high-performance model while retaining the inter-individual variability that is also responsible for the structured progression in the establishment and subsequent evolution of the neonatal gut microbiota.

[0136] To handle high-dimensional microbiome data with small sample sizes during model training, a deep neural network was compiled with a limited number of hidden layers, ranging from one to three, and implemented a neuron dropout regularization technique on the hidden layers. During training, some neurons are randomly assigned a value of zero, or “dropped out,” to prevent co-adaptations between neurons and thus reduce the risk of overfitting. In addition, a comprehensive validation framework was implemented that isolates the test set from hyperparameter optimization through a 10-fold cross-validation scheme to eliminate any potential bias or confounding that may arise during the model comparison process. The data also suffered from class imbalance (1305 non-ECUN and 257 ECUN samples).In the present study, to address the problem of unbalanced classes, the inventors used several metrics such as AUROC, accuracy, TPR-ECUN, and TNRnon-ECUN to estimate the model's performance to signal any bias when classifying one or the other phenotype. Thus, the deep neural network model according to the invention performed very well on the test set, with an AUROC score of 0.987 ± 0.01, an accuracy of 94.6%, a TPR-ECUN of 90.6%, and a TNRnon-ECUN of 95.8%.

[0137] Furthermore, this approach gave rise to a new method for reclassifying serial samples obtained from a single infant into the phenotypic group with the largest number of samples, thus enabling the determination of the infant's final phenotype. In doing so, the inventors increased the rate of correctly classified samples by 4% for ECUN cases and by 1% for unaffected controls, leading to 97% correct classifications (302 out of 313).

[0138] In the present study, the optimized deep neural network was not limited to the prediction of the host phenotype, but also stratified the samples through the characterization of phenotype-specific microbial signatures. The four most significant contributors were Lactobacillus spp. which were associated with the prediction of the control sample.

[0139] The present study also demonstrated the advantages of using complete rRNA genes to characterize the intestinal microbiome of premature infants with the identification of two novel species, namely bacterium_129 (a novel Lactobacillus species) that shared 96.62% identity with L. casei strain Dwan5. and unclassified Enterococcus (a new Enterococcus sp.) that shared 96.53% identity with E. faecalis DSM 713-1.

[0140] In contrast, unclassified Enterobacter, Enterobacteriaceae and Enterococcus faecalis, also identified in previous studies, contributed to the ECUN classification. The ecological role of minor components of the intestinal microbial population of preterm infants, such as uncultured Staphylococcus, Proteus spp. and unclassified Bifidobacterium, were also revealed by SHAP.

[0141] To evaluate the performance of the deep neural network on a completely new dataset from another NICU, the inventors built their own cohort of 56 serially collected stool samples from 19 preterm infants. All samples were characterized by 16S rRNA gene capture hybridization, a technique that allows for a fine-grained description of microbial communities (Gasc and Peyret, 2018). In this cohort, dynamic changes in the gut microbiome composition were observed for each preterm infant indifferent to the phenotypic class. This could be explained by inter-individual variations modulating the gut microbiota of newborns. Phenotypic prediction (based on the previously trained model) of these samples resulted in a TPR-ECUN and TNRnon-ECUN of 82% and 77%, respectively, which increased to 93% and 86% after reclassification of the serial samples.Based on these results, the inventors have achieved a high-performance model capable of effectively classifying samples from different areas and practices of the NICU despite the heterogeneity of the microbiome between cohorts.

[0142] In addition to the samples, a detailed collection of clinical data, including diet, antibiotics, the presence of other pathologies, Bell stages of ECUN or any inconsistent behavior of the child were also noted. The deep neural network successfully predicted the phenotypic samples related to the different ECUN stages (1a and 2a) in the French cohort, highlighting the exceptional ability of the deep neural network to distinguish complex microbial interactions between non-ECUN and ECUN samples.

[0143] Therefore, the method according to the invention demonstrates that phenotype prediction is determined by dominant and low-abundance taxa that tend to covary. This observation suggests the existence of a complex web of ecological relationships that explain how minor species contribute to widespread effects on the overall microbial community.

[0144] In summary, the present study challenged the concept that one or more pathogens cause a single disease. The use of near-full-length SSU rRNA-expressing genes using RiboTaxa to characterize stool samples from preterm infants confirmed a distinct and dynamic microbial composition within and between preterm control and ECUN groups.

[0145] The data demonstrate the model's ability to distinguish complex inter-individual microbial interactions between non-ECUN and ECUN samples, thus providing high-level information on changes in the overall structure of the newborn's intestinal community.

[0146] Among the features contributing to the different predictions, the prediction of ECUN samples was mainly driven by the presence of a higher abundance of Enterococcus faecalis, Veillonella ratti, unclassified Klebsiella, Enterococcus durons, Enterobacter cancerogenus, Clostridium neonatale or C. perfringens.

[0147] These results confirm that the risk of ECUN was not determined by a single species or group of species but by interindividual variability in intestinal microbial composition. Besides microbial species, clinical and demographic data influence sample prediction. Références bibliographiques

[0148] Chakoory O. et al. "RiboTaxa: combined approaches for rRNA genes taxonomie resolution down to the species levelfrom metagenomics data revealing novelties.” NAR Genomics Bioinformatics. 4, lqac070 (2022). - D’Angelo, et al. (2018) “Current status of laboratory and imaging diagnosis of néonatal necrotizing enterocolitis”, Italian Journal of Pédiatrie s. 44 (1). - Koenig et al., (2011). “Succession of microbial consortia in the developing infant gut microbiome”, Proceedings ofthe National Academy of Sciences ofthe United States of America, 108 Suppl 1 (Suppl 1), pp. 4578-4585. - Kopylova E. et al. "SortMeRNA: fast and accurate filtering of ribosomal RNAs in metatranscriptomic data." Bioinformatics 28.24 (2012): 3211-3217. - Masi et al. “Human milk oligosaccharide DSLNT and gut microbiome in preterm infants predicts necrotising enterocolitis", Gut, 70 (12), pp. 2273-2282 (2021) - Miller C.S. et al. "EMIRGE: reconstruction offull-length ribosomal genes from microbial community short read sequencing data." Genome biology 12.5 (2011): 1-14. - Olm et al. (2019a), “Necrotizing enterocolitis is preceded by increased gut bacterial réplication, Klebsiella, and fimbriae-encoding bacteria”, Science Advances, 5 (12), pp. eaax5727. - Pedregosa et al., (2011) Scikit-leam: Machine Learning in Python, The Journal of Machine Learning Research, 12 (null), pp. 2825-2830. - Quast C. et al., 2013). “The SILVA ribosomal RNA gene database project: improved data processing and web-based tools." Nucleic Acids Res. 41, D590-D596 (2013). - Rognes T. et al. "VSEARCH: a versatile open-source tool for metagenomics.” PeerJ 4 (2016): e2584. - Xue Y. et al. "Reconstructing ribosomal genesfrom large scale total RNA meta-transcriptomic data.” Bioinformatics 36.11 (2020): 3365-3371.

Claims

Claims

1. An in vitro method for predictive diagnosis of a condition or pathology, from a biological sample taken from the digestive system and / or from the excretions of a subject, said method comprising the following steps: a) providing the nucleic acid from a plurality of microorganisms present in said biological sample, b) determining by sequencing the nucleotide sequence of at least one nucleic acid chosen from: a gene fragment expressing 16S rRNA, a gene fragment expressing 18S rRNA, a 16S rRNA fragment and an 18S rRNA fragment, of a plurality of microorganisms, to generate a plurality of nucleotide sequences, c) organizing a plurality of nucleotide sequences determined during step b) to reconstruct the nucleotide sequence of at least one gene fragment expressing 16S rRNA, a gene fragment expressing 18S rRNA,a 16S rRNA fragment and an 18S rRNA fragment of a plurality of microorganisms, d) from the results of step c), determination of the identity, taxonomic classification, and relative abundance of a plurality of microorganisms present in said sample, and e) from the characteristics determined during step d), determination, by a previously trained classification model, of the predictive diagnosis of a condition or pathology.,

2. The method of claim 1, wherein said subject is a human newborn.

3. A method according to any preceding claim, wherein said pathology is an intestinal pathology.

4. A method according to any preceding claim, wherein said pathology is necrotizing ulcerative enterocolitis (NEC) in infants.

5. Method according to any one of the preceding claims, in which step e) is carried out from the characteristics determined during step d) and from at least one clinical data characteristic of the subject.

6. Method according to claim 5, in which said at least one clinical data is chosen from: age in number of days since birth, birth weight, gestational age, mode of birth, gender, ethnicity of the mother, diet of the mother, result of the dosage of blood components or markers, the administration of medical treatment, the presence of at least one other pathology or a combination of several of these characteristics.

7. Method according to any one of the preceding claims, characterized in that the classification model comprises: - a previously trained neural network, in particular during supervised learning, - a machine learning algorithm, and - a training data set.

8. Method according to any one of the preceding claims, characterized in that it further comprises a step of determining at least a first profile of a plurality of microorganisms, said first profile being characteristic of a predictive diagnosis of the presence of a condition or pathology of the digestive system or of a defined extra-digestive pathology.

9. Method according to the preceding claim, characterized in that it further comprises a step of determining at least a second profile of a plurality of microorganisms, said second profile being characteristic of a predictive diagnosis of the absence of a state or pathology of the digestive system or of a defined extra-digestive pathology.

10. A method according to any preceding claim, wherein determining the identity, taxonomic classification, and relative abundance of a plurality of microorganisms present in said sample is carried out using a computer tool for determining microbial community structures.

11. Use of a computer tool for determining microbial community structures for, from a nucleotide sequence of at least one gene fragment expressing 16S rRNA, a gene fragment expressing 18S rRNA, a 16S rRNA fragment and an 18S rRNA fragment of a plurality of microorganisms, reconstructed by organizing a plurality of nucleotide sequences, the plurality of nucleotide sequences being generated from the result of sequencing the nucleotide sequence of at least one nucleic acid chosen from: a gene fragment expressing 16S rRNA, a gene fragment expressing 18S rRNA, a 16S rRNA fragment and an 18S rRNA fragment, of a plurality of microorganisms: determine the identity of microorganisms in the microbial community; determine a relative abundance of microorganisms in the microbial community; and determine a predictive diagnosis of a condition or pathology based on the determined identity and relative abundance of microorganisms.

Citation Information

Patent Citations

  • Composition and methods for predicting necrotizing enterocolitis in preterm infants

    US20210254137A1