Method and system for probabilistic typing of microbial strains

EP4721064A1Pending Publication Date: 2026-04-08BIOMERIEUX SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Current genomic typing methods for microorganisms face challenges in achieving high precision while being cost-effective and time-efficient, particularly in distinguishing bacterial strains at a taxonomic level lower than the species level, due to the need for extensive marker sets and complex bioinformatics processing in whole-genome sequencing, or the limitations of using few markers in targeted typing methods.

Method used

A probabilistic typing method using a limited number of genetic markers, where a database of genetic profiles and a phylogenetic tree are used to determine the probability of microbial strain grouping, allowing for precise typing with fewer than 50 markers, and employing PCR amplification for genetic profiling without complete sequencing.

Benefits of technology

This approach enables high-precision microbial typing with reduced costs and time, effectively distinguishing microbial groups and serotypes, as demonstrated by high correct assignment rates and balanced precision in testing with Listeria and Salmonella strains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024064927_05122024_PF_FP_ABST
    Figure EP2024064927_05122024_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a microbiological typing method, which comprises providing a database of genetic profiles of microbial strains, a phylogenetic tree, assignments of the genetic profiles to positions in the phylogenetic tree, and frequencies of variation of the profiles in the tree. The typing of a microorganism comprises measuring the genetic profile of the microorganism, determining probabilities of variation of the genetic profiles belonging to the groups with respect to the measured genetic profile, the probabilities of variation being calculated as a function of the frequencies of variation at a predefined level of the phylogenetic tree and of the belonging of the microorganism to one of the microbial groups if an identification performance criterion is reached for the group, the criterion being reached if at least one probability of variation for the group is greater than a first predefined threshold.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]METHOD AND SYSTEM FOR PROBABILISTIC TYPING OF MICROBIAL STRAINS FIELD OF THE INVENTION The invention relates to the field of genomic typing of microorganisms, namely the characterization of bacterial strains at a taxonomic level below the species level based on their genomes. The invention finds particular application in the epidemiological monitoring of bacterial strains, the monitoring of contamination by bacterial strains and the search for the root causes of bacterial contamination and / or infection, whether their origin is industrial, environmental, veterinary, clinical or other. STATE OF THE ART The genomic typing of microorganisms, for example bacteria, brings together all the techniques for comparing the genomes of bacterial strains with the aim of characterizing the latter at levels below that of the species. Typing makes it possible in particular to determine the degree of similarity of two bacterial strains to each other or toclassify strains into species-level subdivisions, for example clonal groups or groups defined by phenotypic characteristics of the bacteria such as their serotypes or their susceptibility to antibiotics. Typing techniques can be classified into two categories, the first based on the analysis of a very restricted targeted portion of the bacterial genome, such as typing based on MLST markers (for "MultiLocus Sequence Typing") known as "16S rRNA", and the second based on the comparison of genomes obtained by complete sequencing (or WGS for "Whole Genome Sequencing"), for example typing of the cg-MLST type (for "core genome" MLST"). While the latter offer the highest level of precision, the use of complete sequencing has many disadvantages, including duration, cost, or the need to implement complex bioinformatics processing. Conversely, when the first techniques usea very limited number of markers, they allow the use of polymerase chain reaction (or "PCR") amplification platforms for their detection, which are much faster and less expensive. Whatever the technique used, state-of-the-art genomic typing is based on the search for and use of conserved portions of the genome to characterize microorganisms. For example, MLST typing is based on housekeeping genes that exhibit high stability over time, just as cg-MSLT typing directly targets the entire basic genome (or "core genome" in English, these three expressions being used interchangeably). It is observed in this type of approach that obtaining high precision requires a very large number of markers. For example, in order to detect clonal groups of the species P. aeruginosa with an accuracy greater than 90%, cg-MLST typing involves several thousandof markers, making the use of conventional PCR platforms de facto irrelevant, or even impossible. In summary, a user wishing to carry out molecular typing, for example for the purpose of monitoring microbial contamination, must choose between precision and speed. DISCLOSURE OF THE INVENTION The aim of the present invention is to propose molecular typing using a limited number of genetic markers and with high precision. To this end, the subject of the invention is an in vitro method for typing a microorganism included in a sample, comprising: a. the provision of computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the databaseat positions in the phylogenetic tree, of frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; − for at least part of the genetic profiles of the database, the belonging of said genetic profiles to microbial groups distinct from the microbial species, said distinct groups being absent from the phylogenetic tree; b. the measurement of the genetic profile of the microorganism; c. the determination, implemented by computer: − of probabilities of variation of the genetic profiles belonging to said groups compared to the measured genetic profile, said probabilities of variation being calculated as a function of the frequencies of variation at a predefined level of the phylogenetic tree; − of the belonging of the microorganism to one of the microbial groups if an identification performance criterion is reached for said group, said criterion being reached if at least oneprobability of variation for said group is greater than a first predefined threshold. According to one embodiment: − the phylogenetic tree is constructed from a core genome of the species of the microorganism; − the predefined set of genetic markers is chosen from the accessory genome of the species of the microorganism. According to one embodiment, the number of markers of the genetic profiles is less than 50, and preferably between 16 and 30. According to one embodiment: − the phylogenetic tree is constructed by applying clustering based on genetic distances from a core genome of the species of the microorganism; − the genetic distances decrease from the root to the leaves of the tree so as to define a level in the tree comprising between 150 and 350 differences in the core genome. In particular: − the groups of microorganisms correspond to microbial serotypes of the species of the microorganism; − the phylogenetic tree includes at least10 levels, preferably at least 15 levels; − the probabilities of variation of the genetic profiles belonging to said groups compared to the measured genetic profile are calculated as a function of the frequencies of variation of a level of the phylogenic tree comprised between 7 and 15, preferably between 10 and 15, more preferably between 11 and 13. According to one embodiment, the markers of the genetic profiles are selected so as to minimize an entropy or an impurity criterion at a level of the phylogenic tree comprising between 150 and 350 differences in the core genome. According to one embodiment, the identification performance criterion further comprises the difference between the probability of variation for said group and the probabilities of variation for the other groups, the identification performance criterion being achieved if said difference is greater than a second predefined threshold. In particular, the identification performance indicatorcomprises a correct assignment rate and / or a raw precision and / or a balanced precision. According to one embodiment, comprising the search for racial causes during contamination of a food product. According to one embodiment, the measurement of the genetic profile of the microorganism is carried out by means of an amplification of target genetic sequences without implementing complete sequencing, in particular a polymerase chain reaction. The invention also relates to an in vitro typing system for a microorganism included in a sample, comprising: a. computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the database topositions in the phylogenetic tree, of frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; b. means for collecting and storing a measurement of the genetic profile of the microorganism; c. computer-implemented means for determining: − probabilities of variation of the genetic profiles belonging to said groups compared to the measured genetic profile, said probabilities of variation being calculated as a function of the frequencies of variation at a predefined level of the phylogenetic tree; − the belonging of the microorganism to one of the microbial groups if an identification performance criterion is reached for said group, said criterion being reached if at least one probability of variation for said group is greater than a first predefined threshold. According to one embodiment, the system is configured to implement a method of the aforementioned type.The invention also relates to a computer program product recorded on a computer-readable computer medium comprising instructions for carrying out step c. of a method of the aforementioned type. The invention also relates to a composition for the in vitro typing of microorganisms included in a sample by amplification of target nucleotide sequences, in particular by a polymerase chain reaction, said composition comprising a set of oligonucleotides for the amplification of at least the set of genetic markers obtained according to a method of the aforementioned type. The invention also relates to a method for preparing a biological sample comprising, or likely to comprise, a bacterial strain, with a view to typing said strain by a polymerase chain reaction, comprising bringing said sample into contact with a composition of the aforementioned type. According to one embodiment, the method comprisesa lysis of the sample resulting from said contact so as to release genetic material from the bacteria of the bacterial strain. The invention also relates to a method for preparing a sample likely to comprise a microorganism, said method comprising − the extraction of the genetic material included in the sample; − the mixing of said extracted genetic material with a composition comprising a set of primers for the amplification, in particular by polymerase chain reaction, of a set of genetic markers obtained according to a method of the aforementioned type. BRIEF DESCRIPTION OF THE FIGURES The invention will be better understood on reading the following description, given solely by way of example, and carried out in relation to the appended drawings, in which: − Figure 1 is an overall view of a method according to the invention; − Figure 2 is a flowchart detailing a learning phase of a phylogenetic typing according to the invention; − theFigure 3 is a plot illustrating a decreasing genomic distance for the construction of a phylogenetic tree; − Figure 4 is a schematic view of a binary phylogenetic tree; − Figure 5 is a diagram illustrating the calculation of the variation frequencies of a genetic profile according to the invention for a taxon of the phylogenetic tree; − Figure 6 is a flowchart illustrating the production of a PCR kit used for typing according to the invention; − Figure 7 is a flowchart detailing a prediction phase of a phylogenetic type according to the invention; − Figure 8 is a flowchart illustrating the calculation of thresholds used during the prediction phase; − Figure 9 is a flowchart illustrating a learning phase for serotyping according to the invention; − Figure 10 is a flowchart illustrating a serotype prediction phase according to the invention; − Figure 11 is a flowchart illustrating the calculation of thresholds used during the prediction phaseof serotype; − Figure 12 is a plot of the decay of Shannon entropy when selecting genetic markers constituting a genetic profile for typing Listeria; − Figure 13 is a set of plots of frequencies of variation of genetic profiles in the phylogenetic tree of Listeria; − Figure 14 is a plot of the correct serotype assignment rate of Listeria as a function of balanced accuracy; − Figure 15 is a set of plots of the correct phylogenetic assignment rate of Listeria as a function of average accuracy; − Figure 16 is a plot of the decay of Shannon entropy when selecting genetic markers constituting a genetic profile for typing Salmonella; − Figure 17 is a set of plots of frequencies of variation of genetic profiles in the phylogenetic tree of Salmonella; − Figure 18 is a plot of the rate of correct serotype assignment of Salmonella as a function of thebalanced accuracy; − Figure 19 is a set of plots of correct phylogenetic assignment rates of Salmonella as a function of average accuracy; and − Figures 20 to 22 illustrate computer architectures for implementing the typing according to the invention. DETAILED DESCRIPTION OF THE INVENTION A. Definitions In this document, the following definitions apply: − Biological sample: means any sample containing or suspected of containing one or more microorganisms; − Basic genome (or "core genome" or "heart genome", noted "Gb") of a microbial species: means parts of the genome (genes or groups of genes) shared by all the microbial strains belonging to the microbial species(s) considered. For the purposes of the invention, the basic genome can evolve if the set of strains from which it is produced evolves. Preferably, the core genome is constructed from at least 100 different microbial strains, andpreferably from at least 1000 strains. It will be noted that the genomic portions of microbial strains belonging to the basic genome may differ from one another. For example, a particular gene may belong to the basic genome and have several different alleles. In the following, the expression "basic genome" when applied to a microbial strain corresponds to the portion of the genome of the strain belonging to the basic genome of the species. − Accessory genome of a microbial species (noted "Ga"): designates parts of the genome (genes or groups of genes) of a set of bacterial strains which do not belong to the basic genome of said set. These parts of the genome are therefore absent from the genome of at least one of the bacterial strains of the set of bacterial strains considered. For the purposes of the invention, the accessory genome can evolve if the set of strains from which it is produced evolves. Preferably, the accessory genome is constructedfrom at least 100 different microbial strains, and preferably from at least 1000 strains. − Genetic marker: designates a gene or a nucleotide sequence at a specific position in the genome (also called a "locus"), a nucleotide sequence that can be reduced to a single nucleotide, such as two alleles differing by a single nucleotide, or to a polymorphism of a single nucleotide in the context of a mutation (or SNP for "single nucleotide polymorphism"); − Microorganism: designates all microscopic organisms, including bacteria, yeasts and fungi; − Phylogenetic tree of a microbial species: designates a tree representing evolutionary relationships among a set of microbial strains, relationships determined according to their similarities and differences in their core genomes. For the purposes of the invention, a phylogenetic tree can be enriched if the set of strains from which it is produced evolves.preferably, a phylogenetic tree is constructed from at least 100 different microbial strains, and preferably from at least 1000 strains. By convention, the root of a phylogenetic tree is designated as the first level of the tree and the leaves of the tree as the last levels. A level "higher" than another thus corresponds to a level closer to the root. B. Embodiment of the invention A probabilistic typing according to the invention is now described, applied to the monitoring of contamination by bacterial strains of a species within an organization, in particular in premises for the production of food intended for human consumption, as well as the detection of their serotypes. Referring to Figure 1, the typing according to the invention comprises the following main phases: - a phase 10 of constituting a database of complete genomes of strains belonging to the bacterial species; - a phase 20 called"learning" or "training" phase, aimed at generating a phylogenetic tree based on the genomes in the database, choosing an appropriate genetic profile and frequencies of variation of said genetic profile along the tree; - a phase 30 called "prediction" carrying out the typing of an unknown bacterial strain based on the tree and the frequencies produced during phase 20; and - phases possibly completed by a phase 40 called "epidemiological" during which corrective and / or prophylactic measures are implemented within the organization based on the result of the typing of the unknown strain. Preferably, the database of complete genomes is as representative as possible of the genetic diversity of the species. It is advantageously completed over time by adding complete genomes so as to target this diversity if its initial version is incomplete. In this case, the typing according to the invention is implementedwhen said database comprises at least 100 genomes of the species, preferably at least 1000 genomes and even more preferably several thousand genomes. Referring to Figures 2, 3 and 4, the training phase 20 begins with the production of a basic genome in 202 and an accessory genome in 204 from the genomes of the database 200. In a first variant, the genomes comprise annotations categorizing the genes as belonging to one or the other of these genomes. In a second variant, the latter are produced ad hoc from the database, the genes having an allele frequency greater than a threshold, for example 95%, belonging to the basic genome, otherwise to the accessory genome. In a third variant, the basic and accessory genomes are produced from the database by comparison with a reference genome of the species, stored in the database, the genes having an allele frequency compared to thereference genome above a threshold, for example 95%, belonging to the core genome, otherwise to the accessory genome. In a fourth variant, the database is already established for the species as well as a version of its core and accessory genome. The typing according to the invention continues, in 206, by the generation of a phylogenetic tree based on the core genome. Several techniques for constructing such a tree are possible, as described for example in the documents "Phylogenetics Algorithms and Applications" (Hu et al., Ambient Communications and Computer Systems, 2019), "A Review: Phylogeny Construction Methods" (Shaktawat et al., International Journal of Emerging Science and Engineering, 2019) or "Essential Bioinformatics" (chapter 11 by Jin Xiong, Cambridge University, 2011). In a variant of the invention, the tree is constructed from a cg-MLST (for "core-genome multi locus sequencing typing") coding of the basic genomes. These aretranslated into cg-MLST profiles using, for example, BioNumerics software (bioMérieux, Marcy l'Etoile, France), SeqSphere+ (Ridom, Münster, Germany), BIGSdb (available at https: / / github.com / kjolley / BIGSdb and described in Schürch, A. et al. "Whole genome sequencing options for bacterial strain typing and epidemiologic analysis based on single nucleotide polymorphism versus gene-by-gene–based approaches." Clin. Microbiol. Infect., 2018), or chewBBACA (available at https: / / github.com / B-UMMI / chewBBACA and described in Silva M et al., "chewBBACA: a complete suite for gene-by-gene schema creation and strain identification." Microb. Genom., 2018). The advantage of cg-MLST profiles is that they are fixed-length sequences of a few hundred components, aligned with each other, allowing the direct use of tree construction methods. This tree is advantageously generated by applying a succession of groupings by methodsingle-linkage clustering), and therefore based on genetic distances between core genomes, for example calculated using the normalized Hamming distance. Subdividing a taxon in the tree then consists of grouping into the same subgroup core genomes differing from each other by a distance less than a predefined threshold. For example, software such as HierCC (available at https: / / github.com / zheminzhou / pHierCC and described in the document "HierCC: a multi-level clustering scheme for population assignments based on core genome MLST" by Zhou et al., Bioinformatics, 2021) and Pathogenwatch from the Center for Genomic Pathogen Surveillance (available at https: / / pathogen.watch / ) can be used. Once the tree has been constructed, the latter is stored in a database 208. In a particularly preferred variant of the invention, the number of levels of the tree and the clustering distances are chosen so as toto produce at least one level in the tree such that the core genome diversity at this level matches or is close to the clonal diversity observed in the species, so that the taxa at this level make it possible to distinguish different clonal complexes among the bacterial strains corresponding to the complete genomes in the database. For example, for Salmonella and Listeria, the level has between 150 and 350 differences in the core genome of the strains, preferably between 200 and 300 differences. In particular, the clustering distance decreases from the root to the leaves, for example linearly. The learning phase 20 continues with the choice, at 210, of a type of genetic profile in the accessory genome, namely a set of genetic markers of the accessory genome of the species whose absence and presence are measured in a strain to be tested during the prediction phase 30. The choice of the profile consists of choosing a level inthe tree, and preferably one of the levels representing clonal diversity, then to choose a set of markers having a minimum entropy or degree of impurity at this level. The total number N M of markers selected to construct the genetic profile is advantageously selected according to the particular PCR platform which is used to carry out the typing. For example, for the applicant's GENE-UP platform, this number N Mis equal to 23 markers for Salmonella and 16 markers for Listeria monocytogenes. In practice, the set of markers is chosen so as to minimize an entropy criterion (e.g., Rényi, Hartley, Shannon, collision, or min entropy) or a Gini impurity criterion, by implementing a greedy algorithm. As an illustration, and referring to the phylogenetic tree in Figure 4, the selection of a set of markers minimizing the Shannon entropy criterion of level N consisting of Q N = 20 taxa T1, T2,…, T 20 consists of choosing a first marker M1 of the accessory genome minimizing the entropy criterion: où ^^^^ ( ^^^^ ) = ^^^^ ^^^^ with NG the total number of core genomes used to construct the phylogenetic tree and ^^^^ ^^^^the number of times the marker is present in accessory genomes in the i ème taxon T i . Once the M1 marker is selected, this procedure is relaunched to select the second M2 marker in the accessory genome removed from the M1 marker. Thus the j ième marker M j is the marker such that ^^^^ ( ^^^^ ) = ^^^^ ^^^^ with NG the total number of core genomes used to construct the phylogenetic tree and ^^^^ ^^^^ is the number of times the marker ^^^^ ^^^^ is present in accessory genomes in the i ème taxon T i . Alternatively, the stepwise selection of markers is performed by minimizing the conditional entropy of level N according to the formula: Or: − ^^^^ ( ^^^^ ) =^^^^ ^^^^ with NG the total number of core genomes used to construct the phylogenetic tree and ^^^^ ^^^^ the number of times the marker is present in accessory genomes in ème − and ^^^^ ( ^^^^ ^^^^ | ^^^^ ) = ^^ s the i taxon T i ^ ^ ; ^ ^^^ ^ ^^^ ^^^^ with ^^^^ ^^^^ the number of times that the l ième marker set value … , ^^^^ ^^^^−1, ^^^^ ^^^^ �, among the 2 ^^^^ possible values ​​of said set, is present in the ith taxon T i and ^^^^ ^^^^ the number of core genomes included in taxon T i. The inventors have observed that the typing according to the invention requires at most 50 markers, preferably less than 30 markers, more particularly a number between 16 and 30 markers, to obtain high typing performance. The learning phase 20 continues, in 212, by the extraction of the accessory genomes from the genetic profiles consisting of the selected markers, profiles stored in a database 214. The genetic profile of an accessory genome is therefore a vector P of {0,1} ^^^^ ^^^^ , the j ième component P(j) of the vector P corresponding to the value of the marker ^^^^ ^^^^ , this value being equal to 0 when the marker ^^^^ ^^^^is absent and to 1 when it is present. Once the profile base 214 is constituted, frequencies of variation of the profiles along the tree are calculated at 216. Referring to Figure 5, for each taxon T of the tree which is partitioned into subdivisions T1, T2, …, TQ, there is a set of pro associated with the taxon T partitioned into subsets of profiles. ^^^^ ^^^^ fils , ^ genetics{^^^^ , ^^^^ , … , ^^^^} 1 ^^^1 ^^^^11 2 ^^^^ 1 ^^^2^ , … , ^^^ ^^^^^1�,respectively associated with the subtaxa T1, T , …, T Q . For each marker M j of the genetic profile, a frequency of variation ^^^^( ^^^^ ^^^^| ^^^^) of this marker is calculated for the taxon T as being equal to the number, divided by Q, of sub-taxa T1, T2, …, T Q for which the marker M j takes both the value 0 and the value 1, namely a frequency according to the relation: where ^^^^ ^^^^ ^^^^ ^^^^ ( ^^^^ )denotes the cardinality function of a set E. The calculation of frequencies thus produces for each subdivided taxon T, a vector of frequencies ^^^^ ( . | ^^^^ ) de [ 0,1 ] ^^^^ ^^^^ tel que ^^^^(. | ^^^^)( ^^^^) = Once calculated, the frequency vectors are stored in a database 218. The learning phase 20 ends with the calculation, at 220, of a grid of hyperparameters used during the prediction phase 30, a calculation which will be detailed below once the prediction 30 is described, parameters being stored in a database 222. In summary, the learning phase 20 provides a phylogenetic tree in which each complete genome and its genetic profile extracted from its accessory genome is positioned in the tree and each taxon of the tree, excluding the leaves, is associated with a frequency of variation of each genetic marker of the profile.Referring to Figure 6, the invention comprises the design and production 50 of a composition, or "kit", for carrying out an amplification aimed at detecting the selected markers, comprising: - the selection, at 500, of the amplification primers for the amplification of the selected genetic markers: - the production of said primers at 502 and if necessary the production of detection probes so as to detect the presence of a marker in at least one of its variants, and preferentially in the maximum number of its variants; - the production, at 504, of a buffer solution comprising the selected primers, the detection probes if necessary, dNTPs (deoxyribonucleotide triphosphates) providing the energy and the nucleotide bases necessary for the production of amplicons, an amplification enzyme and salts such as Mg ions. 2+or NaCl allowing the enzyme to function properly, at appropriate concentrations. Dyes can be added to this solution. − the production of a lysis buffer comprising beads, for example magnetic or ceramic beads, possibly a dye. The kits produced are then used during the prediction phase 30. Referring to Figure 7, this phase begins with the collection 300 of a biological sample in the organization, for example a food product production plant. During a following step 302, the sample is prepared so as to detect a given pathogen by methods known in the state of the art such as a culture on a dish, an immunological detection of the VIDAS type ® , or even by molecular biology for example with the GENE-UP platform ®. Following the detection of said pathogen, a confirmation step can take place. At the same time, if the colony is identified as belonging to the species, the prediction phase 30 continues at 302 using isolated colonies or an enrichment broth for its characterization by PCR amplification. In particular, the sample undergoes a set of preliminary steps, such as lysis to release the DNA of the bacteria constituting it, optionally extraction of the DNA resulting from the lysis using beads, for example magnetic or ceramic beads. Once the DNA of the bacterial strain has been released, it is added, at 304, to the PCR kit designed to target at least the genetic markers selected during the learning step 20. Finally, the PCR is implemented at 306 in order to measure the genetic profile of the strain, for example using the GENE-UP platform. ® of the Applicant. The measured profile, noted P mes, is then analyzed, at 308, in order to position the bacterial strain in the phylogenetic tree. Starting from a level N of the tree, for example the level directly above the leaf level, the analysis 308 includes the following steps: a. for each taxon of level N, noted ^^^^^^^^, ^^^^(i ème N taxon ième level): a1. for each stored profile, noted ^^^^^^^^ ^^^^ ^^^^, from the database 214 which is assigned to the taxon ^^^^^^^^, ^^^^, the calculation of a probability ^^^^^^^^ ^^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^^^^^ ^^^^ ^^^^) than the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^either a variation of the memorized profile ^^^^^^^^ ^^^^ ^^^^according to the relationship: We will note that if the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^is identical to the profile ^^^^^^^^ ^^^^^ then this probability is equal to 1 (no variation) and decreases rapidly if the profile ^^^ ^^^^^ ^^^^ ^^^^deviates from the profile ^^^^^^^^ ^^^^ ^^^^, and this is all the more so as the frequencies ^^^^^^^^ ^^^^^ ^^^^� ^^^^ ^^^^ � ^^^^^^^^, ^^^^� are weak. a2. selection of the maximum ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^( ^^^^, ^^^^ ) probabilities ^^^^^^^^ ^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^^^^^ ^^^^ ^^^^) calculated for the taxon ^^^^^^^^, ^^^^. b. Assignment, in 310, of the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^to one of the taxa ^^^ ^ ^ ^^^, ^^^^ of level N such that: where ^^^^1 and ^^^^2 are predetermined non-zero positive thresholds. c. If no taxon of level N meets the criterion of step b, move to the next level N-1 of the tree and implement steps a and b for this new level. Once the profile ^^^ ^^^^^ ^^^^ ^^^^assigned to a level of the phylogenetic tree and a taxon of this level, the prediction phase 30 delivers a report to the user, for example in the form of a display on a screen and / or files stored on a computer and / or a message (e.g. email) for the user. If two taxa have identical results at the end of step b, these two taxa are included or an inconclusive typing is noted in the report. The position of the new bacterial strain tested in the phylogenetic tree makes it possible to type the latter in relation to the bacterial strains corresponding to the complete genomes of the database 200. In addition to this type in reference to strains previously observed and capitalized in the different databases, the typing according to the invention allows the relative typing of two bacterial strains if these two strains are assigned to the same taxon.In addition, the invention makes it possible to define the lineage of a strain in relation to those in the databases or in relation to another strain tested. Indeed, the relative position of two strains in the tree makes it possible to characterize their evolution in relation to each other. A path through the phylogenetic tree has been described from the leaves to the root: as long as the assignment relationships are not verified, the tree is ascended. Alternatively, the phylogenetic tree is traversed from the root to the leaves: as long as the assignment relationship is verified, the tree is descended. The values ​​of the thresholds. and ^^^^2 are hyperparameters determined in step 220 of the training phase 30, preferably by means of cross-validation. Referring to Figure 8, a grid of values ​​for the thresholds and ^^^^2 is first determined, and for each pair of thresholds of the grid, a determination of the correct assignment rate and the average precision is implemented by: − the partition of the complete genome base 200 into a training base 800 and a test base 802 according to a cross-validation strategy 804, for example a so-called 10-fold strategy ("ten-folds cross validation" in English); − the implementation of the learning phase 20 on the training base 802; − the implementation of the prediction phase 30 for each genetic profile from the test base 802 and the calculation in 806 of the average precision ACC and the correct assignment rate TA of the genetic profiles of the base 802 in the phylogenetic tree; − the selection, in 808, of the maximum value of the average precisions and the correct assignment rates among the folds of the cross-validation strategy.These two values ​​therefore correspond respectively to the average precision and the correct assignment rate of the phylogenetic prediction phase 30 for the thresholds. and ^^^^2. The pairs of thresholds having the best precision and rates are retained for the phylogenetic typing according to the invention. Alternatively, the user can choose the precision and / or the assignment rate that he desires and the phylogenetic typing selects the pair of thresholds whose precision and rate is closest to those chosen by the user. Preferably, the number of levels in the phylogenetic tree and the law of decrease of the clustering distances, for example affine, are also hyperparameters, the method described in relation to figure 8 being therefore completed by the selection of a grid for these two parameters, and the traversal of said grid jointly with the traversal of the threshold grid and ^^^^2. A preferred variant of the invention is now described, consisting of predicting the membership of a bacterial strain to a particular group of the bacterial species based on its measured genetic profile. In the illustrated example, these groups are serotypes, but this variant applies to any explicit subdivision at a level of the bacterial species. It will be noted that this subdivision of the species does not correspond to the first level of the phylogenetic tree and is defined independently of the latter. Referring to Figure 9, the learning phase 20 described above in relation to Figure 2 is completed by a serotype learning 20A comprising: a. the collection of the serotypes of all or part of the bacterial strains corresponding to the complete genomes, serotypes stored in a database 224; b.the distribution, in 226, of the observed genetic profiles of the database 214 of the serotyped strains among the NS different serotypes ST1, ST2,…ST. NS . The genetic profiles associated with a serotype are hereinafter referred to as "serotyped"; c. the calculation, 228, of the frequencies of variation of the serotyped genetic profiles. In particular, for each marker M j , its variation frequency in an STi serotype is equal to the number of times this marker takes the value 1 out of the total number of genetic profiles assigned to the STi serotype; d. the storage in a database 230 of the variation frequencies serotyped genetic profiles in the different serotypes. If the elements of the learning phase 20 described in relation to Figure 2 remain unchanged, meaning in particular that the type of genetic profiles used is determined according to the phylogenetic tree and not the subdivision into serotypes, the prediction phase 30 can take two forms. In the first, the phase 30 already described is completed by a prediction of the serotype of a new strain to be tested, in parallel with steps 308-312 of Figure 2. In the second form, only said prediction is solely implemented and replaces steps 308-312. Referring to Figure 10, the prediction 30A of the serotype of a new bacterial strain follows steps 300-306 measuring the genetic profile of a bacterial strain of the species contained in a biological sample, and comprises the following steps: a. for each serotype, noted ST i: a1. for each serotyped genetic profile, noted ^^^^ ^^^^ ^^^^ , from the ST group i from database 228, the calculation of a probability ^^^^^^^^ ^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^ ^^^^ ^^^^ ) than the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^either a variation of the memorized profile ^^^^ ^^^^ ^^^^ according to the relationship: a2. selection of the maximum ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^( ^^^^ ^^^^ ^^^^ ) probabilities ^^^^^^^^ ^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^^^^^ ^^^^ ^^^^) calculated for the ST serotype i ; b. Assignment of the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^to one of the ST serotypes j such as : where ^^^^3 and ^^^^4 are predetermined non-zero positive thresholds; c. Production and delivery, at 318, of a report on the serotype detected for the bacterial strain subject to serotype typing. In a variant of serotyping 30A, if the second condition ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ � ^^^^ ^^^ ^ ^ ^^^ � − ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^^ ( ^^^^ ^^^^ ^^^^ ) ≥ ^^^^4 is not satisfied, serotyping returns the first two or three serotypes associated with the probabilities ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^( ^^^^ ^^^^^^^^ )the highest, as well as said probabilities. The values ​​of the thresholds ^^^^3 and ^^^^4 are hyperparameters determined during step 220 of the training phase 30, preferably by means of cross-validation. Referring to Figure 11, a grid of values ​​for the thresholds ^^^^3 and ^^^^4 is first determined, and for each pair of thresholds in the grid, a determination of the correct assignment rate and the average accuracy for serotype typing is implemented by: − partitioning the complete genome databases 200 and serotypes 224 into a training database 1100 and a test database 1102 according to a cross-validation strategy 1104,for example a so-called 10-fold strategy ("ten-folds cross validation" in English); − the implementation of the learning phase 20 - 201 on the training base 1102; − the implementation of the prediction phase 30A for each genetic profile from the test base 1102 and the calculation in 1106 of the average precision ACC and the correct assignment rate TA of the genetic profiles of the base 802 in the serotypes; − the selection, in 1108, of the maximum value of the average precisions and the correct assignment rates among the folds of the cross-validation strategy. These two values ​​therefore correspond respectively to the average precision and the correct assignment rate of the serotype prediction phase 30A for the thresholds ^^^^3 and ^^^^4. The pairs of thresholds presenting the best precision and rates are retained for serotype typing according to the invention. Alternatively,the user can choose the precision and / or the assignment rate that he desires and the serotyping according to the invention selects the pair of thresholds whose precision and rate is closest to those chosen by the user. In the variant just described, the determination of the serotype of a strain to be tested comprises the calculations as described in relation to figures 9 and 10. In another variant, for each taxon of the phylogenetic tree, the probability of each serotype in this taxon is calculated according to the serotype database 224. For example, the probability of a given serotype ^^^^ ^^^, ^ ^ ^^^ in the taxon T iis equal to the number of genomes grouped in this taxon that exhibit this serotype divided by the total number of genomes in the taxon. These probabilities are stored in a reference table or dictionary, and predicting the serotype of a strain to be tested involves first determining the taxon in the phylogenetic tree to which it belongs in the manner described in steps 308-312. Once the taxon has been identified, serotyping the strain involves querying the reference table probabilities associated with said taxon and delivering at least the serotype with the highest probability to the user. In the typing and serotyping variants just described, a calculation of the values ​​^^^^^^^^ ^^^^ ^^^^, ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^^ , and assignment of the measured profile ^^^ ^^^^^ ^^^^ ^^^^is implemented once this profile is measured. Alternatively, the already known profiles are associated directly with the taxon or taxa of the binary phylogenetic tree, for example by means of a reference table. Indeed, having previously carried out the assignment calculation, the taxa are already known. In particular, the profiles ^^^^^^^^ ^^^^ ^^^^ are assigned to their associated taxa by means of the calculations previously described. Thus, if the profile of a new strain is identical to a profile ^^^^^^^^ ^^^^ ^^^^, the reference table is called for this profile and the associated taxa are directly output. Similarly, when a new profile ^^^ ^^^^^ ^^^^ ^^^^is measured and different from all the profiles in the reference table, its assignment is calculated as described above, then this new profile is stored with the taxon(s) identified in the reference table. In this way, we avoid implementing assignment calculations that have already been performed by a simple call to a reference table. Alternatively, rather than a reference table, each already known profile is stored in a memory space dedicated to the taxon(s) of the phylogenetic tree. In other words, the taxa in the tree are assigned "addresses" equal to the known and assigned profiles. D. Examples of implementation D.1. Listeria monocytogenes 37,000 strains of Listeria were collected, completely sequenced and serotyped. The number of levels in the tree is equal to 23, i.e. the total number of targets of the GENE-UP PCR platform ®of the Applicant, and the distance decay law is selected for clustering so as to obtain a diversity similar to the clonal diversity at a level between 10 and 15 in a phylogenetic tree with 24 levels. Figure 12 illustrates the decay of the Shannon entropy of the genetic profile as markers are selected in the accessory genome at level 13 of the phylogenetic tree, the entropy of the final profile being only a few percent of the free entropy at this level. Figure 13 illustrates the frequencies of variation of the profile at the different levels of the tree, with a curve in a graph corresponding to a taxon of the level and the bold curve corresponding to the frequency of the labeled marker.As can be seen in Figure 14, which illustrates the serotype typing among 12 known Listeria serotypes, for a correct serotype assignment rate of 100%, the average balanced accuracy is equal to 90%, demonstrating the effectiveness of this typing according to the invention. It should be noted that it is possible to choose a higher accuracy for a satisfactory lower assignment rate. In particular, for a rate of 80%, the accuracy is close to 95% with the 95% confidence interval contained entirely in the last decile. Figure 15 illustrates the average accuracy of a level as a function of the correct phylogenetic assignment rate. Here again, we note the excellent level of performance of the invention, these two values ​​being respectively greater than 95% and 70% for any level above level 15, and each close to 100% for levels above level 9. D.2. Salmonella 39,000 Salmonella strains were collected, completely sequenced and serotyped.The number of levels in the tree is equal to 23, which is the total number of targets in the GENE-UP PCR platform. ®of the Applicant, and the law of decreasing distances is selected for clustering so as to obtain a diversity similar to the clonal diversity at a level between 10 and 15 in a phylogenetic tree comprising 24 levels. Figure 16 illustrates the decrease in the Shannon entropy of the genetic profile as markers are selected in the accessory genome at level 13, the entropy of the final profile being substantially zero. Figure 17 illustrates the frequencies of variation of the profile at the different levels of the tree, a curve in a graph corresponding to a taxon of the level and the curve in bold corresponding to the average frequency of the level. As can be seen in Figure 18, which illustrates the serotype typing among the more than 2,600 known Salmonella serotypes, for a correct serotype assignment rate of 100%, the average balanced accuracy is greater than 90%, demonstrating the effectiveness of this typing according to the invention.It should be noted that it is possible to choose a higher precision for a satisfactory lower assignment rate. In particular, for a rate of 80%, the precision is close to 100% with the 95% confidence interval contained entirely in the last decile. Figure 19 illustrates the average precision of a level as a function of the correct phylogenetic assignment rate. Here again, we note the excellent level of performance of the invention, these two values ​​being respectively greater than 95% and 80% for any level higher than level 13. E. Preventive and curative measures The solution's ability to easily and quickly identify strains makes it possible to consider its use for the search for root causes during contamination of a finished food product identified during routine controls.This includes, in particular, the search for the strain in the environment of the factory concerned, using surface samples using, for example, swabs, sponges or cloths. The search for the strain can also be carried out in the raw materials involved in the design of the finished product concerned, in order to link the contamination to one or more suppliers and / or carrier(s) and / or storage location(s). It can also make it possible to identify contamination during the laboratory detection phase, either by a laboratory strain or by their DNA, which would have biased the release test that identified the presence of the species concerned. In this case, it would make it possible to establish a plan for retesting the finished product and avoid its destruction.Finally, the solution can be used during a health crisis to quickly characterize the responsible strain(s) and link them to strains identified in incriminated food, cosmetic, veterinary, or other products. Its usefulness therefore covers internal needs of a production plant, for example of food products, investigation needs during health crises or during on-site regulatory audits. The use of this solution thus makes it possible to rationalize and implement appropriate preventive and / or corrective measures to reduce or even eliminate the risks of recurrence of contaminations, and consequently of infections. F.Hardware computer implementation of the invention The learning phase 20, 20A and the prediction phase 30-30A, apart from the steps of sample preparation and measurement of the genetic profile of the strains to be tested, are implemented by computer, namely by means of hardware circuits comprising computer memories (cache, RAM, ROM, etc.) and one or more microprocessors or processors, organized or not in the form of calculation nodes, necessary for the execution of computer instructions stored in the memories for the implementation of said phases. Several architectures are possible, as for example illustrated in figures 20-22. In a first architectural variant (fig.20), a first organization 2000, for example the Applicant, hosts or controls one or more calculation servers 2002 associated with one or more databases 2002 for storing genomes, the phylogenetic tree, phylogenetic frequencies, observed genetic profiles, serotypes and serotype frequencies and the learning phase 20, 20A is implemented by the organization 2000 on its server(s) 2002. A second organization 2006, for example a microbiological laboratory, hosts a PCR platform 2008 connected to, or incorporating, a computer 2010 and implements part of the prediction phase 30, namely all of the steps up to the measurement of the genetic profile of a bacterial strain to be tested. The measured profile is then pushed, through a remote connection network, to the first organization 2000 for implementation on the server 2002 of the remainder of the prediction phase 30, 30A.The report produced is then pushed, through the network 2012, to the second organization 2006 which takes or does not take epidemiological measures depending on the report. A second architectural variant (fig. 21) differs from the first, in that a copy of the database 2004 and the software implementing the prediction phase 30, 30A are downloaded into the second organization which implements, using the computer 2010 or a computing server (not shown), the entire prediction phase 30, 30A. In a third variant (fig. 22), all of the learning phases 20, 20A and prediction 30, 30A are implemented by a single organization 2006. G. Extension of the teaching of the embodiment An embodiment of the invention has been described using PCR to measure the genetic profile of a bacterial strain.Alternatively, this profile is measured by means of a DNA chip targeting said markers in a manner known per se. An embodiment has been described in which the database of complete genomes is independent, at least in its first version, of the bacterial ecology of the organization implementing the typing for its own needs. Alternatively, this database consists essentially of bacterial strains originating solely from the organization. A first minimal version of the database can however be provided to the organization and then supplemented with the organization's strains. An embodiment applied to a bacterial species has been described. The invention also applies to yeasts and fungi. Serotyping has been described. Other subtypes of a species can be characterized, such as for example their resistances and / or sensitivities to microbial agents or groups based on typical sequences.A correct assignment rate and a precision, balanced or not, have been described as criteria for choosing the hyperparameters of the prediction. Other criteria are possible, such as sensitivity and specificity or a confidence interval other than 95%.

Claims

CLAIMS 1. An in vitro method for typing a microorganism included in a sample, comprising: d. providing computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the database to positions in the phylogenetic tree, frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; − for at least part of the genetic profiles of the database, the belonging of said genetic profiles to microbial groups distinct from the microbial species, said distinct groups being absent from the phylogenetic tree; e.measuring the genetic profile of the microorganism; f. determining, implemented by computer: − probabilities of variation of the genetic profiles belonging to said groups compared to the measured genetic profile, said probabilities of variation being calculated as a function of the frequencies of variation at a predefined level of the phylogenetic tree; − the belonging of the microorganism to one of the microbial groups if an identification performance criterion is reached for said group, said criterion being reached if at least one probability of variation for said group is greater than a first predefined threshold.

2. Method according to claim 1, wherein: − the phylogenetic tree is constructed from a core genome of the species of the microorganism; − the predefined set of genetic markers is chosen from the accessory genome of the species of the microorganism. 3.Method according to claim 1 or 2, in which the number of markers of the genetic profiles is less than 50, and preferably between 16 and 30.

4. Method according to one of the preceding claims, in which. − the phylogenetic tree is constructed by applying clustering based on genetic distances from a core genome of the microorganism species; − the genetic distances decrease from the root to the leaves of the tree so as to define a level in the tree comprising between 150 and 350 differences in the core genome.

5. Method according to claim 4, wherein − the groups of microorganisms correspond to microbial serotypes of the microorganism species; − the phylogenetic tree comprises at least 10 levels, preferably at least 15 levels; − the probabilities of variation of the genetic profiles belonging to said groups with respect to the measured genetic profile are calculated as a function of the frequencies of variation of a level of the phylogenetic tree comprised between 7 and 15, preferably between 10 and 15, more preferably between 11 and 13. 6.A method according to any preceding claim, wherein the markers of the genetic profiles are selected so as to minimize an entropy or an impurity criterion at a level of the phylogenetic tree comprising between 150 and 350 differences in the core genome.

7. A method according to any preceding claim, wherein the identification performance criterion further comprises the deviation between the probability of variation for said group and the probabilities of variation for the other groups, the identification performance criterion being met if said deviation is greater than a second predefined threshold.

8. A method according to claim 7, wherein the identification performance indicator comprises a correct assignment rate and / or a raw precision and / or a balanced precision. 9.Method according to any one of the preceding claims, comprising the search for racial causes during contamination of a food product 10. Method according to any one of the preceding claims, in which the measurement of the genetic profile of the microorganism is carried out by means of an amplification of target genetic sequences without carrying out complete sequencing, in particular a polymerase chain reaction.

11. System for in vitro typing of a microorganism included in a sample, comprising: d. computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the database to positions in the phylogenetic tree, frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; e. means for collecting and storing a measurement of the genetic profile of the microorganism; f.computer-implemented means for determining: − probabilities of variation of the genetic profiles belonging to said groups with respect to the measured genetic profile, said probabilities of variation being calculated as a function of the frequencies of variation at a predefined level of the phylogenetic tree; − the belonging of the microorganism to one of the microbial groups if an identification performance criterion is reached for said group, said criterion being reached if at least one probability of variation for said group is greater than a first predefined threshold.

12. System according to claim 11, configured to implement a method according to any one of claims 2 to 10.

13. Computer program product recorded on a computer-readable computer medium comprising instructions for executing step c. of a method according to any one of claims 1 to 10. 14.Composition for the in vitro typing of microorganisms included in a sample by amplification of target nucleotide sequences, in particular by a polymerase chain reaction, said composition comprising a set of oligonucleotides for the amplification of at least the set of genetic markers obtained according to claim 6.

15. A method for preparing a biological sample comprising, or capable of comprising, a bacterial strain, for the purpose of typing said strain by a polymerase chain reaction, comprising bringing said sample into contact with a composition according to claim 14.

16. A method according to claim 15, comprising lysing the sample resulting from said contacting so as to release genetic material from the bacteria of the bacterial strain.

17. A method for preparing a sample capable of comprising a microorganism, said method comprising − extracting the genetic material comprised in the sample; − mixing said extracted genetic material with a composition comprising a set of primers for the amplification, in particular by polymerase chain reaction, of a set of genetic markers obtained according to a method according to claim 6.