Method and system for probabilistic typing of microbial strains
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2026-04-08
AI Technical Summary
Current genomic typing methods for microorganisms face challenges in achieving high precision while being cost-effective and time-efficient, particularly in distinguishing bacterial strains at a taxonomic level lower than the species level, as they often require extensive sequencing or a large number of genetic markers, making conventional PCR platforms impractical.
A method and system that uses a limited number of genetic markers to identify microorganisms by constructing a phylogenetic tree from a core genome and selecting markers from the accessory genome, allowing for precise classification using probabilistic typing and amplification platforms like PCR, enabling accurate assignment to several hundred subdivisions with a low number of markers.
This approach allows for high precision in classifying microorganisms with a significantly reduced number of genetic markers, achieving over 80% assignment rate in certain species with confidence intervals, making it feasible for rapid and reliable typing using existing amplification platforms, such as PCR.
Smart Images

Figure EP2024064923_05122024_PF_FP_ABST
Abstract
Description
[0001]METHOD AND SYSTEM FOR PROBABILISTIC TYPING OF MICROBIAL STRAINS FIELD OF THE INVENTION The invention relates to the field of genomic typing of microorganisms, namely the characterization of bacterial strains at a taxonomic level below the species level based on their genomes. The invention finds particular application in the epidemiological monitoring of bacterial strains, the monitoring of contamination by bacterial strains and the search for the root causes of bacterial contamination and / or infection, whether their origin is industrial, environmental, veterinary, clinical or other. STATE OF THE ART The genomic typing of microorganisms, for example bacteria, brings together all the techniques for comparing the genomes of bacterial strains with the aim of characterizing the latter at levels below that of the species.Typing can be used to determine the degree of similarity between two bacterial strains or to classify strains into species-level subdivisions, for example, clonal groups or groups defined by phenotypic characteristics of the bacteria such as their serotypes or their susceptibility to antibiotics. Typing techniques can be classified into two categories: the first based on the analysis of a very restricted targeted portion of the bacterial genome, such as typing using MLST (for "MultiLocus Sequence Typing") or "16S rRNA", and the second based on the comparison of genomes obtained by complete sequencing (or WGS for "Whole Genome Sequencing"), for example, cg-MLST (for "core genome" MLST).While the latter offer the highest level of accuracy, the use of complete sequencing has many disadvantages, including duration, cost, and the need to implement complex bioinformatics processing. Conversely, when the former techniques use a very limited number of markers, they allow the use of polymerase chain reaction (or "PCR") amplification platforms for their detection, which are much faster and less expensive. Regardless of the technique used, state-of-the-art genomic typing is based on the search for and use of conserved portions of the genome to characterize microorganisms.For example, MLST typing is based on housekeeping genes that exhibit high stability over time, while cg-MSLT typing directly targets the entire core genome (or "core genome" in English, these three expressions are used interchangeably). It is observed in this type of approach that obtaining high accuracy requires a very large number of markers. For example, in order to detect clonal groups of the species P. aeruginosa with an accuracy greater than 90%, cg-MLST typing involves several thousand markers, making the use of conventional PCR platforms de facto irrelevant, if not impossible. In summary, a user wishing to perform molecular typing, for example for the purpose of monitoring microbial contamination, must choose between accuracy and speed.STATEMENT OF THE INVENTION The purpose of the present invention is to propose molecular typing using a limited number of genetic markers and with high precision. To this end, the subject of the invention is an in vitro method for identifying a microorganism included in a sample, comprising: a. the provision of computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the database to positions in the phylogenetic tree, frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; b. the measurement of the genetic profile of the microorganism; c.the determination, implemented by computer, for at least one level of the set of predefined levels of the phylogenetic tree: − of a probability of variation of each genetic profile of the database belonging to said level compared to the measured genetic profile, said probability of variation being calculated as a function of the frequencies of variation at said level; − of the belonging of the microorganism to a subdivision of said level if a first identification performance criterion is reached for said subdivision, said first criterion being reached if at least one probability of variation for said subdivision is greater than a first predefined threshold, method in which: − the phylogenetic tree is constructed from a core genome of the species of the microorganism; and − the predefined set of genetic markers is chosen from an accessory genome of the species of the microorganism.The invention is based on the principle that the variation frequencies of the accessory genome, when structured by a phylogenetic tree built on the basic genome, effectively explain the diversity of a microbial species. Applying this principle, the invention consists of building a precise and flexible genomic "sensor" of this information to make it simple, fast and reliable for the user. This sensor measures the profile of a strain to be tested and estimates its probability that it belongs to a taxon of the tree according to the variation frequencies of the profile of this taxon. The invention thus allows classification accuracy in several hundred subdivisions of the species. For example, an assignment rate in an 18-level tree greater than 80% with a confidence interval of 95% can be obtained in Listeria monocytogenes and Salmonella with 23 and 13 markers respectively.This very low number of genetic targets allows the use of amplification platforms such as PCR platforms, for example GENE-UP platforms. ®and FilmArray marketed by the Applicant, or isothermal amplification platforms, or the use of DNA chips. According to one embodiment, if the first identification performance criterion at said level is not reached then step c) is implemented for the level directly above in the phylogenic tree. According to one embodiment, the number of markers of the genetic profiles is less than 50, and preferably between 16 and 30. According to one embodiment: − the phylogenetic tree is constructed by applying clustering based on genetic distances; − the genetic distances decrease from the root to the leaves of the tree so as to define a level in the tree comprising between 150 and 350 differences in the core genome. According to one embodiment, the set of predefined markers is selected so as to minimize an entropy or impurity criterion of said set at a level of the phylogenetic tree.In particular, the markers of the genetic profiles are selected so as to minimize an entropy or an impurity criterion at the level comprising between 150 and 350 differences in the core genome. According to one embodiment, the first identification performance criterion further comprises the difference between the probability of variation for the subdivision and the probabilities of variation for the other subdivisions of said level, the first identification performance criterion being achieved if said difference is greater than a second predefined threshold. In particular, the second identification performance criterion further comprises the difference between the probability of variation for said group and the probabilities of variation for the other groups, the second identification performance criterion being achieved if said difference is greater than a fourth predefined threshold. According to one embodiment, the database comprises pairs of first and second thresholds, and foreach of said thresholds at least one identification performance indicator, said performance indicator comprising at least one indicator related to the specificity and / or the sensitivity of the identification method. In particular, the identification performance indicator comprises a correct assignment rate and / or a raw precision and / or a balanced precision. Advantageously, the performance indicators for the pairs of first and second thresholds are displayed on a screen, in which a user selects a displayed identification performance indicator value, and in which the method sets the first and second thresholds to their values corresponding to the selected performance indicator value. According to one embodiment, the database comprises for at least part of the genetic profiles, the membership of said genetic profiles to distinct microbial groups, said distinct groups being absent from the phylogenetic tree, themethod comprising the determination, implemented by computer: − of probabilities of variation of the genetic profiles belonging to said groups compared to the measured genetic profile, said probabilities of variation being calculated as a function of the frequencies of variation at a predefined level of the phylogenetic tree; − of the belonging of the microorganism to one of the groups of microorganisms if subdivision of said level if a second identification performance criterion is reached for said group, said second criterion being reached if at least one probability of variation for said group is greater than a third predefined threshold. In particular, the second identification performance criterion further comprises the difference between the probability of variation for said group and the probabilities of variation for the other groups, the second identification performance criterion being reached if said difference is greater than a fourth predefined threshold. In particular, themicrobial groups are serotypes or sequence types. According to one embodiment, the method further comprises the search for racial causes during contamination of a food product. According to one embodiment, the measurement of the genetic profile of the microorganism is carried out by means of an amplification of target genetic sequences without implementing complete sequencing, in particular a polymerase chain reaction. The invention also relates to an in vitro typing system for a microorganism included in a sample, comprising: a. computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the database topositions in the phylogenetic tree, of frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; b. means for collecting and storing a measurement of the genetic profile of the microorganism; c. computer-implemented means for determining for at least one level of the set of predefined levels of the phylogenetic tree: − a probability of variation of each genetic profile of the database belonging to said level compared to the measured genetic profile, said probability of variation being calculated as a function of the frequencies of variation at said level; − the belonging of the microorganism to a subdivision of said level if a first identification performance criterion is reached for said subdivision, said first criterion being reached if at least one probability of variation for said subdivision is greater than a first predefined threshold. Such a systemis configured to implement a method of the aforementioned type. The invention also relates to a computer program product recorded on a computer-readable computer medium comprising instructions for executing step c. of a method of the aforementioned type. The invention also relates to a computer program product recorded on a computer-readable computer medium comprising instructions for constructing a phylogenetic tree based on a core genome and the variation frequencies of a set of markers of an accessory genome at levels of the phylogenetic tree. The invention also relates to a composition for the in vitro typing of microorganisms included in a sample by amplification of target nucleotide sequences, in particular by a polymerase chain reaction, said composition comprising a set of oligonucleotides for the amplification of at least the set of markersgenetic material obtained in accordance with the aforementioned method. The invention also relates to a method for preparing a biological sample comprising, or likely to comprise, a bacterial strain, with a view to typing said strain by a polymerase chain reaction, comprising bringing said sample into contact with a aforementioned composition. In particular, the method comprises lysing the sample resulting from said contacting so as to release genetic material from the bacteria of the bacterial strain. BRIEF DESCRIPTION OF THE FIGURES The invention will be better understood from reading the following description, given solely by way of example, and carried out in relation to the appended drawings, in which: − Figure 1 is an overview of a method according to the invention; − Figure 2 is a flowchart detailing a learning phase of a phylogenetic typing according to the invention; − Figure 3 is a plot illustrating a genomic distancedecreasing for the construction of a phylogenetic tree; − Figure 4 is a schematic view of a binary phylogenetic tree; − Figure 5 is a diagram illustrating the calculation of the variation frequencies of a genetic profile according to the invention for a taxon of the phylogenetic tree; − Figure 6 is a flowchart illustrating the production of a PCR kit used for typing according to the invention; − Figure 7 is a flowchart detailing a phase of prediction of a phylogenetic type according to the invention; − Figure 8 is a flowchart illustrating the calculation of thresholds used during the prediction phase; − Figure 9 is a flowchart illustrating a learning phase for serotyping according to the invention; − Figure 10 is a flowchart illustrating a serotype prediction phase according to the invention; − Figure 11 is a flowchart illustrating the calculation of thresholds used during the serotype prediction phase; − Figure 12 is a plot of theShannon entropy decay during the selection of genetic markers constituting a genetic profile for Listeria typing; − Figure 13 is a set of plots of frequencies of variation of genetic profiles in the Listeria phylogenetic tree; − Figure 14 is a plot of the correct serotype assignment rate of Listeria as a function of balanced accuracy; − Figure 15 is a set of plots of correct phylogenetic assignment rate of Listeria as a function of average accuracy; − Figure 16 is a plot of Shannon entropy decay during the selection of genetic markers constituting a genetic profile for Salmonella typing; − Figure 17 is a set of plots of frequencies of variation of genetic profiles in the Salmonella phylogenetic tree; − Figure 18 is a plot of the correct Salmonella serotype assignment rate versus balanced accuracy; − Figure 19 is aset of plots of correct phylogenetic assignment rates of Salmonella as a function of average accuracy; and − Figures 20 to 22 illustrate computer architectures for implementing the typing according to the invention. DETAILED DESCRIPTION OF THE INVENTION A. Definitions In this document, the following definitions apply: − Biological sample: means any sample containing or suspected of containing one or more microorganisms; − Basic genome (or "core genome" or "heart genome", noted "Gb") of a microbial species: means parts of the genome (genes or groups of genes) shared by all the microbial strains belonging to the microbial species(s) considered. For the purposes of the invention, the basic genome can evolve if the set of strains from which it is produced evolves. Preferably, the core genome is constructed from at least 100 different microbial strains, and preferably from at least 1000strains. It will be noted that the genomic portions of microbial strains belonging to the basic genome may differ from one another. For example, a particular gene may belong to the basic genome and have several different alleles. In the following, the expression "basic genome" when applied to a microbial strain corresponds to the portion of the genome of the strain belonging to the basic genome of the species. − Accessory genome of a microbial species (noted "Ga"): designates parts of the genome (genes or groups of genes) of a set of bacterial strains which do not belong to the basic genome of said set. These parts of the genome are therefore absent from the genome of at least one of the bacterial strains of the set of bacterial strains considered. For the purposes of the invention, the accessory genome can evolve if the set of strains from which it is produced evolves. Preferably, the accessory genome is constructed from at least 100 strainsdifferent microbial strains, and preferably from at least 1000 strains. − Genetic marker: designates a gene or a nucleotide sequence at a specific position in the genome (also called a "locus"), a nucleotide sequence which can be reduced to a single nucleotide such as, for example, two alleles differing by a single nucleotide, or to a polymorphism of a single nucleotide in the context of a mutation (or SNP for "single nucleotide polymorphism"); − Microorganism: designates all microscopic organisms including in particular bacteria, yeasts and fungi; − Phylogenetic tree of a microbial species: designates a tree representing evolutionary relationships among a set of microbial strains, relationships determined according to their similarities and differences in their core genomes. For the purposes of the invention, a phylogenetic tree can be enriched if the set of strains from which it is produced evolves. Preferably a treephylogenetic tree is constructed from at least 100 different microbial strains, and preferably from at least 1000 strains. By convention, the root of a phylogenetic tree is designated as the first level of the tree and the leaves of the tree as the last levels. A level "higher" than another thus corresponds to a level closer to the root. B. Embodiment of the invention A probabilistic typing according to the invention is now described, applied to the monitoring of contamination by bacterial strains of a species within an organization, in particular in premises for the production of food intended for human consumption, as well as the detection of their serotypes. Referring to Figure 1, the typing according to the invention comprises the following main phases: - a phase 10 of constituting a database of complete genomes of strains belonging to the bacterial species; - a phase 20 called "learning" or"training" phase, aimed at generating a phylogenetic tree based on the genomes in the database, choosing an appropriate genetic profile and frequencies of variation of said genetic profile along the tree; - a phase 30 called "prediction" performing the typing of an unknown bacterial strain based on the tree and the frequencies produced during phase 20; and - phases possibly completed by a phase 40 called "epidemiological" during which corrective and / or prophylactic measures are implemented within the organization based on the result of the typing of the unknown strain. Preferably, the database of complete genomes is as representative as possible of the genetic diversity of the species. It is advantageously completed over time by adding complete genomes so as to target this diversity if its initial version is incomplete. In this case, the typing according to the invention is implemented when said databasecomprises at least 100 genomes of the species, preferably at least 1000 genomes and even more preferably several thousand genomes. Referring to Figures 2, 3 and 4, the training phase 20 begins with the production of a base genome at 202 and an accessory genome at 204 from the genomes of the database 200. In a first variant, the genomes include annotations categorizing the genes as belonging to one or the other of these genomes. In a second variant, the latter are produced ad hoc from the database, the genes having an allele frequency greater than a threshold, for example 95%, belonging to the base genome, otherwise to the accessory genome. In a third variant, the base and accessory genomes are produced from the database by comparison with a reference genome of the species, stored in the database, the genes having an allele frequency compared to the reference genomegreater than a threshold, for example 95%, belonging to the basic genome, otherwise to the accessory genome. In a fourth variant, the database is already constituted for the species as well as a version of its basic and accessory genome. The typing according to the invention continues, in 206, by the generation of a phylogenetic tree based on the basic genome. Several techniques for constructing such a tree are possible, as described for example in the documents "Phylogenetics Algorithms and Applications" (Hu et al., Ambient Communications and Computer Systems, 2019), "A Review: Phylogeny Construction Methods " (Shaktawat et al., International Journal of Emerging Science and Engineering, 2019) or "Essential Bioinformatics" (chapter 11 of Jin Xiong, Cambridge University, 2011). In a variant of the invention, the tree is constructed from a cg-MLST (for "core-genome multi locus sequencing typing") coding of the basic genomes. These are translated into profilescg-MLST using, for example, BioNumerics software (bioMérieux, Marcy l'Etoile, France), SeqSphere+ (Ridom, Münster, Germany), BIGSdb (available at https: / / github.com / kjolley / BIGSdb and described in Schürch, A. et al. "Whole genome sequencing options for bacterial strain typing and epidemiologic analysis based on single nucleotide polymorphism versus gene-by-gene–based approaches." Clin. Microbiol. Infect., 2018), or chewBBACA (available at https: / / github.com / B-UMMI / chewBBACA and described in Silva M et al., "chewBBACA: a complete suite for gene-by-gene schema creation and strain identification." Microb. Genom., 2018). The advantage of cg-MLST profiles is that they are fixed-length sequences of a few hundred components, aligned with each other, allowing the direct use of tree construction methods. This tree is advantageously generated by applying a succession of groupings using the simple link method.("Single-linkage clustering" in English), and therefore based on genetic distances between core genomes, for example calculated using the normalized Hamming distance. The subdivision of a taxon of the tree then consists of grouping into the same subgroup the core genomes differing from each other by a distance lower than a predefined threshold. For example, software such as HierCC (accessible at https: / / github.com / zheminzhou / pHierCC and described in the document "HierCC: a multi-level clustering scheme for population assignments based on core genome MLST" by Zhou et al., Bioinformatics, 2021) and Pathogenwatch from the Center for Genomic Pathogen Surveillance (accessible at https: / / pathogen.watch / ) can be used. Once the tree has been constructed, the latter is stored in a database 208. In a particularly preferred variant of the invention, the number of levels of the tree and the clustering distances are chosen so as to produce at leastleast one level in the tree such that the core genome diversity at this level matches or is close to the clonal diversity observed in the species, so that the taxa at this level make it possible to distinguish different clonal complexes among the bacterial strains corresponding to the complete genomes in the database. For example, for Salmonella and Listeria, the level has between 150 and 350 differences in the core genome of the strains, preferably between 200 and 300 differences. In particular, the clustering distance decreases from the root to the leaves, for example linearly. The learning phase 20 continues with the choice, at 210, of a type of genetic profile in the accessory genome, namely a set of genetic markers of the accessory genome of the species whose absence and presence are measured in a strain to be tested during the prediction phase 30. The choice of the profile consists of choosing a level in the tree, andpreferentially one of the levels representing clonal diversity, then to choose a set of markers having a minimum entropy or degree of impurity at this level. The total number N M of markers selected to construct the genetic profile is advantageously selected according to the particular PCR platform which is used to carry out the typing. For example, for the applicant's GENE-UP platform, this number N Mis equal to 23 markers for Salmonella and 16 markers for Listeria monocytogenes. In practice, the set of markers is chosen so as to minimize an entropy criterion (e.g., Rényi, Hartley, Shannon, collision, or min entropy) or a Gini impurity criterion, by implementing a greedy algorithm. As an illustration, and referring to the phylogenetic tree in Figure 4, the selection of a set of markers minimizing the Shannon entropy criterion of level N consisting of Q N = 20 taxa T1, T2,…, T 20 consists of choosing a first marker M1 of the accessory genome minimizing the entropy criterion: ^ ^^^ ^^^^ où ^^^^ ( ^^^^ ) = ^^^^ ^^^^ ^^^^ ^^^^ with NG the total number of core genomes used to build the tree and ^^^^ ^^^^the number of times the marker is present in accessory genomes in the i ème taxon T i . Once the M1 marker is selected, this procedure is relaunched to select the second M2 marker in the accessory genome removed from the M1 marker. Thus the j ième marker M j is the marker such that ^ ^^^ ^^^^ ^^^^ ( ^^^^ ) = ^^^^ ^^^^ ^^^^ ^^^^ with NG the total number of core genomes used to build the tree and ^^^^ ^^^^ is the number of times the marker ^^^^ ^^^^ is present in accessory genomes in the i ème taxon T i . Alternatively, the stepwise selection of markers is performed by minimizing the conditional entropy of level N according to the formula: ^^^^ ^^^^ 2 ^^^^ Or: − ^^^^ ( ^^^^ ) = ^^^^ ^^^^^^^^ ^^^^ with NG the total number of core genomes used to build the tree and ^^^^ ^^^^ the number of times the marker is present in accessory genomes in the i ème taxo − and ^^^^ ( ^^^^ ^^^^ | ^^^^ ) = ^ n T i ^ ^^ ; ^ ^^^ ^ ^^^ ^^^^ with ^^^^ ^^^^ the number of times that the l ième value of the marker set � ^^^^1, ^^^^2, … , ^^^^ ^^^^−1, ^^^^ ^^^^ �, among the 2 ^^^^ possible values of said set, is present in the taxon T i and ^^^^ ^^^^ the number of core genomes included in taxon T i. The inventors have observed that the typing according to the invention requires at most 50 markers, preferably less than 30 markers, more particularly a number between 16 and 30 markers, to obtain high typing performance. The learning phase 20 continues, in 212, by the extraction of the accessory genomes from the genetic profiles consisting of the selected markers, profiles stored in a database 214. The genetic profile of an accessory genome is therefore a vector P of {0,1} ^^^^ ^^^^ , the j ième component P(j) of the vector P corresponding to the value of the marker ^^^^ ^^^^ , this value being equal to 0 when the marker ^^^^ ^^^^is absent and to 1 when it is present. Once the profile base 214 is constituted, frequencies of variation of the profiles along the tree are calculated at 216. Referring to Figure 5, for each taxon T of the tree which is partitioned into subdivisions T1, T2, …, TQ, there is a set of pro^ ^^^ fils ^ ^^^ génétiq ^ u ^ ^^ es {^^^^ , ^^^^ , … , ^^^^} associated with taxon T partitioned into subsets 1 1 11 2 ^^^^ s of profiles� ^^^1^ , ^^^2^ , … , ^^^ ^^^^^1�,^^^^ ^^^^2 , ^^^^ ^^^^2 , … , ^^^ ^ ^ ^^^^2 ^ ^^ ^^^^ ^^^^ ^^^^ , ^^^^ ^^^^ ^^^^ , … , ^^^^ ^^^^ ^^^^ respectively associated with the subtaxa T1, T, …, genetics, a frequency of ^^^^( ^^^^ ^^^^| ^^^^) of this marker is calculated for the taxon T as being equal to the number, divided by Q, of sub-taxa T1, T2, …, T Q for which the marker M jtakes both the value 0 and the value 1, namely a frequency according to the relation: ^^^^ ^^^^ ^^^^ ^^^^ ^^^^ ^^^^ \∃ ( ^^^^, ^^^^), ^^^^ ≠ ^^^^, ^^^^ ^^^^ ^^^^ ^ ^^^ ( ^^^^) = 0 & ^^^^ ^^^^ ^^^^ ^ ^^^ ( ^^^^) = 1 where ^^^^ ^^^^ ^^^^ ^^^^ ( ^^^^ ) denotes the cardinality function of a set E. The calculation of frequencies thus produces for each subdivided taxon T, a vector of frequencies ^^^^ ( . | ^^^^ ) de [ 0,1 ] ^^^^ ^^^^ tel que ^^^^(. | ^^^^)( ^^^^) = ^^^^� ^^^^ ^^^^ � ^^^^�. Once calculated, the frequency vectors are stored in a database The learning phase 20 ends with the calculation, at 220, of a grid of hyperparameters used during the prediction phase 30, a calculation which will be detailed below once the prediction 30 has been described, the parameters being stored in a database 222. In summary, the learning phase 20 provides a phylogenetic tree in which each complete genome and its genetic profile extracted from its accessory genome is positioned in the tree and each taxon of the tree, excluding the leaves, is associated with a frequency of variation of each genetic marker of the profile.Referring to Figure 6, the invention comprises the design and production 50 of a composition, or "kit", for carrying out an amplification aimed at detecting the selected markers, comprising: - the selection, at 500, of the amplification primers for the amplification of the selected genetic markers: - the production of said primers at 502 and if necessary the production of detection probes so as to detect the presence of a marker in at least one of its variants, and preferably in the maximum of its variants; - the production, at 504, of a buffer solution comprising the selected primers, the detection probes if necessary, dNTPs (deoxyribonucleotide triphosphates) providing the energy and the nucleotide bases necessary for the production of amplicons, an amplification enzyme and salts such as Mg ions. 2+or NaCl allowing the enzyme to function properly, at appropriate concentrations. Dyes can be added to this solution. − the production of a lysis buffer comprising beads, for example magnetic or ceramic beads, possibly a dye. The kits produced are then used during the prediction phase 30. Referring to Figure 7, this phase begins with the collection 300 of a biological sample in the organization, for example a food product production plant. During a following step 302, the sample is prepared so as to detect a given pathogen by methods known in the state of the art such as a culture on a dish, an immunological detection of the VIDAS type ® , or even by molecular biology for example with the GENE-UP platform ®. Following the detection of said pathogen, a confirmation step can take place. At the same time, if the colony is identified as belonging to the species, the prediction phase 30 continues at 302 using isolated colonies or an enrichment broth for its characterization by PCR amplification. In particular, the sample undergoes a set of preliminary steps, such as lysis to release the DNA of the bacteria constituting it, optionally extraction of the DNA resulting from the lysis using beads, for example magnetic or ceramic beads. Once the DNA of the bacterial strain has been released, it is added, at 304, to the PCR kit designed to target at least the genetic markers selected during the learning step 20. Finally, the PCR is implemented at 306 in order to measure the genetic profile of the strain, for example using the GENE-UP platform. ® of the Applicant. The measured profile, noted P mes, is then analyzed, at 308, in order to position the bacterial strain in the phylogenetic tree. Starting from a level N of the tree, for example the level directly above the leaf level, the analysis 308 includes the following steps: a. for each taxon of level N, noted ^^^^^^^^, ^^^^(i ème N taxon ième level): a1. for each stored profile, noted ^^^^^^^^ ^^^^ ^^^^, from the database 214 which is assigned to the taxon ^^^^^^^^, ^^^^, the calculation of a probability ^^^^^^^^ ^^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^^^^^ ^^^^ ^^^^) than the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^either a variation of the memorized profile ^^^^^^^^ ^^^^ ^^^^according to the relationship: ^^^^^^^^ ^^^^^^^^ ^ ^ ^^^ ^^^^ ^^^^ ^^^^ ^^^^ = ^^^^ ^^^^ ^^^^ ^^^^ ^^^^ We will note that if the measured profile ^^^ ^^^^^ ^^^^ ^^^^is identical to the profile ^^^^^^^^ ^^^^^ then this probability is equal to 1 (no variation) and decreases rapidly if the profile ^^^ ^ ^^^^ ^^^^ ^^^^deviates from the profile ^^^^^^^^ ^^^^ ^^^^, and this is all the more so as the frequencies ^^^^^^^^ ^^^^^ ^^^^� ^^^^ ^^^^ � ^^^^^^^^, ^^^^� are weak. a2. selection of the maximum ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^( ^^^^, ^^^^ ) probabilities ^^^^^^^^ ^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^^^^^ ^^^^ ^^^^) calculated for the taxon ^^^^^^^^, ^^^^. b. Assignment, in 310, of the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^to one of the taxa ^^^ ^ ^ ^^^, ^^^^ of level N such that: ^^^ ^ ^ ^^^, ^^^^ = ^^^^ ^^^^ ^^^^max ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^^^^^ ( ^^^^, ^^^^) ^^^^ ^^^ where ^^^^1 and ^^^^2 are predetermined non-zero positive thresholds. c. If no taxon at level N meets the criterion of step b, move to the next level N-1 of the tree and implement steps a and b for this new level. Once the profile ^^^ ^^^^^ ^^^^ ^^^^assigned to a level of the phylogenetic tree and a taxon of this level, the prediction phase 30 delivers a report to the user, for example in the form of a display on a screen and / or files stored on a computer and / or a message (e.g. email) for the user. If two taxa have identical results at the end of step b, these two taxa are included or an inconclusive typing is noted in the report. The position of the new bacterial strain tested in the phylogenetic tree makes it possible to type the latter in relation to the bacterial strains corresponding to the complete genomes of the database 200. In addition to this type in reference to strains previously observed and capitalized in the different databases, the typing according to the invention allows the relative typing of two bacterial strains if these two strains are assigned to the same taxon.In addition, the invention makes it possible to define the lineage of a strain in relation to those of the databases or in relation to another strain tested. Indeed, the relative position of two strains in the tree makes it possible to characterize their evolution in relation to each other. A path through the phylogenetic tree has been described from the leaves to the root: as long as the assignment relationships are not verified, the tree is ascended. Alternatively, the phylogenetic tree is traversed from the root to the leaves: as long as the assignment relationship is verified, the tree is descended. The values of the thresholds ^^^^1 and ^^^^2 are hyperparameters determined during step 220 of the learning phase. preferably by means of cross-validation. Referring to Figure 8, a grid of values for thresholds ^^^^1 and ^^^^2 is first determined, and for each pair of thresholds in the grid, a of the correct assignment rate and the average precision is implemented by: − the partition of the complete genome base 200 into a training base 800 and a test base 802 according to a cross-validation strategy 804, for example a so-called 10-fold strategy ("ten-folds cross validation" in English); − the implementation of the learning phase 20 on the training base 802; − the implementation of the prediction phase 30 for each genetic profile from the test base 802 and the calculation in 806 of the average precision ACC and the correct assignment rate TA of the genetic profiles of the base 802 in the phylogenetic tree; − the selection, in 808, of the maximum value of the average precisions and the correct assignment rates among the folds of the cross-validation strategy. These two values therefore correspond respectively to the average precision and the correct assignment rate of the phylogenetic prediction phase 30 for the thresholds ^^^^1 and ^^^^2.The pairs of thresholds presenting the best precision and rates are retained. the phylogenetic typing according to the invention. Alternatively, the user can choose the precision and / or the assignment rate that he desires and the phylogenetic typing selects the pair of thresholds whose precision and rate are closest to those chosen by the user. Preferably, the number of levels in the phylogenetic tree and the law of decrease of the clustering distances, for example affine, are also hyperparameters, the method described in relation to figure 8 being therefore completed by the selection of a grid for these two parameters, and the traversal of said grid jointly with the traversal of the grid of thresholds ^^^^1 and ^^^^2. A preferred variant of the invention is now described, which consists of predicting the membership of a bacterial strain in a particular group of the bacterial species based on its measured genetic profile. In the illustrated example, these groups are serotypes, but this variant applies to any explicit subdivision at a level of the bacterial species. It will be noted that this subdivision of the species does not correspond to the first level of the phylogenetic tree and is defined independently of the latter. Referring to Figure 9, the learning phase 20 described above in relation to Figure 2 is completed by a serotype learning 20A comprising: a. the collection of the serotypes of all or part of the bacterial strains corresponding to the complete genomes, serotypes stored in a database 224; b.the distribution, in 226, of the observed genetic profiles of the database 214 of the serotyped strains among the NS different serotypes ST1, ST2,…ST. NS . The genetic profiles associated with a serotype are hereinafter referred to as "serotyped"; c. the calculation, 228, of the frequencies of variation of the serotyped genetic profiles. In particular, for each marker M j , its variation frequency ^^^^� ^^^^ ^^^^ � ^^^^ ^^^^ ^^^^ � in a STi serotype is equal to the number of times this marker takes the total number of genetic profiles assigned to the STi serotype;. d. storage in a database 230 of the variation frequencies ^^^^� ^^^^ ^^^^ � ^^^^ ^^^^ ^^^^ � serotyped genetic profiles in the different serotypes. If the elements of the learning phase 20 described in relation to Figure 2 remain unchanged, meaning in particular that the type of genetic profiles used is determined according to the phylogenetic tree and not the subdivision into serotypes, the prediction phase 30 can take two forms. In the first, the phase 30 already described is supplemented by a prediction of the serotype of a new strain to be tested, in parallel with steps 308-312 of Figure 2. In the second form, only said prediction is solely implemented and replaces steps 308-312. Referring to Figure 10, the prediction 30A of the serotype of a new bacterial strain follows steps 300-306 measuring the genetic profile of a bacterial strain of the species contained in a biological sample, and comprises the following steps: a. for each serotype, noted ST i : a1. for each serotyped genetic profile, noted ^^^^ ^^^^ ^^^^ , from the ST group ifrom database 228, the calculation of a probability ^^^^^^^^ ^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^ ^^^^ ^^^^ ) than the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^either a variation of the memorized profile ^^^^ ^^^^ ^^^^ according to the relationship: ^^^^^^^^ ^^^^^^^^ ^ ^ ^^^ ^^^^ ^^^^ ^^^^ = ^^^^ ^^^^ ^^^^ ^^^^ ^^^^ ^^^^ a2. selection of the maximum ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^( ^^^^ ^^^^ ^^^^ ) probabilities ^^^^^^^^ ^^^^ ^^^^( ^^^ ^ ^^^^ ^^^^ ^^^^, ^^^^^^^^ ^^^^ ^^^^) calculated for the ST serotype i ; b. Assignment of the measured profile ^^^ ^ ^^^^ ^^^^ ^^^^to one of the ST serotypes j such as: ^^^^ ^^^ ^ ^ ^^^ = ^^^^ ^^^^ ^^^^max ^^^^^^^ ^^^^ ^^^^( ^^^^ ^ ) ^^^^^^^ ^^^^ ^^^^^^^ ^^^^ where ^^^^3 and ^^^^4 are predetermined non-zero positive thresholds; c. Production and delivery, at 318, of a report on the serotype detected for the bacterial strain subject to serotype typing. In a variant of serotyping 30A, if the second condition ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ � ^^^^ ^^^ ^ ^ ^^^ � − ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^^ ( ^^^^ ^^^^ ^^^^ ) ≥ ^^^^4 is not satisfied, serotyping returns the first two or three serotypes associated with the probabilities ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^( ^^^^ ^^^^^^^^ )the highest, as well as said probabilities. The values of the thresholds ^^^^3 and ^^^^4 are hyperparameters determined during step 220 of the training phase 30, preferably by means of cross-validation. Referring to Figure 11, a grid of values for the thresholds ^^^^3 and ^^^^4 is first determined, and for each pair of thresholds in the grid, a determination of the correct assignment rate and the average accuracy for serotype typing is implemented by: − partitioning the complete genome databases 200 and serotypes 224 into a training database 1100 and a test database 1102 according to a cross-validation strategy 1104,for example a so-called 10-fold strategy ("ten-folds cross validation" in English); − the implementation of the learning phase 20 - 201 on the training base 1102; − the implementation of the prediction phase 30A for each genetic profile from the test base 1102 and the calculation in 1106 of the average precision ACC and the correct assignment rate TA of the genetic profiles of the base 802 in the serotypes; − the selection, in 1108, of the maximum value of the average precisions and the correct assignment rates among the folds of the cross-validation strategy. These two values therefore correspond respectively to the average precision and the correct assignment rate of the serotype prediction phase 30A for the thresholds ^^^^3 and ^^^^4. The pairs of thresholds presenting the best precision and rates are retained for serotype typing according to the invention. Alternatively,the user can choose the precision and / or the assignment rate that he desires and the serotyping according to the invention selects the pair of thresholds whose precision and rate is closest to those chosen by the user. In the variant just described, the determination of the serotype of a strain to be tested comprises the calculations as described in relation to figures 9 and 10. In another variant, for each taxon of the phylogenetic tree, the probability of each serotype in this taxon is calculated according to the serotype database 224. For example, the probability of a given serotype ^^^^ ^^^, ^ ^ ^^^ in the taxon T iis equal to the number of genomes grouped in this taxon that exhibit this serotype divided by the total number of genomes in the taxon. These probabilities are stored in a reference table or dictionary, and predicting the serotype of a strain to be tested involves first determining the taxon in the phylogenetic tree to which it belongs in the manner described in steps 308-312. Once the taxon has been identified, serotyping the strain involves querying the probabilities in the reference table associated with said taxon and delivering at least the serotype with the highest probability to the user. In the typing and serotyping variants just described, a calculation of the values ^^^^^^^^ ^^^^ ^^^^, ^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^^ , and assignment of the measured profile ^^^ ^^^^^ ^^^^ ^^^^is implemented once this profile is measured. Alternatively, the already known profiles are associated directly with the taxon or taxa of the binary phylogenetic tree, for example by means of a reference table. Indeed, having previously carried out the assignment calculation, the taxa are already known. In particular, the profiles ^^^^^^^^ ^^^^ ^^^^ are assigned to their associated taxa by means of the calculations previously described. Thus, if the profile of a new strain is identical to a profile ^^^^^^^^ ^^^^ ^^^^, the reference table is called for this profile and the associated taxa are directly output. Similarly, when a new profile ^^^ ^^^^^ ^^^^ ^^^^is measured and different from all the profiles in the reference table, its assignment is calculated as described above, then this new profile is stored with the taxon(s) identified in the reference table. In this way, we avoid implementing assignment calculations that have already been performed by a simple call to a reference table. Alternatively, rather than a reference table, each already known profile is stored in a memory space dedicated to the taxon(s) of the phylogenetic tree. In other words, the taxa in the tree are assigned "addresses" equal to the known and assigned profiles. D. Examples of implementation D.1. Listeria monocytogenes 37,000 strains of Listeria were collected, completely sequenced and serotyped. The number of levels in the tree is equal to 23, i.e. the total number of targets of the GENE-UP PCR platform ®of the Applicant, and the distance decay law is selected for clustering so as to obtain a diversity similar to the clonal diversity at a level between 10 and 15 in a phylogenetic tree with 24 levels. Figure 12 illustrates the decay of the Shannon entropy of the genetic profile as markers are selected in the accessory genome at level 13 of the phylogenetic tree, the entropy of the final profile being only a few percent of the free entropy at this level. Figure 13 illustrates the frequencies of variation of the profile at the different levels of the tree, with a curve in a graph corresponding to a taxon of the level and the bold curve corresponding to the frequency of the labeled marker.As can be seen in Figure 14, which illustrates the serotype typing among 12 known Listeria serotypes, for a correct serotype assignment rate of 100%, the average balanced accuracy is equal to 90%, demonstrating the effectiveness of this typing according to the invention. It should be noted that it is possible to choose a higher accuracy for a satisfactory lower assignment rate. In particular, for a rate of 80%, the accuracy is close to 95% with the 95% confidence interval contained entirely in the last decile. Figure 15 illustrates the average accuracy of a level as a function of the correct phylogenetic assignment rate. Here again, we note the excellent level of performance of the invention, these two values being respectively greater than 95% and 70% for any level above level 15, and each close to 100% for levels above level 9. D.2. Salmonella 39,000 Salmonella strains were collected, completely sequenced and serotyped.The number of levels in the tree is equal to 23, which is the total number of targets in the GENE-UP PCR platform. ®of the Applicant, and the law of decreasing distances is selected for clustering so as to obtain a diversity similar to the clonal diversity at a level between 10 and 15 in a phylogenetic tree comprising 24 levels. Figure 16 illustrates the decrease in the Shannon entropy of the genetic profile as markers are selected in the accessory genome at level 13, the entropy of the final profile being substantially zero. Figure 17 illustrates the frequencies of variation of the profile at the different levels of the tree, a curve in a graph corresponding to a taxon of the level and the curve in bold corresponding to the average frequency of the level. As can be seen in Figure 18, which illustrates the serotype typing among the more than 2,600 known Salmonella serotypes, for a correct serotype assignment rate of 100%, the average balanced accuracy is greater than 90%, demonstrating the effectiveness of this typing according to the invention.It should be noted that it is possible to choose a higher precision for a satisfactory lower assignment rate. In particular, for a rate of 80%, the precision is close to 100% with the 95% confidence interval contained entirely in the last decile. Figure 19 illustrates the average precision of a level as a function of the correct phylogenetic assignment rate. Here again, we note the excellent level of performance of the invention, these two values being respectively greater than 95% and 80% for any level higher than level 13. E. Preventive and curative measures The solution's ability to easily and quickly identify strains makes it possible to consider its use for the search for root causes during contamination of a finished food product identified during routine controls.This includes, in particular, the search for the strain in the environment of the factory concerned, using surface samples using, for example, swabs, sponges or cloths. The search for the strain can also be carried out in the raw materials involved in the design of the finished product concerned, in order to link the contamination to one or more suppliers and / or carrier(s) and / or storage location(s). It can also make it possible to identify contamination during the laboratory detection phase, either by a laboratory strain or by their DNA, which would have biased the release test that identified the presence of the species concerned. In this case, it would make it possible to establish a plan for retesting the finished product and avoid its destruction.Finally, the solution can be used during a health crisis to quickly characterize the responsible strain(s) and link them to strains identified in incriminated food, cosmetic, veterinary, or other products. Its usefulness therefore covers internal needs of a production plant, for example of food products, investigation needs during health crises or during on-site regulatory audits. The use of this solution thus makes it possible to rationalize and implement appropriate preventive and / or corrective measures to reduce or even eliminate the risks of recurrence of contaminations, and consequently of infections. F.Hardware computer implementation of the invention The learning phase 20, 20A and the prediction phase 30-30A, apart from the steps of sample preparation and measurement of the genetic profile of the strains to be tested, are implemented by computer, namely by means of hardware circuits comprising computer memories (cache, RAM, ROM, etc.) and one or more microprocessors or processors, organized or not in the form of calculation nodes, necessary for the execution of computer instructions stored in the memories for the implementation of said phases. Several architectures are possible, as for example illustrated in figures 20-22. In a first architectural variant (fig.20), a first organization 2000, for example the Applicant, hosts or controls one or more calculation servers 2002 associated with one or more databases 2002 for storing genomes, the phylogenetic tree, phylogenetic frequencies, observed genetic profiles, serotypes and serotype frequencies and the learning phase 20, 20A is implemented by the organization 2000 on its server(s) 2002. A second organization 2006, for example a microbiological laboratory, hosts a PCR platform 2008 connected to, or incorporating, a computer 2010 and implements part of the prediction phase 30, namely all of the steps up to the measurement of the genetic profile of a bacterial strain to be tested. The measured profile is then pushed, through a remote connection network, to the first organization 2000 for implementation on the server 2002 of the remainder of the prediction phase 30, 30A.The report produced is then pushed, through the network 2012, to the second organization 2006 which takes or does not take epidemiological measures depending on the report. A second architectural variant (fig. 21) differs from the first, in that a copy of the database 2004 and the software implementing the prediction phase 30, 30A are downloaded into the second organization which implements, using the computer 2010 or a computing server (not shown), the entire prediction phase 30, 30A. In a third variant (fig. 22), all of the learning phases 20, 20A and prediction 30, 30A are implemented by a single organization 2006. G. Extension of the teaching of the embodiment An embodiment of the invention has been described using PCR to measure the genetic profile of a bacterial strain.Alternatively, this profile is measured by means of a DNA chip targeting said markers in a manner known per se. An embodiment has been described in which the database of complete genomes is independent, at least in its first version, of the bacterial ecology of the organization implementing the typing for its own needs. Alternatively, this database consists essentially of bacterial strains originating solely from the organization. A first minimal version of the database can however be provided to the organization and then supplemented with the organization's strains. An embodiment applied to a bacterial species has been described. The invention also applies to yeasts and fungi. Serotyping has been described. Other subtypes of a species can be characterized, such as for example their resistances and / or sensitivities to microbial agents or groups based on typical sequences.A correct assignment rate and a precision, balanced or not, have been described as criteria for choosing the hyperparameters of the prediction. Other criteria are possible, such as sensitivity and specificity or a confidence interval other than 95%.
Claims
CLAIMS 1. An in vitro method for typing a microorganism included in a sample, comprising: a. providing computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the database to positions in the phylogenetic tree, frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; b. measuring the genetic profile of the microorganism; c.the determination, implemented by computer, for at least one level of the set of predefined levels of the phylogenetic tree: − of a probability of variation of each genetic profile of the database belonging to said level compared to the measured genetic profile, said probability of variation being calculated as a function of the frequencies of variation at said level; − of the belonging of the microorganism to a subdivision of said level if a first identification performance criterion is reached for said subdivision, said first criterion being reached if at least one probability of variation for said subdivision is greater than a first predefined threshold. method in which: − the phylogenetic tree is constructed from a core genome of the species of the microorganism; and − the predefined set of genetic markers is chosen from an accessory genome of the species of the microorganism. 2.Method according to claim 1, wherein if the first identification performance criterion at said level is not met then step c) is implemented for the level directly above in the phylogenetic tree.
3. Method according to any one of the preceding claims, wherein the number of markers of the genetic profiles is less than 50, and preferably between 16 and 30.
4. Method according to any one of the preceding claims, wherein: − the phylogenetic tree is constructed by applying clustering based on genetic distances; − the genetic distances decrease from the root to the leaves of the tree so as to define a level in the tree comprising between 150 and 350 differences in the core genome.
5. Method according to any one of the preceding claims, wherein the set of predefined markers is selected so as to minimize an entropy or impurity criterion of said set at a level of the phylogenetic tree.
6. Method according to claims 4 and 5, wherein the markers of the genetic profiles are selected so as to minimize an entropy or an impurity criterion at the level comprising between 150 and 350 differences in the core genome. 7.
8. Method according to any one of the preceding claims, wherein the first identification performance criterion further comprises the deviation between the probability of variation for the subdivision and the probabilities of variation for the other subdivisions of said level, the first identification performance criterion being met if said deviation is greater than a second predefined threshold.
8. Method according to claim 7, wherein the second identification performance criterion further comprises the deviation between the probability of variation for said group and the probabilities of variation for the other groups, the second identification performance criterion being met if said deviation is greater than a fourth predefined threshold. 9.Method according to any one of the preceding claims, in which the database comprises pairs of first and second thresholds, and for each of said thresholds at least one identification performance indicator, said performance indicator comprising at least one indicator linked to the specificity and / or the sensitivity of the identification method.
10. Method according to claim 9, in which the identification performance indicator comprises a correct assignment rate and / or a raw precision and / or a balanced precision.
11. The method of claim 9 or 10, wherein the performance indicators for the pairs of first and second thresholds are displayed on a screen, wherein a user selects a displayed identifying performance indicator value, and wherein the method sets the first and second thresholds to their values corresponding to the selected performance indicator value. 12.Method according to any one of the preceding claims, in which the database comprises for at least part of the genetic profiles, the membership of said genetic profiles to distinct microbial groups, said distinct groups being absent from the phylogenetic tree, the method comprising the determination, implemented by computer: − of probabilities of variation of the genetic profiles belonging to said groups compared to the measured genetic profile, said probabilities of variation being calculated as a function of the frequencies of variation at a predefined level of the phylogenetic tree; − of the membership of the microorganism to one of the groups of microorganisms if subdivision of said level if a second identification performance criterion is reached for said group, said second criterion being reached if at least one probability of variation for said group is greater than a third predefined threshold. 13.Method according to claim 12, wherein the second identification performance criterion further comprises the deviation between the probability of variation for said group and the probabilities of variation for the other groups, the second identification performance criterion being achieved if said deviation is greater than a fourth predefined threshold.
14. Method according to claim 12 or 13, wherein the microbial groups are serotypes.
15. Method according to any one of the preceding claims, comprising the search for racial causes during contamination of a food product.
16. Method according to any one of the preceding claims, wherein the measurement of the genetic profile of the microorganism is carried out by means of an amplification of target genetic sequences without carrying out complete sequencing, in particular a polymerase chain reaction.
17. System for in vitro typing of a microorganism included in a sample, comprising: d. computer storage means comprising: − a database of genetic profiles of microbial strains belonging to the microbial species of said microorganism, said profiles consisting of the presence or absence of a predefined set of genetic markers; − a phylogenetic tree of the species of said microorganism, assignments of the genetic profiles of the database to positions in the phylogenetic tree, frequencies of variation of the genetic profiles of the microbial strains of the database for predefined levels of the phylogenetic tree; e. means for collecting and storing a measurement of the genetic profile of the microorganism; f.computer-implemented means for determining, for at least one level of the set of predefined levels of the phylogenetic tree: − a probability of variation of each genetic profile of the database belonging to said level with respect to the measured genetic profile, said probability of variation being calculated as a function of the frequencies of variation at said level; − the belonging of the microorganism to a subdivision of said level if a first identification performance criterion is reached for said subdivision, said first criterion being reached if at least one probability of variation for said subdivision is greater than a first predefined threshold.
18. System according to claim 16, configured to implement a method according to any one of claims 2 to 16.
19. Computer program product recorded on a computer-readable computer medium comprising instructions for executing step c.of a method according to any one of claims 1 to 16.
20. Computer program product recorded on a computer-readable computer medium comprising instructions for the construction of a phylogenetic tree based on a core genome and the variation frequencies of a set of markers of an accessory genome at levels of the phylogenetic tree.
21. Composition for the in vitro typing of microorganisms included in a sample by amplification of target nucleotide sequences, in particular by a reaction of. polymerase chain reaction, said composition comprising a set of oligonucleotides for the amplification of at least the set of genetic markers obtained according to claim 5 or 6.
22. Method for preparing a biological sample comprising, or likely to comprise a bacterial strain, with a view to typing said strain by a polymerase chain reaction, comprising bringing said sample into contact with a composition according to claim 21.
23. Method according to claim 22, comprising lysis of the sample resulting from said contacting so as to release genetic material from the bacteria of the bacterial strain.