Method for identifying molecules in a complex mixture and related system

The method automatically identifies molecules in complex mixtures by integrating the results from multiple software modules, addressing the inefficiencies and unreliabilities of current methods and achieving improved accuracy and speed.

WO2025104664A1PCT designated stage expired Publication Date: 2025-05-22NAICONS SRL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/061360
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-11-14
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Current methods for identifying molecules in complex mixtures are inefficient and unreliable due to the high number of molecules present, requiring unacceptably long times for manual annotation and relying on limited and non-customizable databases.

Method used

A method that automatically identifies molecules in complex mixtures by comparing the results from multiple software modules (CD, MD, MQ) using a specific implementation logic, allowing for quick and reliable identification even in mixtures with a very large number of different molecules.

Benefits of technology

The method significantly improves the accuracy and reliability of molecule identification in complex mixtures, reducing the time required for annotation and overcoming the limitations of existing databases and software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024061360_22052025_PF_FP_ABST
    Figure IB2024061360_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method (500) for automatically identifying molecules included in one or more complex mixtures (ES) by means of an electronic system (300). The system comprises: - a high--resolution mass spectrometry, HRMS, apparatus (301), connected to a high-pressure liquid chromatograph, HPLC, (302); such a liquid chromatograph is adapted to receive as input such one or more complex mixtures to measure a first retention time (Rt1) of each molecule; the mass spectrometry apparatus is configured to measure, for each molecule, a mass / charge ratio or first spectrum (MS1) and mass / charge ratios of molecular fragments generated during a fragmentation process of the molecule or second spectrum (MS2); - an electronic processing unit (303) operatively associated with the mass spectrometry apparatus to receive, as input, for each molecule, data (10) representative of the first and second spectra and the first retention time. The method comprises the steps of: - processing, by a first software module (20), such input data to group the different adducts of the same molecule in a single compound and to annotate each compound based on a comparison of such data with data of known molecules recorded in one or more databases associated with the first software module; the first software module generates a first annotation (25) of the structure of each molecule; - processing, by a second software module (40), such input data (10) to generate (42) a second annotation of the structure of each molecule; the second software module is configured to compare, based on machine learning techniques, the first spectrum of a molecule and the related associated second spectrum, with those "in silico H" of a database of known molecules associated with the second software module in which the second spectra were calculated; - processing, by a third software module (60), such input data to generate (51) a third annotation of the structure of each molecule, said third software module (60) being configured to compare, based on machine learning techniques, the second spectrum of the molecule with second spectra of known molecules present in a respective database associated with the third software module; - processing, by a fourth software module (30), the first and second spectra of at least a part of the compounds processed by said first software module to generate a prediction (31) of the molecular formula of the molecules; - calculating, by a fifth software module (80), a predicted retention time value : (Rt2) of a database of known molecules (DB); for each molecule of the one or more complex mixtures, the method includes comparing (100) : - the first annotation of the structure of the molecule, - the second annotation of the structure of the molecule, - the third annotation of the structure of the molecule, - the prediction of the molecular formula of the molecule, wherein the comparison step comprises a step of checking a consistency (103) of the annotations generated by the first, second, and third software modules based on : an evaluation of a first parameter, representative of the biological source of the molecule, predicted by the first, second, and third software modules, with respect to the information provided by the database of known molecules, a comparison of a second parameter, representative of the molecular formula of the molecule, predicted by the first, second, and third software modules with respect to the prediction provided by the fourth software module, an evaluation of a third parameter representative of the predicted retention time value with respect to the experimental one,to select (200) which of the aforesaid first, second, and third annotations of the structure of the molecule reliably identifies the molecule.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR IDENTIFYING MOLECULES IN A COMPLEX MIXTURE

[0002] AND RELATED SYSTEM

[0003] DESCRIPTION

[0004] TECHNOLOGICAL BACKGROUND OF THE INVENTION

[0005] Field of application

[0006] The present invention generally relates to systems for identifying molecules in a complex mixture. In particular, the invention relates to an innovative method for automatically carrying out processing on complex mixtures or extracts which can comprise a very large number of different molecules, in order to reliably identify molecules in such mixtures.

[0007] Prior art

[0008] As is known, molecules produced by living organisms, generally referred to as "natural products", are widely used to make drugs adapted to treat and prevent diseases in humans, animals, and plants. Most living organisms produce such natural products, and in most cases, a single organism can produce hundreds of different molecules.

[0009] Living organisms that are evolutionarily distant from one another tend to produce mutually different molecules (e.g., plants vs. bacteria, bacteria vs. fungi, etc.) and, within the same group (plants or bacteria, for example) , different organisms tend to produce different molecules. However, there are cases of identical molecules produced by distant organisms (e.g., geosmin) , therefore any separation of molecules according to the biological source of origin represents a probabilistic and non-absolute aspect. The entire chemical complexity present in an organism can be detected by treating the entire organism, or portions thereof, with appropriate methods (e.g., solvents, resins, etc.) to prepare extracts .

[0010] Therefore, the extracts, or more generally the complex mixtures, are mixtures of different molecules of unknown type and concentration. Knowing which molecules are present in an extract library would make the use thereof more efficient. However, this is not a trivial task .

[0011] In order to analyze the molecules present in a mixture, it is usual to employ an apparatus referred to as a mass spectrometer, in particular a High Resolution Mass Spectrometer (HRMS) associated with analytical techniques of high pressure liquid chromatography (HPLC) or ultra high pressure liquid chromatography (UHPLC) . In particular, the HRMS mass spectrometer associated with liquid chromatography, used as a detector apparatus, is associated with an HPLC (or UHPLC) which carries out a chromatographic separation of the molecules and provides a chromatogram in which, for each molecule, the mass / charge ratio (m / z) (referred to as the spectrum MSI or scan MSI) is measured, as well as the values of the mass / charge ratios of the molecular fragments generated during a subsequent fragmentation process of the molecule (also referred to as the spectrum MS2 or scan MS2) .

[0012] In more detail, during liquid chromatography or LC analysis, the molecules of the complex mixture are retained differently in the column and distributed based on the physicochemical properties thereof, so as to obtain a specific retention time (Rt) for each molecule. Once output from the chromatograph, the molecules of the mixture are sent as input to the mass spectrometer. Such an instrument measures one or more mass charge ratio m / z values for each molecule, which are recorded by the spectrometer as first spectrum MSI. Furthermore, by fragmenting the signals present in the spectrum MSI of the molecules, the spectrometer records a second spectrum MS2 which collects the signals corresponding to the fragments of the molecule.

[0013] Usually, those skilled in mass spectrometry are able to identify a molecule by observing the chromatographic behavior thereof , mass / charge ratio value ( s ) and the fragmentation profile and evaluating the correspondence of these values with those of known molecules listed in speci fic databases . The process thus described, termed as " annotation" by those skilled in the art , is of the "manual" type .

[0014] However, carrying out the manual annotation is only possible for a limited number of complex mixtures due to the high number of molecules present in such mixtures , which would require unacceptably long times i f the process were exclusively manual , and due to the vastness of the reference databases .

[0015] Therefore , today there is a need to achieve methods for automating annotation operations so as to process , with high productivity, a large number of extracts or mixtures of high molecular complexity .

[0016] To this end, computer products have been developed, executable by an electronic processing unit operatively associated with the mass spectrometer, which allow carrying out annotations on an extract automatically by processing the information of the first MS I and second MS2 spectra associated with the molecules generated by the high-resolution mass spectrometer (HRMS ) .

[0017] A first software , known with the trade name Compound Discoverer™ or CD by Thermo Scientific, is configured to process data from mass spectrometers known with the trade name Thermo Scientific Orbitrap™. Such mass spectrometers are configured to provide output data in a ".raw" file format.

[0018] Starting from the values of the first MSI and second MS2 spectra and from the information on the retention times Rt of the molecules in an extract, the first software CD works by first creating a mass "feature", i.e., a chromatographic peak, with m / z ratios defined and unique by MSI, MS2 and retention times Rt and then collecting the "features" deriving from different adducts of the same molecule in a single compound followed by the related quantification in the analyzed extracts. The annotation of each compound is carried out by comparing the relative mass and fragmentation profiles with internal databases associated with such a first software module CD (for example MZVault, CD-MassList, ChemSpider, mzCloud) and, at the same time, calculating a most likely molecular formula (raw formula or elemental composition) for the molecule to be identified which is provided as output.

[0019] In this respect, as known to those skilled in the art, it is worth noting that a mass spectrometer is capable of identifying molecules only if they acquire a positive (or negative) charge. The charge can be acquired by the molecule by associating with an ion, e.g., H+ proton, Na+ sodium, K+ potassium, NH4 + ammonium. Such a charged molecule is referred to as an adduct. Each individual molecule can associate with one or more ions and thus be identified, in terms of mass, in more than one charged form, i.e., as adducts of mutually different type. Each adduct will have a different mass, given by the mass of the molecule plus that of the charged ion.

[0020] The first software CD is capable of discriminating that different adducts actually represent different forms of the same molecule. The software CD is thus capable of providing a unique annotation for the set of adducts.

[0021] If on the one hand such first software CD allows analyzing several extracts at the same time, the limit of such software lies in the use of a closed-type database, mzCloud, the content of which cannot be increased by the user. The fact of not being able to adapt such a database to one's own target negatively affects the reliability of the annotation.

[0022] A second software, known as MolDiscovery or MD (Nature Communications 2021, doi . org / 10.1038 / s41467-021- 23986-0) , operates by comparing the first spectrum MSI of a molecule and the fragmentation profile thereof, or associated second spectrum MS2, with those (in silico) of a customizable database of molecules, in which the fragmentation spectra have been calculated, and not obtained from experimental fragmentation data.

[0023] In particular, the software MD consists of an algorithm, based on machine learning, capable of associating a fragmentation in silico (MS2) starting from a chemical structure. The public spectra MS2 library of the webbased mass spectrometry ecosystem GNPS (Global Natural Products Social Molecular Networking) containing thousands of annotated experimental spectra MS2 related to small molecules, is used to train the model.

[0024] The output provided by the second software MD is the best annotated molecule and the associated annotation score. The limitations of such a second known software MD is that it is not capable of recognizing different adducts originating from the same molecule and the use of theoretical fragmentations which can be very different from the real ones. Furthermore, it is impossible to replace the dataset to train the model.

[0025] Unlike the first software CD, the second software MD is not capable of discriminating that different adducts actually represent different forms of the same molecule. Therefore, MD can provide different annotations for different adducts assuming, erroneously, that such adducts are different molecules.

[0026] Furthermore, unlike the software CD, the second software MD is configured to process only files in ".mzML" format.

[0027] A third software, known as MS2Query or MQ (Nature Communications 2023, doi . org / 10.1038 / s41467-023-37446- 4) , is a machine-learning-based tool which works by comparing the spectrum MS2 of a molecule with those in a database of previously annotated molecules, for example the MQLibrary database.

[0028] Note that thus a MQLibrary database provides the third software MQ with a set of real second spectra MS2 to train it to recognize the spectra of previously unexamined molecules. The public spectra MS2 library of the web-based mass spectrometry ecosystem GNPS (Global Natural Products Social Molecular Networking) containing thousands of annotated experimental spectra MS2 related to small molecules, is used to train the model.

[0029] It is worth noting that, unlike the second software MD, the user can add or replace the annotated spectra MS2 and generate a new model.

[0030] Furthermore, using a machine learning process, such third software MQ is configured to refine the predictions once the training set has been enriched with additional molecules. As in the case of the software MD, the third software MQ is not capable of grouping several adducts originating from the same molecule . Furthermore , unlike the software CD, the third software MQ is configured to process only files in " . mzML" format . The output provided by such a third software MQ is the best annotated molecule and the associated annotation score .

[0031] Therefore , in light of the limitations of the known software methodologies described above , there is an increasing need to devise a methodology that allows carrying out the annotation of each detected molecule of a complex mixture or extract in a more accurate and reliable manner, in particular in the case of complex mixtures comprising a very high number of di f ferent molecules . SUMMARY OF THE INVENTION

[0032] Therefore , it is the general obj ect of the present invention to provide a method for automatically identi fying molecules in a complex mixture , of the type comprising a very large number of molecules which are di f ferent from one another, which allows such an identi fication to be carried out more accurately and reliably, obviating, at least partially, the limitations described with reference to the aforementioned known annotation methodologies . This obj ect is achieved by a method for automatically identifying molecules in a complex mixture, according to claim 1 .

[0033] In particular, it is a task of the invention to provide a method for identi fying molecules which allows increasing the reliability of the identification of the molecules contained in one or more complex mixtures based on a comparison of the results of the annotation operations carried out by each of the aforesaid first CD, second MD, and third MQ software, i . e . , of the predictions provided by such software on the identity of the molecules .

[0034] It is a further task of the invention to provide a method for identi fying molecules in a complex mixture which, using a speci fic implementation logic, allows carrying out the aforesaid comparison quickly in the case of mixtures which can comprise a very large number of mutually dif ferent molecules .

[0035] Preferred and advantageous embodiments of the aforesaid automatic identi fication method are the subj ect of the dependent claims .

[0036] The present invention also relates to an electronic system, according to claim 13 , in which such an electronic system is configured to implement the method of the present invention . BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Further features and advantages of the method for automatically identi fying molecules in a complex mixture according to the invention will be apparent from the following description of preferred embodiments , given by way of non-limiting indication, with reference to the accompanying drawings , in which :

[0038] Figure 1 shows , with a flow diagram, a general embodiment of the method for automatically identi fying molecules in a complex mixture of the invention;

[0039] Figure 2 shows , with a flow diagram, an example of the operating steps of a comparison and processing step of the method in Figure 1 carried out on the annotation results provided by at least four di f ferent software modules ;

[0040] Figure 3 illustrates , by way of example , a block diagram of an electronic system configured to implement the method of automatically identi fying molecules in a complex mixture in Figures 1-2 and 4 ;

[0041] Figure 4 illustrates , with a flow diagram, a particular example of a logic for selecting the annotations provided by the di f ferent software modules in the comparison and processing step in Figure 2 carried out by means of a decision tree .

[0042] Similar or equivalent elements in the aforesaid figures are indicated by the same reference numerals . DETAILED DESCRIPTION

[0043] With reference to Figure 3 , reference numeral 300 is used to indicate an electronic system adapted to implement the method 500 for automatically identi fying molecules in at least one complex mixture ES , in accordance with the present invention . Below, for simplicity of disclosure, the complex mixture ES will also be indicated by the term extract which, as known, is an example of complex mixture . However, note that the method of the present invention is advantageously also applicable to all other types of complex mixtures which are not extracts .

[0044] In particular, such an electronic system 300 comprises a high-resolution mass spectrometry, HRMS , apparatus 301 , connected to an instrument 302 for liquid chromatography analysis , i . e . , a high-pressure liquid chromatograph HPLC, or ultra high-pressure liquid chromatograph UHPLC, carried out on one or more extracts ES sent as input to the system 300 .

[0045] Note that , within the scope of the present invention, a mass spectrometer is said to be at high resolution when such a mass spectrometer has a mass accuracy which ensures an error less than 30 ppm, i . e . , the mass value obtained with such a spectrometer has a maximum error of 30 ppm; therefore , such a spectrometer allows the distinction of two masses which di f fer by 0 . 0006- 0 . 06 Da for molecules with a mass of 200 to 2000 Da, respectively . Preferably, a high-resolution mass spectrometer employable in the present invention has a mass accuracy which ensures an error below 5 ppm .

[0046] The mass spectrometer 301 of the system 300 is associated with the liquid chromatograph HPLC or UHPLC

[0047] 302 which carries out a chromatographic separation of the molecules of the at least one extract ES and provides a chromatogram in which, for each molecule , the mass / charge ratio (m / z ) or first spectrum MS I and the values of the mass / charge ratios of the molecular fragments generated during a subsequent fragmentation process or second spectrum MS2 are measured .

[0048] The electronic system 300 further comprises an electronic processing unit 303 operatively associated with the mass spectrometer 301 . The mass spectrometer 301 is configured to generate as output numerical values representative of the first MS I and second MS2 spectra of each molecule adapted to be sent as input to the aforesaid electronic processing unit 303 of the system 300 .

[0049] In an embodiment, the electronic processing unit

[0050] 303 comprises at least one processor 3031 , for example a microprocessor or microcontroller, and a memory block 3032, 3032' associated with the processor to store instructions. In particular, such a memory block 3032, 3032' is connected to the processor 3031 through a data line or communication bus 3033 (e.g., PCI) and consists of a service memory 3032, of the volatile type (e.g., SDRAM type) , and of a system memory 3032' of a nonvolatile type (e.g., SSD or eMMC type) .

[0051] Furthermore, for example, such an electronic processing unit 303 comprises input / output interface means 3034 connected to the at least one processor 3031 and to the memory block 3032, 3032' through the communication bus 3033 to allow an operator close to the electronic system 300 to interact directly with the processing unit 303.

[0052] In an embodiment, the electronic processing unit 303 is associated with a wired or wireless data communication interface 3035 configured to connect such a processing unit 303 to a data communication network 3036, e.g., the Internet, to allow an operator to remotely interact with the processing unit 303 or to receive updates.

[0053] For example, such a data communication interface 3035 is embodied in a wired data communication block, e.g., of the Ethernet type, or in a radio frequency (RF) block, operating according to the Bluetooth Low Energy, BLE, Wi- Fi standard or according to the 3G, 4G, 5G mobile radio communication standards or in LoRa (Long Range ) and LoRaWAN technology .

[0054] With reference to Figures 1-2 and 4 , the operating steps of the method 500 for automatically identi fying molecules in at least one extract ES , according to the present invention, implemented through the system 300 , are described in greater detail below .

[0055] In an embodiment, the electronic processing unit 303 is arranged to execute the codes of an application program implementing the method 500 of the present invention .

[0056] In a particular embodiment , the microprocessor 3031 of the electronic processing unit 303 is configured to load and execute the application program codes implementing the method of the present invention .

[0057] All the flow diagrams in Figures 1 , 2 and 4 used to describe the method 500 of the invention include a symbolic start step STR and a symbolic end step ED .

[0058] Note that the method 500 of the invention includes executing a first software module 20 , for example the known software module Compound Discoverer™ or CD mentioned above , configured to process data 10 from the mass spectrometer 301 . Such data 10 comprise , for example , first MS I and second MS2 spectra values and information on the retention times of the molecules in an examined extract ES, i.e., the experimental retention times Rt 1.

[0059] In particular, starting from the aforesaid first MSI and second MS2 spectra values and from the information on the experimental retention times Rtl of the molecules, the first software module 20 is configured, firstly, to create a mass "feature", i.e., a chromatographic peak, with m / z ratios defined and unique from MSI, MS2 and retention times Rtl. Next, CD works to collect the "features" deriving from different adducts of the same molecule in a single compound, then proceeding with the relative quantification in the analyzed extracts. The annotation of each compound is carried out by comparing the relative mass and fragmentation profiles with internal databases associated with such a first software module CD (for example MZVault, CD-MassList, ChemSpider, mzCloud) and, at the same time, calculating a most likely molecular formula for the molecule to be identified which is provided as output.

[0060] Note that the values of the experimental retention times Rtl are not evaluated for the annotation, but represent metadata accompanying the compound.

[0061] In more detail, for each molecule contained in the extract ES , such a first software module 20 is configured to generate as output 25 the annotation considered most reliable among those proposed by the databases to which it has access .

[0062] Furthermore , the method 500 of the invention includes executing a second software module 40 , for example the software module MolDiscovery or MD mentioned above , configured to process , for each molecule , the output data 10 from the mass spectrometer 301 .

[0063] In particular, the second software module 40 is configured to compare the first spectrum MS I of a molecule and the related fragmentation profile , or associated second spectrum MS2 , with those " in silico" of a customi zable database of molecules , for example the MDLibrary database , in which the fragmentation spectra have been calculated, and not obtained from experimental fragmentation data .

[0064] In particular, as mentioned above , the software MD consists of an algorithm, based on machine learning, capable of associating a fragmentation in silico (MS2 ) starting from a chemical structure . In other words , MD is capable of predicting the spectrum MS2 from a chemical structure based on a probabilistic model elaborated starting from annotated experimental data ( GNPS ) . The output 42 provided by the second software module 40 is , for each compound, usually a list of best-annotated molecules and the associated annotation score .

[0065] Furthermore , the method 500 of the invention includes executing a third software module 60 , for example the software module known as MS2Query or MQ mentioned above , configured to process , for each molecule , the output data 10 from the mass spectrometer 301 . In particular, the third software module 60 is configured to recogni ze the second spectrum MS2 of a molecule based on the model created based on previously annotated fragmentation learning techniques , for example the MQLibrary database , based on machine-learning techniques . As mentioned above , such an MQLibrary database provides the third software module 60 with a set of real second spectra MS2 to train it to recogni ze the spectra of previously unexamined molecules . The output 51 provided by such a third software module 60 is , for each compound, usually a representative list of the best annotated molecule and the associated annotation score . In addition, the method 500 of the invention includes executing a fourth software module 30 , for example a software known with the trade name Sirius or SI . In particular, such a fourth software module is used for the calculation of molecular formulas from high- resolution mass spectrometry data . In other words , the fourth software module 30 is capable of predicting a molecular formula from the first MS I and second MS2 spectra values .

[0066] Advantageously, the method 500 of the invention includes using a main database DB (NAICONS Mass List ) comprising data of about 170 , 000 chemical structures of known molecules ( for a total of about 439, 000 entries ) . In particular, through such a main database DB, the method of the invention 500 allows verifying whether the mass values of the molecules annotated by the first 20 , second 40 , and third 60 software modules are already associated with some known molecule to retrieve further information : biological source and predicted retention time Rt2 .

[0067] Furthermore , the method 500 of the invention includes executing a further software module or fifth software module 80 . Such a fi fth software module 80 , referred to as "NAI-Rt2" , is configured to predict the retention time Rt for molecules in mass spectrometry analysis with liquid chromatography . In particular, such a fi fth software module 80 performs the calculation of a retention time Rt2 on the main database of known molecules DB mentioned above to detect potentially incorrect annotations provided by any of the aforementioned first 20 , second 40 , and third 60 software modules .

[0068] Such a fi fth software module 80 was developed by the Applicant to carry out functions similar to those of the known Retip software , but was advantageously rewritten in " Python" language and optimi zed by replacing the XGBoost algorithm with an updated version of "Histogrambased Gradient Boosting Regression Tree (part of scikit : https : / / scikit- learn . org / stable / modules / generated / sklearn . ensemble . Hist GradientBoostingRegressor . html# ski earn- ens emble- histgradientboostingregressor ) .

[0069] A preferred embodiment of the method 500 of the invention will be described in more detail with reference to Figure 1 .

[0070] In an initial step, the method 500 of the invention includes analyzing each extract ES using the aforementioned liquid chromatography 302 coupled with high-resolution mass spectrometer (HRMS ) 301 . During the chromatograph analysis 302 , the molecules of the extract or mixture ES are retained di fferently in the column and distributed in a nineteen-minute cycle based on the physicochemical properties of the molecules . The molecules are thus assigned a specific experimental retention time or first retention time , indicated by reference sign Rtl below .

[0071] The output data 10 from the mass spectrometer 301 , i . e . , the values of the first spectrum MSI , second spectrum MS2 and the information on the measured retention times Rtl of all the molecules of the examined extract ES are grouped into a file in a proprietary format of the software manufacturer to be provided as input to the aforementioned first software module 20 , which is configured to process them .

[0072] In an embodiment, the input data 10 to the first software module 20 further comprise libraries of second spectra MS2 of known molecules at dif ferent annotation levels , for example the aforementioned library mzVault .

[0073] Following the processing of the data 10 received as input, the first software module 20 is adapted to generate a respective database 21 , or " . cdresult" file, containing the results produced by processing the set of data files in proprietary format and, in addition, information on the analysis settings used to process such data .

[0074] The first software module 20 is capable of grouping several adducts of the same molecule , identifying them with the same "compound" . For each compound (marked with a unique ID number ) the first software module 20 is configured to associate the most likely molecular formula and to compare the second spectra MS2 of the compounds with the libraries mzVault of previously recorded and recognized molecules .

[0075] I f one or more of the results present in the database 21 indicates that one or more molecules of the examined extract ES corresponds to an already known molecule contained in the library mzVault, then the annotation carried out by the first software module 20 is successful and the method 500 returns related correspondence information 22 . Such an annotation further includes a subsequent filtering step to select only the compounds which in addition to being contained in the library mzVault also have values above threshold values in terms of peak area .

[0076] The method 500 includes a step of filtering 23 the " . cdresults" file to select the compounds to be used in the annotation process , eliminating those compounds which : have been recogni zed by the library mzVault ; have not associated at least one spectrum MS2 , thus they only have spectra MSI ; have a first retention time value Rtl dif ferent from a speci fic range ; have signals shared with "white" extracts , i . e . , containing only solvent ; have properties outside of predetermined threshold values in terms of peak area and in relation to peak shape .

[0077] In the present invention, for example, the thresholds are set to a minimum area of 107and a minimum peak value of 8 and a Rtl between 0 . 5 and 14 minutes .

[0078] Following the filtering 23 , the selected compounds are grouped 24 in a "compounds . j son" file .

[0079] With reference to such grouped compounds 24 , the method of the invention 500 includes performing further operating steps concomitant therewith . One of such operating steps comprises performing, by the first software module 20 , a " compounds2summary" step 25 which includes extracting the information of the grouped compounds 24 .

[0080] In particular, the software 20 provides an annotation by selecting it from among the possible proposals from the databases to which it has access , represented by an InChlKey datum value of the annotated molecule and a Smiles ( Simpli fied Molecular Input Line Entry System) datum, adapted to describe the molecular structure by means of an ASCI I string . The software completes the missing information by obtaining it from the main database DB in relation to the biological source of the annotated molecule ( source ) , at the predicted retention time Rt2 , to provide complete information on which to carry out the subsequent comparison operation 100 .

[0081] Note that , for each compound, the first software module 20 : can assign a molecular formula ; or can report a correspondence with the library mzVault and then produce an annotation; or can report only partial information, for example only a molecular formula without assigning a name or a chemical structure .

[0082] Still with reference to the grouped compounds 24 , the method of the invention 500 includes performing 26, 30 , 31 a step of assigning a molecular formula to the grouped compounds 24 .

[0083] In relation to the operating step of assigning a molecular formula, the method 500 of the invention includes processing the grouped compounds 24 by the fourth software module 30 . Before performing such a processing, the method 500 includes using the output of the "compounds . j son" file 24 and extracting the related MS I and MS2 and saving them by means of the module 26 in " . mgf" format , which can be processed by the aforesaid fourth software module 30 . In particular, a command ( e . g . , the run sirius command) allows executing the fourth software module 30 on the input module 26 in " . mgf" format to generate a prediction of the molecular formula of molecules having a mass less than 850 Daltons . The result of such a processing, converted into a " . tsv" file using a " sirius2summary function 31 , consists of filtering the results obtained to select only the best molecular formulas predicted for a relative pair of signals MS I and MS2 . Following the aforesaid filtering 31 , the molecular formula calculated by the fourth software module 30 is provided as input datum for the subsequent comparison and processing operation 100 .

[0084] In addition to the calculation of the molecular formula, the method of the invention 500 also includes calculating a respective predicted retention time, indicated by reference sign Rt2 , of each molecule by the fi fth software module 80 , e . g . , the software "NAI-Rt2" mentioned above .

[0085] In particular, such fi fth software 80 uses a machine learning algorithm, which is trained and tested, for example using 385 manually annotated molecules . The method 500 then includes that the molecules of the entire DB ( 170 , 000 molecules ) are processed to calculate the predicted retention time Rt2 . At the end of such a processing, the results are added to the DB and are provided as input data in the comparison step 100 . Furthermore, in an embodiment, it is possible that, for each molecule present in the DB, information on the source of isolation is added. In an embodiment, the complex mixtures come from actinomycetes and given the homogeneous biological origin thereof, it is possible to insert further ranking criteria for the annotations (Rt2, molecular formula) . Such information becomes available as input datum in the comparison step 100.

[0086] The method 500 of the invention further includes a step 11 in which the output data 10 from the mass spectrometer 301, in proprietary format, is converted into a " .mzML" file, to be sent as input to both the second 40 and third 60 software modules. For example, the conversion software used for this purpose is the software: "ThermoRawFileReader " released by Thermo and used by means of a docker wrapper (http : / / compomics . github . io / pro j ects / The rmoRawFile Parser ) made by Compomics (https: / / www.compomics.com / ) .

[0087] In more detail, the second 40 and third 60 software modules are configured to perform parallel annotations on the ".mzML" file. The processing performed by the second software module 40 generates as output a single ".tsv" file for a group of samples / extracts (e.g., 80 extracts) . The processing performed by the third software module 60 generates as output a ".tsv" file for each sample. Note that in the " . tsv" files , each line shows the annotation related to an "Extract-second spectrum MS2- scan" pair . Multiple spectra MS2 and thus multiple annotations can be present for each molecule .

[0088] In the case of the second software module 40 , the annotations are not performed for a few second spectra MS2 when the corresponding parent ion has an experimental neutral mass which is not present in the library of such a software module 40 .

[0089] In relation to the second software module 40 , the method 500 then includes a step of filtering 41 the data contained in the generated " . tsv" file . In particular, compounds are selected with annotation score having a greater reliability than a predetermined threshold, e . g . , score >=50 .

[0090] A step of grouping 42 such data using the "moldiscovery2 summary" functionality is also included .

[0091] In particular, the software 40 provides one or more predictions on the structure of the annotated molecule, represented by the Smiles value and name of the annotated molecule and score ; then the module 42 completes the missing information by deriving it from the main database DB in relation to the biological source of the annotated molecule ( source ) , at the predicted retention time Rt2 , while the molecular formula and InChlKey are completed by the Soupy program by means of the open-source tool RDkit starting from the Smiles value; therefore, the complete information is available on which to carry out the subsequent comparison operation 100.

[0092] In relation to the third software module 60, the method 500 includes a step of filtering the data contained in the generated ".tsv" file.

[0093] In particular, compounds with annotation scores having a reliability greater than a certain threshold, e.g., score >=0.3, are selected, proceeding with the selection only of compounds which have a difference between experimental MSI (m / z) and MSI present in the library (m / z) not exceeding 2 ppm.

[0094] Also in the ".tsv" files generated as output by such a third software module 60, each line corresponds to an annotation related to an "Extract-second spectrum MS2- scan" pair.

[0095] Also for this software, multiple spectra MS2 and thus multiple annotations can be present for each molecule.

[0096] Since the third software module 60 is configured to process all second spectra MS2 of the molecules, independently of the associated parent ion, all the scans MS2 return an annotation.

[0097] Since the third software module 60, during processing, is adapted to rename the scan numbers of the files in proprietary format , the method 500 of the invention includes an additional functional module 50 "mzml2 summary" configured to extract the results of the processing performed by the third software module 60 with the file in proprietary format .

[0098] At this point, the data of the " . tsv" file obtained with the filtering and output of the additional functional module 50 "mzml2summary" are j oined through the functionality 51 which aligns the results of the third software module 60 to the extract-scan number pairs . Such a functionality 51 returns a single " . tsv" file , "ms2query . tsv" , per group of samples / extracts to be sent to the next comparison and processing step 100 . Such a single file "ms2query . tsv" also comprises further information .

[0099] In particular, the software module 60 provides one or more predictions on the structure of the annotated molecule , represented by the value of the InChlKey datum, the name of the annotated molecule or by the Smiles datum, and the functionality 51 completes the missing information, obtaining it from the main database DB in relation to the biological source of the annotated molecule ( source ) , at the predicted retention time Rt2 , while the molecular formula is completed starting from the Smiles datum by means of the open-source tool RDkit ; if only InChlKey is available , the Smiles datum is obtained by querying PubChem; therefore, the information on which to carry out the next comparison operation 100 is provided .

[0100] Compounds having an identification code CompoundID for which at least one of the software modules 20 , 40 , 60 has generated an annotation are provided to the next comparison and processing step 100 of the method . In contrast , the compounds which none of the aforesaid software modules 20 , 40 , 60 managed to annotate are classi fied by the method 500 as "unknown" .

[0101] Note that the results obtained from the processing performed by the first 20 and fourth 30 software modules are expressed in terms of compounds . In particular, only one molecular formula per compound is obtained from the fourth software module 30 . Instead, the annotations from the second 40 and third 60 software modules refer to the " extract-second spectrum MS2-scan" pair .

[0102] Therefore , in order to compare the results of the four annotations carried out, the aforesaid comparison and processing step 100 of the method 500 of the invention includes performing a step of aligning 101 the outputs provided by each software module 20 , 30 , 40 , 60 . Such an alignment 101 occurs using the identification code CompoundID, Extract and MS2 Scan and is performed only on the filtered signals.

[0103] Furthermore, since the second 40 and third 60 software modules are configured to provide more than one annotation for each compound, the aforesaid comparison and processing step 100 of the method 500 of the invention includes carrying out a step of compressing the results 102 or "squashing" (i.e., reducing a group of information to a single piece of information which represents them) to choose the most reliable annotation to be inserted into a decision tree during the comparison step 100.

[0104] With reference to Figure 2, the comparison and processing step 100 of the method 500 for automatically identifying molecules of the invention is described below in greater detail. As mentioned, such a comparison and processing step 100 is carried out on the annotation results provided by the four different analysis software modules 20, 30, 40, 60 described above.

[0105] In particular, such annotation results, in ".tsv" format, are provided by the functional modules 42, 51, 31, and 25. The comparison and processing module 100 receives as input all the summaries mentioned above and returns the most reliable annotation, based on a decision tree, as will be detailed below .

[0106] In the alignment step 101 , it is possible to align the summaries provided by the functional modules 42 , 51 , 31 , and 25 . In fact, the second 40 and third 60 software modules produce a set of values for each Compound ID, while the first 20 software module produces an annotation per compound ( set of scans ) and the fourth 30 software module produces a plurality of molecular formulas per pair MSI , MS2 . Note that of the molecular formulas produced by the fourth software module 30 , the method of the invention includes selecting only one molecular formula .

[0107] Since the first software module 20 generates an annotation for each compound, it is possible to use the information returned by such first software 20 to align the results provided by all the other software .

[0108] Such an alignment includes using a multi-level index system (Multilndex ) in which the first level comprises the Compound IDs , the second level the extracts and the third level the scan numbers .

[0109] In particular, such an alignment step 101 of the method of the invention includes grouping, for each Compound ID, all the extracts in which such an ID has been detected and for each of such extracts , all the scans in which the

[0110] ID was observed . Furthermore , so as to simplify the comparison of the outputs provided by the aforesaid four software modules 20 , 30 , 40 , 60 , the comparison step 100 of the method 500 includes carrying out a squashing step 102 to reduce the outputs of the second 40 and third 60 software modules to a single set of annotation values for each compound .

[0111] In an embodiment, such a " squashing" step 102 includes a step of selecting the best annotation / scan related to the second 40 and third 60 software modules , as detailed in the example reported below .

[0112] In particular, once the output values of the aforesaid software 40 , 60 have been calculated, the

[0113] " squashing" step 102 includes considering four characteri zing parameters for each annotation / scan, i . e . :

[0114] • in the case of the second software module 40 (MD) : A_MD_ScoreBinned, A_MD_SiriusRank, A_MD_InChIKeyFreq and CD_Scan;

[0115] • in the case of the third software module 60 (MQ) : A_MQ_ScoreBinned, A_MQ_SiriusRank, A_MQ_InChIKeyFreq and CD_Scan .

[0116] Such parameters are used to sort the annotations / scans lexicographically . In other words , at first, the annotations / scans are sorted with respect to the first parameter ( ScoreBinned) . I f two or more annotations / scans have the same first parameter, such annotations / scans are sorted with respect to the second parameter . I f two or more annotations / scans have the same second parameter, such annotations / scans are sorted with respect to the third parameter . The method includes proceeding in a similar manner up to the fourth parameter .

[0117] Note that the column names of the parameters selected in this step comprise the prefix A, followed by the name of the relevant software .

[0118] Furthermore, the grouping of the values of such parameters is carried out so that the first three parameters , ScoreBinned, SiriusRank, InChlKeyFreq take decreasing values , while the fourth parameter CD_Scan takes increasing values . The first set of values for each compound is thus maintained .

[0119] Note that with the compression or " squashing" step 102 , the output information of the second 40 and third 60 software modules is reduced to a single annotation for each individual index identi fied as CD_ID .

[0120] For example , with reference to Table A and Table B reported below, the following is observed in relation to the third software module 60 (MQ) .

[0121] In the case of the same compound identi fier, CD_ID, which is assigned by the first software module 20 , equal to 255 , the third software module 60 reports di f ferent results when the extract, CD_extract, and the scan number, CD_scan, vary.

[0122] TABLE A

[0123] TABLE B

[0124] As indicated in Tables A-B the results of the annotation on the different extracts are ordered by the "squashing" function based on the evaluation of the following criteria which are, in order of relevance:

[0125] 1) a score provided by the third software module 60 = Binned Score;

[0126] 2) a correspondence of the molecule of the extract with the molecular formula predicted by the fourth software module 30 = Sirius Rank;

[0127] 3) a frequency of InchiKey of each extract;

[0128] 4) the scan number = CD_Scan. The aforesaid criteria are evaluated in sequence and the switching from one to the next occurs in the case of equal values associated with the same criterion.

[0129] With reference to the method of assigning a value to the above criteria, the following is noted.

[0130] In relation to the step of scoring the Binned Score or binning step, for the third software module 60 (MQ) , which has a score range of [0.3 - 1] it is possible to divide the score into twenty uniform bins. For the second software module 40 (MD) , which has a score range of [50 - 220] it is possible to divide the score into three bins for scores below 100 and seventeen bins for scores above 100.

[0131] In relation to the Sirius Rank scoring step, the molecular formula calculated by the third software module 60 (MQ) (or by the second software 40 - MD) is compared with that predicted by the fourth software module 30 and a selectable rank is assigned as {-1,0,1} . More specifically:

[0132] • -1 if both the third software module 60 (MQ) (or the second software 40 - MD) and the fourth software module 30 predict a molecular formula, but the formulas are different;

[0133] • 0 if the third software module 60 (MQ) (or the second software 40 - MD) or the fourth software module 30 was not capable of predicting a molecular formula ;

[0134] • 1 i f both the third software module 60 (MQ) ( or the second software 40 - MD) and the fourth software module 30 predict the same molecular formula .

[0135] In relation to the assignment step of the Frequency of the InChlKey predicted for each CompoundID, when the annotations / scans originated for the same CompoundID result in the same annotation, this is weighed more with respect to when the annotations are divergent . This is carried out by counting the frequency of the InchlKey predicted for each compound ID, ignoring any missing predictions .

[0136] In relation to the CD_Scan scoring step, it is possible to select the annotation carried out on the lower scan number .

[0137] In particular, with reference to Tables A and B, the first two sorting criteria generate " equal merit" since six predictions have the same Binned Score value (between 0 . 895 and 0 . 93 in the Tables ) and an agreement (value 0 in the Tables ) is never obtained with the prediction provided by the fourth software module 30 . Furthermore , the first two predictions also show the same frequency of InChlKey (value 3 in the Tables ) : i . e . , three times the third software module 60 (MQ) predicted the same molecule within the compound 255 . Therefore, the CD_Scan scan number is used to select the most representative prediction and the extract with the lowest scan number, i.e., CD_scan = 184, is selected. Such an extract, corresponding to the first line of Table A, is that returned by the "Squashing" operation

[0138] 102 (Table B) .

[0139] After the aforesaid "Squashing" operation 102, the comparison and processing step 100 of the method 500 of the invention includes adding further information by combining the data from the different software modules 20, 40, 60.

[0140] Advantageously, a step of checking the consistency

[0141] 103 of the annotations (predictions) produced by the first 20, second 40, and third 60 software modules is included. In such a consistency check step 103, each of the annotation results provided by the software modules 20, 40, 60 is assigned a score using as discriminating parameters: a first parameter, representative of the biological source, a second parameter, representative of the molecular formula, predicted by each software 20, 40, 60 and compared with the prediction provided by the fourth software module 30, and the predicted retention time value Rt2.

[0142] In other words, for each annotation, each of such parameters can take a tuple of three values {-1,0,1} . In an embodiment, with reference to the "biological source" parameter: the consistency of the biological source is 1 if the biological source field in the main database DB of known molecules takes the value "True"; the consistency of the biological source is 0 if no molecular structure InChlKey has been predicted by one of said software modules 20, 40, 60 or if the predicted molecular structure InChlKey is not present in the main database DB of known molecules; the consistency of the biological source is -1 if the biological source field in the main database DB takes the value "False".

[0143] In an embodiment, with reference to the "molecular formula" parameter, based on the comparison of each of the software modules 20, 40, 60 with the fourth software module 30: the consistency of the molecular formula is 0 if the software module 20, 40, 60 being compared does not include a molecular formula; the consistency of the molecular formula is 1 if the two predicted molecular formulas are equal; the consistency of the molecular formula is -1 if the two predicted molecular formulas are different.

[0144] In an embodiment, with reference to the parameter "predicted retention time value Rt2": the consistency of the predicted retention time is 1 if the value of the predicted retention time Rt2 deviates from the value of the first retention time (experimental retention time) Rtl by less than 1.4 minutes ; the consistency of the predicted retention time is -1 if the value of the predicted retention time Rt2 deviates from the value of the first retention time Rtl by more than 1.4 minutes; the consistency of the retention time is 0 if no chemical structure InChlKey has been predicted by one of said software modules 20, 40, 60 or if the predicted InChiKey is not present in the main database DB of known molecules.

[0145] Note that if no InChlKey value is predicted for any of the software 20, 40, 60, the consistency value is (0,0,0) .

[0146] In a particular embodiment of the step of checking the consistency 103 of the annotations, the method includes a step of grouping 103' the three-value tuples representative of the consistency of the software modules 20, 40, 60 into classes (not mutually exclusive) . Such a grouping step 103' is based on the following logic: a class T groups the tuples having at least two 1 (in any position) ; a class A is the subclass of class T the tuples of which have a 1 in the first position (i.e., the consistency of the source is 1) and at least one other 1; a class 0 contains tuples which only have a 1 in the first position, thus they are only consistent with the biological source parameter ("Source") .

[0147] Such classes are used in a subsequent step of selecting the annotations 200 of the method which, by means of a decision tree, allows determining which of the annotations generated by the software modules 20, 40, 60 is the most correct to identify a molecule of the extract ES .

[0148] Such a selection step 200 is described in detail based on the decision tree shown in Figure 4.

[0149] With reference to the selection step 200 of the method in question, it is possible to preliminarily distinguish whether one or more of the aforesaid software modules 20, 40, 60 are capable of actually predicting a molecule, thus a molecular structure value InChlKey (not a "not a number" or nan) can be attributed to each of such software, from whether no annotation is made, i.e., it is not possible to attribute a value InChlKey to one or more of the aforesaid software modules .

[0150] In the following example, it will be assumed, in a first case 201, that a molecular structure value InChlKey is attributable to three or two of the aforesaid software modules 20, 40, 60. In a second case 202, it will be assumed that a structure value InChlKey is attributable only to one of the aforesaid software, while the others have not provided any predictions.

[0151] In a step 203, the predicted values InChlKey for each software are compared in pairs. The result of such a comparison 204 can include two or three identical values InChlKey.

[0152] In this case, the name of the annotation returned by the process, i.e., the molecule considered as the winner of the process, is selected 205 based on a ranking of the names of the annotations provided by the software modules 20, 40, 60 used in the method. For example, such a ranking of names, from the most preferred to the least preferred, includes:

[0153] I) annotation name provided by the second software module 40 (MD) ,

[0154] II) annotation name provided by the first software module 20 (CD) ,

[0155] III) annotation name provided by the third software module 60 (MQ) .

[0156] Since the third software module 60 (MQ) generally returns the annotation names having greater lexical complexity with respect to the names provided by the other software 20 , 40 , such a software 60 was chosen as the software with the lowest ranking in step 205 . The annotation name provided by the third software module 60 is thus not selected in the case of two or three identical InChlKey .

[0157] The annotation resulting from such a selection 205 is defined in Class 1 .

[0158] I f the comparison 206 of values InChlKey results in no identity, the method includes a step of comparing 207 the names of the compounds predicted by the software 20 , 40 , 60 involved . In particular, the names predicted by the software modules 20 , 40 , 60 are compared in pairs .

[0159] I f the similarity of the names provided by two software is greater than a predetermined threshold, for example 75% , both software compared to each other are added to a list . After all the comparisons , this list can comprise two or all three of the software mentioned above . It should be noted that since the similarity of the names is not transitive , i f the list contains the names of all three software modules 20 , 40 , 60 , it does not necessarily mean that all the names provided are similar in pairs : there can still be two pairs of similar names and a pair of non-similar names .

[0160] In this case , the name of the annotation considered as the winner of the process is selected 208 based on a ranking of similarity of the names of the annotations provided by the software modules 20 , 40 , 60 used in the method . For example , such a ranking of names , from the most preferred to the least preferred, includes :

[0161] I ) annotation name provided by the second software module 40 (MD) ,

[0162] I I ) annotation name provided by the first software module 20 ( CD) ,

[0163] I I I ) annotation name provided by the third software module 60 (MQ) .

[0164] For the above reasons , the third software module 60 (MQ) was chosen as the software with the lowest ranking in step 208 and will never be selected based on name similarity .

[0165] The annotation resulting from such a selection 208 is defined in Class 2 .

[0166] Instead, i f the similarity of the names provided by two software is lower than said 75% threshold, the method includes a step of evaluating 209 the consistency of the annotations provided by the di f ferent software modules 20, 40, 60 based on a ranking determined by a comparison of the aforementioned parameters, indicated from the most relevant to the least relevant, i.e. : first parameter (biological source) , second parameter (molecular formula) , third parameter (predicted retention time value Rt2) .

[0167] In this second case, the name of the annotation considered as the winner of the process is selected as follows .

[0168] Firstly, annotations are only considered when they are in consistency class T or 0.

[0169] In the case of at least 2 consistencies 210 and only one software having two consistencies, the name of the annotation provided by such software is returned 211 by the process and the annotation is defined in Class 3a .

[0170] If there are 2 (or more) consistencies 210 and the best two are even, such a tie is resolved 212 based on an order of consistency based on a classification of the software modules 20, 40, 60 including, for example, from the most preferred to the least preferred:

[0171] I) third software module 60 (MQ) ,

[0172] II) second software module 40 (MD) ,

[0173] III) first software module 20 (CD) .

[0174] The name of the annotation resulting from such a tie 212' is returned by the process and the annotation is defined in Class 3b.

[0175] If there is a single consistency 213 or no consistency, the software best positioned in terms of the first parameter, representative of the biological source (BS) , is selected 214. The name of the annotation provided by such software is returned by the process 215 and the annotation is defined in Class 4a.

[0176] In the case of tie also with reference to the first biological source parameter (BS) , such a tie is resolved 216 based on an order of consistency based on a classification of the software modules 20, 40, 60 including, for example, from the most preferred to the least preferred:

[0177] I) third software module 60 (MQ) ,

[0178] II) second software module 40 (MD) ,

[0179] III) first software module 20 (CD) .

[0180] The name of the annotation resulting from such a tie 216' is returned by the process and the annotation is defined in Class 4b.

[0181] In the other cases 217, the resulting annotation 218 of the process, also defined in Class 5a, remains undecided .

[0182] With reference to case 202 in which a molecular structure value InChlKey is attributable to only one of the aforesaid software, while the others have not given any predictions, the selection step 200 of the method includes establishing 219 two types of winners depending on whether the annotation has a consistency class A or 0.

[0183] In the case of a software module in class A, such software is selected 220 if it includes a tuple with a 1 in the first position (i.e., the consistency of the biological source BS is 1) and at least one other 1. The name of the annotation 220' resulting from such software is returned by the process and the annotation is defined in Class 3c.

[0184] Instead, in the case of a software in class 0, such software is selected 221 if it includes a tuple with only a 1 in the first position (i.e., the consistency of the biological source BS is 1) . The name of the annotation 221' resulting from such software is returned by the process and the annotation is defined in Class 4c.

[0185] In the other cases 222, the annotation resulting from the process remains undecided 222' , also defined in Class 5b.

[0186] In an embodiment (not shown in the figures) , the method 500 of the invention includes a final step of calculating the exact neutral mass resulting from the selected annotation and the experimental neutral mass obtained from the first spectrum MSI of the selected Compound_ID and the related adduct type. After the comparison step 100 described above, if the difference between such two values, measured in ppm, is greater than two, the value of the annotation selected in the corresponding step 200 is excluded from the final list.

[0187] With the method 500 described above, the following information is provided for each compound contained in an extract examined: the software module 20 (CD) , 40 (MD) , 60 (MQ) the annotation of which was selected as final; the annotation class, which shows the confidence of the chosen annotation; a set of warnings which highlight discordant annotations or consistencies.

[0188] In an embodiment (not shown in the figures) , once the different annotation classes have been established, the method of the invention includes creating a ".json" file using the IT tool "catalogue_export " . This file contains all the information related to the annotations including compoundID (which in the ".cdresult" file is associated with the first MSI and second MS2 spectra, retention time, extracts in which it was detected) , the annotated molecule (with molecular formula, SMILES, InChlKey, biological source and predicted retention time Rt2) and confidence class of the annotation. Such a ".json" file also contains information about the "unknown" molecules, such as compound ID (and associated metadata) and molecular formula calculated by the fourth software module, if present.

[0189] Furthermore, the file reports the identifying IDs of the compounds which have matched with the libraries used in the processing by the first software module 20.

[0190] The contents of such a ".json" file together with the ".cdresults" file are used as input information for a further software tool, referred to as "archive_import " , adapted to import all the new annotations and attributes into an archive .

[0191] Once uploaded to the archive, an updated version of the fragmentation libraries related to the different annotation classes (including the "unknown" class) can be exported from the archive using the tool vaults_export . These libraries will be added to the next work of the first software module CD 20 which processes a new data set of extracts / samples .

[0192] The method of the invention for automatically identifying molecules in an extract has multiple advantages and achieves the intended objects.

[0193] It is the main advantage of the present invention to provide a method 500 and a related system 300 capable of analyzing new extracts or complex molecules and predicting the various structures of the molecules composing them alone . As the annotated data available to the system in the various libraries increases , both the ability to annotate signals of a complex mixture and the reliability of automatic annotation predictions will increase , by improving the models .

[0194] Furthermore , the Applicant highlights that a possible advantageous application of the method for automatically detecting and identi fying molecules in one or more extracts of the invention is to allow creating a web-based searchable database containing the molecules automatically annotated according to the method of the invention .

[0195] In particular, the invention also relates to an electronically searchable , web-based database adapted to contain data representative of a plurality of molecules automatically identi fied from one or more complex mixtures (ES ) by means of the method of the invention .

[0196] The use of such a database allows those skilled in mass spectrometry to drastically reduce the time necessary for the annotation process , avoiding iterative operations , such as the evaluation of the consistency of the chromatographic behavior, of the m / z value , i . e . , the first spectrum MS I , and of the fragmentation profile , the second spectrum MS2 with those listed in speci fic databases , focusing only on what those skilled consider most useful or that really needs to be manually carried out .

[0197] At the same time , providing such a database of automatically annotated molecules allows users not skilled in mass spectrometry to easily navigate between molecules , exploring the output results . Those skilled in the art may make changes and adaptations to the embodiments of the method and system described above or can replace elements with others which are functionally equivalent in order to meet contingent needs without departing from the scope of the following claims . Each of the features described as belonging to a possible embodiment can be made irrespective of the other embodiments described .

[0198] -k 'k 'k

Claims

CLAIMS1. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) by means of an electronic system (300) comprising:- a high-resolution mass spectrometry, HRMS, apparatus (301) connected to a high or ultra-high pressure liquid chromatograph, HPLC or UHPLC, (302) , said liquid chromatograph (302) being adapted to receive said one or more complex mixtures (ES) to measure a first retention time (Rtl) of each molecule, said high-resolution mass spectrometry, HRMS, apparatus (301) being configured to measure, for each molecule of said complex mixture, a mass / charge ratio or first spectrum (MSI) and mass / charge ratios of molecular fragments generated during a process of fragmentation of the molecule or second spectrum (MS2) ; an electronic processing unit (303) operatively associated with said mass spectrometry apparatus (301) to receive, as input, for each molecule, data (10) representative of said first (MSI) and second (MS2) spectra and said first retention time (Rtl) , the method (500) comprising the following steps performed by the electronic processing unit (303) , processing, by a first software module (20) , said input data (10) to group the different adducts of thesame molecule into a single compound and to annotate each compound based on a comparison of said data (10) with data of known molecules recorded in one or more databases associated with said first software module (20) , said first software module (20) generating a first annotation (25) of the structure of each molecule for the aforesaid compounds ; processing, by a second software module (40) , said input data (10) to generate (42) a second annotation of the structure of each molecule, said second software module (40) being configured to compare, based on machine learning techniques, the first spectrum (MSI) of a molecule and the related second spectrum (MS2) , with those "in silico" of a database of known molecules associated with the second software module (40) in which said second spectra (MS2) were calculated; processing, by a third software module (60) , said input data (10) to generate (51) a third annotation of the structure of each molecule, said third software module (60) being configured to compare, based on machine learning techniques, the second spectrum (MS2) of the molecule with second spectra of known molecules present in a respective database associated with the third software module (60) ; processing, by a fourth software module (30) , thefirst (MSI) and second (MS2) spectra of at least a part of the compounds processed by said first software module (20) to generate a prediction (31) of the molecular formula of the molecules; calculating, by a fifth software module (80) , a predicted retention time value (Rt2) of a database (DB) of known molecules; for each molecule of the one or more complex mixtures (ES) , the method includes comparing (100) : the first annotation of the structure of the molecule generated by the first software module (20) , the second annotation of the structure of the molecule generated by the second software module (40) , the third annotation of the structure of the molecule generated by the third software module (60) , the prediction (31) of the molecular formula of the molecule generated by the fourth software module (30) , wherein the comparison step (100) comprises a step of checking a consistency (103) of predictions associated with the annotations generated by the first (20) , second (40) , and third (60) software modules based on: an evaluation of a first parameter, representative of the biological source of the molecule, predicted by the first (20) , second (40) , and third (60) software modules, a comparison of a second parameter, representative of themolecular formula of the molecule, predicted by the first (20) , second (40) , and third (60) software modules with respect to the prediction provided by the fourth software module (30) , an evaluation of a third parameter representative of the predicted retention time value (Rt2) , to select (200) which of the aforesaid first, second, and third annotations of the structure of the molecule reliably identifies the molecule.

2. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to claim 1, wherein said first parameter, representative of the biological source of the molecule, takes a tuple of three values {-1,0,1} where: the consistency of the first parameter is 1 if the biological source field in the database of known molecules (DB) takes the value "True"; the consistency of the first parameter is 0 if a molecular structure (InChlKey) has not been predicted by one of said first (20) , second (40) , and third (60) software modules or if the predicted molecular structure (InChlKey) is not present in the database (DB) of known molecules; the consistency of the first parameter is -1 if the field of the biological source in the database(DB) of known molecules takes the value "False".

3. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to claim 1 or 2, wherein said second parameter, representative of the molecular formula, takes a tuple of three values {-1,0,1} and, based on the comparison of the prediction generated by each of said first (20) , second (40) , and third (60) software modules with the prediction of the fourth software module (30) , the method provides that : the consistency of the second parameter is 0 if the software module (20, 40, 60) being compared does not include a molecular formula; the consistency of the second parameter is 1 if the two predicted molecular formulas are equal; the consistency of the second parameter is -1 if the two predicted molecular formulas are different.

4. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to any one of the preceding claims, wherein said third parameter, representative of the predicted retention time value (Rt2) , takes a tuple of three values {-1,0,1}, where: the consistency of the third parameter is 1 if the value of the predicted retention time (Rt2)deviates from the value of the first retention time(Rtl) by less than 1.4 minutes; the consistency of the third parameter is -1 if the value of the predicted retention time (Rt2) deviates from the value of the first retention time (Rtl) by more than 1.4 minutes; the consistency of the third parameter is 0 if no molecular structure (InChlKey) has been predicted by one of said first (20) , second (40) , and third (60) software modules or if the predicted molecular structure (InChiKey) is not present in the database (DB) of known molecules.

5. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to any one of the preceding claims, wherein said step of checking a consistency (103) further comprises a step of grouping ( 103 ’ ) into classes the three-value tuples representative of the consistency of the predictions generated by the first (20) , second (40) , and third (60) software modules, based on the logic: a class T groups the tuples having at least two 1 in any position; a class A is a subclass of class T the tuples of which have a 1 in the first position and at least one other 1;a class 0 contains tuples which only have a 1 in the first position.

6. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to any one of the preceding claims, wherein said selection step (200) comprises the further steps of : predicting (201) a molecular structure value (InChlKey) by three or two of said first (20) , second (40) , and third (60) software modules; or predicting (202) a molecular structure value (InChlKey) by only one of said first (20) , second (40) , and third (60) software modules.

7. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to the preceding claim, wherein said prediction step (201) comprises the further steps of: comparing in pairs (203) the molecular structure values (InChlKey) predicted by each of said first (20) , second (40) , and third (60) software modules; if the result of said comparison (204) returns two or three identical molecular structure values (InChlKey) , the method includes selecting (205) the name of the annotation returned by the process based on a ranking of the names of the predictions provided by said first(20) , second (40) , and third (60) software modules, from the most preferred to the least preferred, given by: annotation name provided by the second software module (40) , annotation name provided by the first software module (20) , annotation name provided by the third software module ( 0 ) .

8. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to claim 6, wherein said prediction step (201) comprises the further steps of: comparing in pairs (203) the molecular structure values (InChlKey) predicted for each software module; if the result of such a comparison (206) of molecular structure values (InChlKey) results in no identity, the method includes a step of comparing (207) in pairs the names of the compounds predicted by said first (20) , second (40) , and third (60) software modules: if the similarity of the names provided by two of said software modules is greater than a predetermined threshold (75%) , both software compared to each other are added to a list, the annotation name returned by the process is selected (208) based on a similarity ranking of the annotationnames provided by the software modules (20, 40, 60) , from the most preferred to the least preferred, including : annotation name provided by the second software module (40) , annotation name provided by the first software module (20) , annotation name provided by the third software module ( 0 ) ; if the similarity of the names provided by two of said software modules is lower than the preset threshold (75%) , the method includes a step of evaluating (209) the consistencies of the predictions generated by said software modules (20, 40, 60) based on a ranking, from the most relevant to the least relevant, determined by comparing the first parameter, representative of the biological source, the second parameter, representative of the molecular formula, the third parameter, representative of the predicted retention time (Rt2) .

9. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to claim 8, wherein said step of evaluating (209) the consistencies comprises:- in the case of at least two consistencies (210) and only one of said software modules (20, 40, 60) havingtwo consistencies, the annotation name provided by such a software module is returned by the process (211) ;- in the case of at least two consistencies (210) and the best two are even, the method includes solving such a tie based on an order of consistency based on a classification (212) of the software modules (20, 40, 60) including, from the most preferred to the least preferred :- third software module (60) ,- second software module (40) ,- first software module (20) , the annotation name resulting from such a tie is returned by the process ( 212 ’ ) .

10. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to claim 8, wherein said step of evaluating (209) the consistencies further comprises:- if a single consistency or no consistency is present (213) , the method includes the further steps of:- selecting (214) from said first (20) , second (40) , and third (60) software modules the software module best positioned in terms of first parameter, representative of the biological source, the annotation name provided by the selected software is returned by the process (215) ;in the case of a tie also in terms of said first parameter, the method includes solving such a tie based on (216) an order of consistency based on a classification of the software modules (20, 40, 60) which includes, from the most preferred to the least preferred :- third software module (60) ,- second software module (40) ,- first software module (20) , the annotation name resulting from such a tie is returned by the process (216' ) ; in the other cases (217) , the resulting annotation (218) of the process remains undecided.

11. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to claim 6, wherein the step (202) of predicting a molecular structure value (InChlKey) by only one of said first (20) , second (40) , and third (60) software modules comprises a further step of establishing (219) whether the annotation generated by said software module has a consistency class A or 0;- if the annotation generated by said software module is in class A, such a software module is selected (220) if it includes a tuple with a 1 in the first position and at least one other 1the name of the annotation (220' ) generated by such a software module is returned by the process; if the annotation generated by said software module is in class 0, such a software module is selected (221) if it includes a tuple with only a 1 in the first position, the annotation name resulting from such software ( 221 ’ ) is returned by the process; in the other cases (222) , the resulting annotation (222' ) of the process remains undecided.

12. A method (500) for automatically identifying molecules included in one or more complex mixtures (ES) according to claim 8, wherein said similarity threshold of the names provided by the software modules is 75%.

13. An electronic system (300) for automatically identifying molecules included in one or more complex mixtures (ES) , comprising:- a high-resolution mass spectrometry, HRMS, apparatus (301) connected to a high or ultra-high pressure liquid chromatograph, HPLC or UHPLC, (302) , said liquid chromatograph (302) being adapted to receive said one or more complex mixtures (ES) to measure a first retention time value (Rtl) of each molecule, said high-resolution mass spectrometry, HRMS, apparatus (301) being configured to measure, for each molecule ofsaid complex mixture, a mass / charge ratio or first spectrum (MSI) and mass / charge ratios of molecular fragments generated during a process of fragmentation of the molecule or second spectrum (MS2) ; an electronic processing unit (303) operatively associated with said mass spectrometry apparatus (301) to receive, as input, for each molecule, data (10) representative of said first (MSI) and second (MS2) spectra and said first retention time value (Rtl) , said electronic processing unit (303) comprising at least one processor (3031) and a memory block (3032, 3032' ) associated with the processor for storing instructions, said processor and said memory block being configured to carry out the steps of the method according to claims 1-12.

14. A computer program comprising an application code loaded on a memory block (3032, 3032' ) and executable by at least one electronic processing unit (303) of an electronic system (300) for automatically identifying molecules in one or more complex mixtures (ES) , to implement the method according to claims 1-12.

15. An electronically searchable, web-based database adapted to contain data representative of a plurality of molecules automatically identified starting from one or more complex mixtures (ES) by means of theidentification method according to claims 1-12.