Method for reconstructing a mass spectrum
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2026-04-08
AI Technical Summary
Current mass spectrometry techniques, such as tandem mass spectrometry, are limited by instrumental constraints, requiring complex and costly instruments, and are inefficient in differentiating ions with close m/z ratios, necessitating prior knowledge of the sample and being slower due to sequential data acquisition.
A method that reconstructs an individual mass spectrum from a series of mass spectra using correspondence scores and a grouping method, allowing for the identification of precursor ions and their fragments without the need for tandem mass spectrometry devices, utilizing in-source collision-induced dissociation and classification algorithms to group ions precisely.
Enables the reconstruction of individual mass spectra from mixed samples using single-stage mass spectrometers, reducing costs and complexity, and overcoming limitations related to measurement resolution and ionization sources, while identifying isobaric species and working with various types of mass spectra.
Smart Images

Figure EP2024065342_12122024_PF_FP_ABST
Abstract
Description
METHOD FOR RECONSTRUCTING A MASS SPECTRUM TECHNICAL FIELD OF THE INVENTION [1] The present invention generally relates to a method of reconstructing a mass spectrum. [2] It relates in particular to a method for reconstructing an individual mass spectrum of at least one precursor ion from the exploitation of a series of mass spectra. [3] It also relates to a device implementing such a method. STATE OF THE ART [4] Mass spectrometry (MS) is an analytical technique for detecting ions from a sample and analyzing these ions based on their ratio (m / z), where m represents the mass of an ion and z its electrical charge. Mass spectrometry is used in many applications to analyze, identify, and characterize the chemical structure of ionized molecules. [5] A mass spectrometer typically includes an ionization source to form ions from a sample to be analyzed, an analyzer that separates the ions based on their m / z ratio, and a detector. A mass spectrum is obtained by recording the abundance of ions as a function of their mass-to-charge ratio (m / z). However, simple mass spectrometry only allows the measurement of the m / z ratio and cannot differentiate between ions with identical m / z ratios, particularly due to the resolution limitations of the spectrometer used. In addition, measuring the m / z ratio does not provide information on the composition or structure of the ions. [6] It is known to use tandem mass spectroscopy to study ions of interest in a sample more precisely. Such a technique makes it possible to provide an individual mass spectrum for a given precursor ion after activation. To this end, tandem mass spectrometry includes, after ionization of the molecules in the sample, a step of selection by a mass filter of an ion of interest. This selected ion is then subjected to an activation step which consists of increasing its internal energy which can lead to its fragmentation. The ions produced by the fragmentation of the selected ion (also called precursor ion or parent ion) are then analyzed in one or more other mass spectrometry step(s) on the fragment ions thus generated, in which the mass analysis steps can be separated spatially or temporally depending on the instrument used. [7] For example, tandem mass spectrometry can be implemented by isolating an ion of interest in an ion trap by supplying it with an amount of internal energy sufficient through collision with a gas for it to fragment. Detection of the products of this fragmentation can provide information on the nature and / or structure of the precursor ion. Tandem mass spectrometry is the basis for applications of mass spectrometry in structural analysis and in particular the sequencing of proteins and other biopolymers (such as sugars or nucleic acids). [8] Tandem mass spectrometry is functional but requires more sophisticated instruments, usually including two mass analyzers coupled in series, which adds complexity to the spectrometers and also increases their costs and maintenance. [9] Tandem mass spectrometry is also limited by instrumental constraints, in particular the fineness of selection which in certain cases is not sufficient to separate two ions of interest with very close masses (very close m / z ratio). In this case, two different ions but with close m / z ratios will be selected and activated simultaneously. The tandem mass spectrum obtained will therefore be the sum of the tandem mass spectra of the two ions selected simultaneously, which complicates the interpretation of the measurements.
[0010] Tandem mass spectrometry can also be slower in terms of analysis speed compared to other methods due to the sequential acquisition of data. This is because the ions of interest are selected and fragmented one after the other.
[0011] Furthermore, in tandem mass spectrometry, a criterion for choosing the ions to be analyzed in a sample is necessary. In this case, it is often necessary to have a list of ions of interest that one wishes to fragment, which amounts to having a priori knowledge of the sample. Alternatively, it is also possible to choose to fragment all ions whose intensity exceeds a certain threshold, again set according to the sample. Prior knowledge of the sample is therefore necessary.
[0012] Data independent acquisition (DIA) mass spectrometry is an advanced technique used particularly in proteomics (see document US2017 / 0032948). However, it requires systematic fragmentation of all ions present in a series of m / z ratio windows covering a range of m / z ratios.
[0013] The invention aims to remedy at least one of the aforementioned drawbacks. PRESENTATION OF THE INVENTION
[0014] In order to overcome the aforementioned drawback of the state of the art, the present invention proposes a method for reconstructing a mass spectrum, said method comprising the following steps: - acquisition of a plurality of mass spectra measured on an ionized sample comprising at least one precursor ion and fragment ions, - determining a plurality of series of abundance values of the ionized sample from the plurality of mass spectra, each series of abundance values of the ionized sample being determined from the plurality of mass spectra, each series of abundance values comprising N abundance values where N is an integer greater than or equal to two; - arranging the plurality of abundance value series into pairs of abundance value series, each pair of abundance value series comprising a first abundance value series and a second abundance value series, - for each pair of abundance value series, calculating a correspondence score between the first abundance value series and the second abundance value series, so as to obtain correspondence scores for all pairs of abundance value series, the correspondence scores providing a correspondence score map between all pairs of abundance value series, - determination, based on a mass-to-charge ratio, of at least one precursor ion and the fragment ions associated with each determined precursor ion, by applying a grouping method to the calculated correspondence score card, - reconstruction of an individual mass spectrum of each determined precursor ion and each of the associated fragment ions from the correspondence score card.
[0015] Thus, thanks to the invention, in particular to the different steps of the method according to the invention, it is possible to go back to the mass spectrum of an individual ion present in a mixture (i.e. a sample which may contain several ions and fragments) without necessarily using a tandem mass spectroscopy device. As a result, the present invention can carry out tandem mass spectrometry analyses from a series of simple mass spectra having been acquired by a mass spectroscopy device which is not provided with a second mass filter (here an ion trap) but in which either the desorption / ionization step causes the fragmentation of the species, or by the implementation of an in-source activation step by collision-induced dissociation (or "in-source Collision-induced dissociation also known in English as in-source CID") making it possible to activate all of the ionized species.As a result, the process is easier to implement and less expensive compared to a process requiring working from tandem mass spectrometers.
[0016] Furthermore, although this method works from a series of mass spectra, the calculation step and the use of a classification algorithm make it possible to precisely group at least one precursor ion and the fragments associated with this precursor ion. Such a method is therefore functional on mass spectra acquired with devices whose selection width is sometimes too large to isolate an ion precisely. The invention can also make it possible to identify the presence of isobaric species in a tandem mass spectrometry experiment. Therefore, the method according to the invention also makes it possible to work on all types of acquired mass spectra. In other words, the method makes it possible to overcome the technical limitations of mass spectrum devices, which may be linked to the measurement resolution or to constraints linked to the ionization sources used.
[0017] Other advantageous and non-limiting characteristics of the method according to the invention, taken individually or in all technically possible combinations.
[0018] In one embodiment, the step of determining a series of abundance values comprises determining a series of ranges of mass-to-charge ratio values for the plurality of measured mass spectra of size NxM with M an integer greater than 1, and for each mass spectrum of the plurality of mass spectra, a step of integrating abundance values belonging to the same range of mass-to-charge ratio values of the mass spectrum considered, to provide, for each range of mass-to-charge ratio values of the spectrum considered, an integrated abundance value associated with a mass-to-charge ratio determined in each range of mass-to-charge ratio values, the integrated abundance values over the series of ranges of mass-to-charge ratio values forming a series of abundance values of the ionized sample.
[0019] In one embodiment, the match score is based on at least one of the following: - mutual information between the first series of abundance values and the second series of abundance values, said mutual information being based on an entropy of the first series of abundance values and on an entropy of the second series of abundance values; and / or - a correlation between the first series of abundance values and the second series of abundance values.
[0020] In one embodiment, for the mutual information, the entropy of the first series of abundance values is an entropy normalized by a first invariance parameter as a function of a first median value determined on this first series of abundance values and the entropy of the second series of abundance values is an entropy normalized by a second invariance parameter as a function of a second median value determined on this second series of abundance values.
[0021] In one embodiment, the correspondence score being based on mutual information, the method comprises a step of estimating the entropy of the first set of abundance values and a step of estimating the entropy of the second set of abundance values using a k-nearest neighbors algorithm.
[0022] In a particular and advantageous example, the k-nearest neighbors algorithm is applied to a single neighbor.
[0023] In one embodiment, when at least two matching scores based on two different elements are calculated in the calculating step, several score cards are provided following the calculating step, said method further comprising, for each score card, a step of filtering said score card according to a contrast of the considered score card.
[0024] In one embodiment, the method further comprises a step of processing the scorecard before applying the clustering method, said processing step using at least one of the following algorithms: - standardization algorithm, - algorithm for processing diagonal elements of the scorecard, - application of a square function.
[0025] In one embodiment, the score map is a two-dimensional matrix, wherein the clustering method determines the at least one precursor ion and its associated fragment ions from match scores located outside a diagonal of the score map.
[0026] In one embodiment, the clustering method uses at least one unsupervised learning algorithm.
[0027] In one embodiment, the clustering method is based on at least one of the following algorithms: - a hierarchical ascending classification algorithm, - a k-nearest neighbors algorithm, - a k-means algorithm, - a stochastically distributed neighbor incorporation algorithm.
[0028] In one embodiment, the method further comprises, after the reconstruction step, a step of analyzing the individual mass spectrum of the determined precursor ion and the associated fragment ions to characterize the chemical composition of the determined precursor ion and the chemical composition of the fragment ions associated with the precursor ion considered.
[0029] In one embodiment, the method comprises a step of processing the acquired mass spectra comprising at least one of the following calculations: - filtering of the mass spectrum considered by a Gaussian or normal kernel, - a normalization of the mass spectrum considered, - a smoothing of the mass spectrum considered, - a return to the baseline.
[0030] In one embodiment, the method comprises a step of processing the series of abundance values comprising at least one of the following calculations: - filtering the series of abundance values given by a Kalman filter; - filtering the series of abundance values given by a Fourier filter, - a filtering by moving average of the given series of abundance values.
[0031] In one embodiment, the method comprises a compressed acquisition step to reduce the number of acquisitions.
[0032] In one embodiment, the plurality of mass spectra comprises at least 1000 mass spectra, preferably between 2000 and 8000 mass spectra.
[0033] In an exemplary embodiment, an external variation by chromatographic separation, separation by ion mobility and / or by discriminating transmission of an ion optic is applied during the acquisition step so as to modify or modulate the abundance of the ions and the number of mass spectra of the plurality of mass spectra is less than or equal to 10, for example equal to 5 mass spectra.
[0034] Of course, the various features, variants and embodiments of the invention may be combined with each other in various combinations to the extent that they are not incompatible or mutually exclusive. DETAILED DESCRIPTION OF THE INVENTION
[0035] The description which follows with reference to the appended drawings, given as non-limiting examples, will make it clear what the invention consists of and how it can be implemented.
[0036] On the attached drawings:
[0037] Figure 1 is a schematic representation of a device for reconstructing an individual mass spectrum of at least one precursor ion according to the present disclosure;
[0038] Figure 2 is a schematic representation of a mass spectroscopy device used in the device according to the present disclosure;
[0039] Figure 3 is a schematic representation of an example of a method according to the present disclosure;
[0040] Figure 4 is a schematic representation of an embodiment of an acquisition step of a method according to the present disclosure.
[0041] Figure 5 is an example of an average mass spectrum determined in an acquisition step of a method according to the present disclosure;
[0042] Figure 6 is a schematic representation of m / z ratio value ranges determined from an average mass spectrum obtained in an acquisition step of a method according to the present disclosure;
[0043] Figure 7 is a representation of an example of a first series of abundance values and a second series of abundance values obtained in an arranging step of a method according to the present disclosure;
[0044] Figure 8 is a representation of a first series of abundance values and a second series of abundance values illustrated in Figure 7 for which the abundance values of each of the series have been filtered by a Kalman filter in a method according to the present disclosure;
[0045] Figure 9 is a schematic representation of a determining step included in a step of calculating a matching score in a method according to the present disclosure, said score being based on a calculation of mutual information;
[0046] Figure 10 is a succession of curves each illustrating a cloud of points of a plurality of mutual information determined from non-invariant entropies;
[0047] Figure 11 is a succession of curves each illustrating a point cloud of a plurality of mutual information determined from invariant entropies according to the present disclosure;
[0048] Figure 12 is a score map obtained at a step of calculating a matching score based on mutual information by applying a method according to the present disclosure;
[0049] Figure 13 is a score map obtained at a step of calculating a correspondence score based on partial correlation by applying a method according to the present disclosure;
[0050] Figure 14 is a score map obtained at a step of calculating a correspondence score based on partial correlation by applying a method according to the present disclosure, the correlation scores being determined from pairs of series of abundance values filtered by a Kalman filter;
[0051] Figure 15 is a schematic representation of a dendrogram obtained by a clustering method used in a method according to the present disclosure;
[0052] Figure 16 is an exemplary representation of an individual mass spectrum of a precursor ion of an ionized sample, said individual mass spectrum being obtained according to a method according to the present disclosure;
[0053] Figure 17 is an exemplary representation of another individual mass spectrum of another precursor ion of an ionized sample, said individual mass spectrum being obtained according to a method according to the present disclosure;
[0054] Figure 18 is a schematic representation of a two-dimensional map obtained by a clustering method used in a method according to the present disclosure;
[0055] Figure 19 is a schematic representation of the two-dimensional map obtained from Figure 18 to which another clustering algorithm has been applied.
[0056] There will be described with the aid of FIG. 1 and FIG. 2 a device 100 for reconstructing a mass spectrum of an individual ion present in a sample according to the present disclosure.
[0057] The device illustrated in Figure 1 comprises a measurement module 110 configured to determine at least one mass spectrum 2, preferably a plurality of mass spectra.
[0058] For this purpose, the measurement module 110 comprises a mass spectrometry device 10 as illustrated in FIG. 2. The mass spectroscopy device 10 may comprise a single-stage mass spectrometer device (as opposed to a tandem mass spectrometer).
[0059] In the example illustrated in Figure 2, the mass spectroscopy device 10 is a single-stage device. It comprises an ionization source 11 for forming ions from the sample 1 to be analyzed and an activation means 12 configured to fragment the ions. In the present disclosure, the chemical composition of the sample is unknown. Therefore, it comprises a mixture comprising one or more ions suspended in a solution. In other words, the sample may comprise different molecular species.
[0060] In the present disclosure, sample 1 is for example a solution containing molecules. This sample 1 is then placed in the gas phase and is then ionized by the ionization source 11.
[0061] The ionization source 11 comprises at least one of the following ionization sources, namely: an electron impact ionization source, an electrospray source, a chemical ionization source, a photoionization source, a laser ionization source, a matrix-assisted laser induced desorption ionization source (known by the acronym MALDI source), a surface-assisted desorption ionization (SELDI) source, an atmospheric pressure MALDI source, a matrix-assisted fast evaporation ionization source, any other method involving a matrix or a surface, any method involving a laser, a plasma desorption ionization source, an atmospheric pressure chemical ionization source or an atmospheric pressure photoionization source, an atmospheric pressure ionization source.
[0062] In a preferred embodiment of the invention, the ionization source 11 comprises an electrospray source. These ions are then activated by the activation means 12 configured to fragment the ionized ions in the gas phase in an in-source activation step (CID). All the ions are activated simultaneously and can produce electrically charged fragments or fragment ions, also referred to as fragments hereinafter.
[0063] In these different variants, the activation means 12 comprises at least one of the following means: activation by collision, activation by absorption of electromagnetic radiation, thermal activation, etc.
[0064] In another embodiment, the mass spectrometry device 10 is devoid of activation means 12. In this case, the ionization source can also be configured to fragment the ions of the sample. In other words, this type of ionization source 11 provides sufficient energy during ionization to produce fragments. In practice, such an ionization source 11 comprises for example an ionization source 11 by electron impact or an ionization source 11 using laser desorption ionization (LDI) technology.
[0065] In another embodiment (not shown), the mass spectrometry device 10 comprises a first mass filter for selecting a range of m / z ratios of interest, an activation means 12 and a second mass filter for analyzing fragmentation products resulting from the activation. In this embodiment, one or more precursors are selected and activated simultaneously. These ions will each produce one or more fragment ions. The recorded tandem mass spectrum therefore includes all of the fragments generated and possibly the precursor ions. Such a mass spectroscopy device is also called a two-stage mass spectrometry device.
[0066] The ions obtained at the output of the activation means 12 are called ionized samples. They include, in particular, at least one ion, called a precursor ion.
[0067] In the present disclosure, a precursor ion or parent ion is an ion that can be decomposed into different fragments. Therefore, a fragment ion (or fragment) is a part or chemical structure of the parent ion.
[0068] The mass spectroscopy device 10 also includes a mass filter 13 (also called analyzer) which separates at least one precursor ion as well as the different fragments of the ionized sample according to their m / z ratio.
[0069] In the present disclosure, the abundance of each ion can be used to obtain quantitative information about the concentration of the molecular species in the sample. The mass spectroscopy device 10 also includes a detector 14 configured to detect ionized species (ionized ions). Mass spectrum 2 is obtained by recording the abundance of ions as a function of their mass-to-charge ratio (m / z).
[0070] A mass spectrum 2 thus provides information on the chemical composition of the sample, in particular on the number of species present, typically ions and fragments, and on their respective mass.
[0071] As illustrated in Figure 2, the mass spectroscopy device lacks a second mass filter required to perform a tandem mass spectrometry measurement.
[0072] In the present disclosure, the mass spectrometry device 10 is also configured to acquire a time series of mass spectra 2. By time series, it is meant that the mass spectrometry device 10 successively acquires (over time) mass spectra. In practice, the sample 1 used to acquire these mass spectra 2 is identical for each acquisition of mass spectrum 2 by the mass spectrometry device 10. However, the ionized sample of each acquisition (of mass spectrum) may vary during the acquisitions.
[0073] Optionally, the mass spectrometry device 10 may also comprise a computer (not shown) comprising a memory for storing the mass spectra acquired at the output of the detector.
[0074] The device 100 illustrated in FIG. 1 also comprises a processing unit 120.
[0075] By processing unit 120 is meant a computer or a processor or a central processing unit (CPU) or any electronic device making it possible to implement a succession of commands and / or calculations. Typically, the processing unit 120 comprises a processor, a memory and various input and output interfaces. For this purpose, the processing unit 120 is configured to control the operations and functions performed by the device 100. The input and output interface may be the connection point allowing the device 100 to receive data and send commands or data. The processing unit 120 may be connected to other processing units to exchange information with them.
[0076] The processing unit 120 comprises for this purpose one or more (calculation) modules configured to carry out a succession of commands and / or calculations which will be explained below with figure 3. Typically, the processing unit is configured to implement the method explained below.
[0077] The device 100 may also comprise a communication module 130 configured to exchange data with elements connected to the device 100 and / or elements internal to the device 100. By exchanging data is meant sending and / or receiving data, the received data corresponding in this example to input data which may comprise a multitude of mass spectra, for example a part of the series of mass spectra or the entire series of mass spectra acquired by the measurement module 110.
[0078] In this example, the communication module 130 receives the mass spectra determined by the measurement module 110 and stores them in the storage module 140.
[0079] The communication module 130 is, optionally in this example, connected to a database 4, external to the device 100. The database 4 typically comprises previously acquired mass spectra.
[0080] The database 4 is used in particular when the device 100 does not include a measuring device 110. Therefore, in this case, the device 100 does not directly measure mass spectra 2 but extracts, from the database 4, mass spectra 2 which have been previously recorded. Therefore, it is possible to work with previous mass spectra, no measurement then being necessary. In this case, these received mass spectra (from the database 4) have preferably all been acquired by the same mass spectroscopy device 10.
[0081] In the example considered, the communication module 130 receives mass spectra 2 from the measuring device 110 (for example a series of mass spectra) and from the database 4. Similarly, these received mass spectra were acquired by the same mass spectroscopy device 10. In other words, the mass spectra of the database were previously acquired by the measuring device 110 and recorded in the database, external to the device 100. In a variant, this database is included in the measuring device 110 (memory of the device 10) but remains different from the memory of the measuring device 110.
[0082] To receive the data from the database 4, the communication module 130 is also configured to be connected to at least one communication network 5 which may be a wireless communication network or a wired communication network. In practice, the wireless communication network uses at least one of the following technologies: WiFi technology, Bluetooth technology, 2G / 3G / 4G / 5G technology, etc. The wired communication network uses ADSL technology, etc. Thus, in the device 100, the mass spectra measured by the measuring device 110 or extracted from the database 4 are received by the communication module 130 which transfers this data to the storage module 140 of the device 100.
[0083] The storage module 140 comprises, for example, a memory and / or a hard disk.
[0084] The storage module 140 stores in particular computer program instructions designed so that the processing unit 120 implements at least some of the steps of the method described below with reference to FIG. 3 when these instructions are executed by the processing unit 120.
[0085] All mass spectra received by the communication module 130 are concatenated to form a single series of mass spectra. This single series of mass spectra generally comprises at least 1000 mass spectra, preferably between 2000 and 8000 mass spectra in the purely stochastic case. Note that all experimental conditions can vary between spectra of the same series: pressure, temperature, voltage, etc. It is sufficient that the abundance relationships between precursors and fragments are preserved. Conversely, experimental conditions in which there would be no more fragmentations do not allow the method to be applied. This is the only exclusion of the present method.
[0086] Furthermore, if external perturbations of the abundances are applied (for example, variations in concentrations due to chromatographic separation, ion mobility separation, or any other perturbation such as discriminating transmission of ion optics), then the number of measurements required is reduced compared to the stochastic case. In the non-stochastic case where an external perturbation modifies or modulates the abundances of the ions, the number of mass spectra can be reduced to 5 spectra to form a series of measurements usable on a reduced number of acquisitions.
[0087] Thus, the mass spectra measured by the measuring device and / or the mass spectra extracted from the database 4 are concatenated in the storage module 140 to form a plurality of mass spectra.
[0088] The storage module 140 thus stores this plurality of mass spectra.
[0089] Preferably, the plurality of mass spectra is a time series of mass spectra. In other words, all the mass spectra of the series of mass spectra are measured by the same mass spectroscopy device 10 for the same sample 1. Such a configuration has the advantage of obtaining a more precise series of mass spectra because it is directly acquired with the same mass spectroscopy device 10. In a variant, the series of mass spectra can be acquired by several similar mass spectra devices 10 and under similar measurement conditions. This has the advantage of being faster to obtain the series of mass spectra.
[0090] Figure 3 represents an example of a method implemented by the device 100.
[0091] This method 1000 is a method of reconstructing a mass spectrum of at least one precursor ion from a plurality of mass spectra measured on the ionized sample comprising at least one precursor ion and fragment ions.
[0092] In practice, the method 1000 is implemented by computer, for example by being implemented by the device 100 illustrated in FIG. 1.
[0093] The method 1000 can also be implemented in microcontrollers or programmable gate arrays (e.g. FPGAs).
[0094] This method 1000 comprises a step 1010 of acquiring a series of spectra mass on the ionized sample comprising at least one precursor ion and fragment ions. In the present disclosure, acquisition means the direct measurement of mass spectra and / or the reception of mass spectra.
[0095] For this purpose, the acquisition step 1010 comprises a preliminary (optional) step 1001 for recovering the plurality of mass spectra. In practice, the preliminary step 1001 begins with a step 1002 of measuring mass spectra by the measuring unit 110 of the device 100 and / or a step 1004 of receiving mass spectra from the database 4. An example of a preliminary step is illustrated in FIG. 4. In this example, the preliminary step comprises the step 1002 of measuring mass spectra by the measuring unit 110 and the step 1004 of receiving mass spectra from the database 4. In this example, all of these mass spectra (from the measuring module 110 and the database 4) are concatenated in a concatenation step 1006 in the memory module 4 to form the plurality of mass spectra of the at least one precursor ion and fragment ions.
[0096] This plurality of mass spectra (or series of mass spectra) is then detected and acquired 1008, by the processing unit 120 of the device 100, to reconstruct a mass spectrum of at least one precursor ion from this series of mass spectra.
[0097] The method 1000 illustrated in FIG. 3 then optionally comprises a step 1015 of processing the plurality of acquired mass spectra. Typically, the processing unit 120 processes all the mass spectra of the plurality of mass spectra.
[0098] Different treatments can be implemented on the mass spectra of the mass spectrum series. The treatments include at least one of the treatments below: - smoothing, for example, using a Savitzky-Golay algorithm with a baseline correction algorithm, etc.; - normalization, for example by dividing each acquired mass spectrum by a constant a which is a function of the acquired mass spectrum. In practice, the constant a corresponds to the abundance (intensity) of a peak in the mass spectrum, in particular to the abundance of the most intense peak (i.e. the one with the highest intensity). Thus, for normalization, each peak in the mass spectrum considered is divided by the abundance of the most intense peak in the mass spectrum; - peak extraction known in English as peak picking algorithm; - elimination or filtering, for example using a signal-to-noise ratio to remove irrelevant mass spectra; - deisotoping, - grouping of data, for example using techniques known as binning; - filtering, for example by convolving each spectrum to a particular kernel.
[0099] In a non-limiting manner, the processing unit 120 may filter each mass spectrum of the plurality of mass spectra in order to attenuate noise on each spectrum related to the measurement of these mass spectra which may impact the abundances of the m / z ratios. Here, typically, the processing unit 120 determines the signal-to-noise ratio of each mass spectrum of the plurality of mass spectra and eliminates, for example, the mass spectra comprising a signal-to-noise ratio less than or equal to 0.5.
[0100] The processing step 1015 may also include compressed acquisition processing, also known as “compressed sensing”. This processing may be carried out as an alternative to the processing listed below or in combination thereof.
[0101] After filtering, the processing unit 120 normalizes each mass spectrum of the plurality of acquired mass spectra. For this, for each mass spectrum, the processing unit 120 determines the peak having a maximum abundance (for example via the peak picking algorithm) and then divides the abundances of the other peaks of the given mass spectrum by the previously determined maximum value. The normalization constant a is made to vary from one mass spectrum to another.
[0102] After normalization, the processing unit 120 can perform peak alignment processing to select ranges of given m / z ratios. This has the advantage of subsequently working on only a portion of the mass spectra in the mass spectrum series.
[0103] Following the processing step 1015 (if present in the method 1000), the method 1000 then comprises a step 1020 of determining a plurality of series of abundance values of the ionized sample. In practice, in the determining step 1020, a series of abundance values of the at least one precursor ion and a series of abundance values of each fragment ion are determined. In the present disclosure, each series of abundance values comprises N abundance values where N is an integer greater than or equal to two. In practice, each integer assigned from 1 to N corresponds to a mass spectrum of the plurality of mass spectra.
[0104] In a first variant, the determination step 1020 may be implemented by the processing unit 120 as follows. For example, the processing unit 120 scans each mass spectrum used as input data of the determination step 1020 and for each mass spectrum, retrieves and records an abundance value for each m / z ratio, thereby creating a series of abundance values (i.e. equivalent to a list of abundance values) for each m / z ratio of the mass spectra of the plurality of mass spectra. Here, by series of abundance values, is meant a succession of abundance values obtained from the same given m / z ratio. Thus, typically, for a given m / z ratio, each value in the series of abundance values corresponds to an acquisition (i.e. a measurement), i.e. to a given mass spectrum N. There are therefore, in a series of abundance values, as many abundance values as there are mass spectra used in the determination step 1020. In this example, there are therefore as many series of abundance values as there are m / z ratios detected in the mass spectra of the mass spectrum series. Each series of abundance values is typically recorded in the memory of the processing unit 120 and / or in the storage module 140.
[0105] In a second variant, the abundance series values correspond to abundance values integrated over a range of m / z ratio values.
[0106] To obtain this type of abundance value series, the processing unit 120 may first calculate an average mass spectrum from the plurality of mass spectra by averaging, for each m / z ratio of the mass spectra of the plurality of mass spectra, the abundance values. An average spectrum is thus obtained.
[0107] For example, Figure 5 illustrates an average mass spectrum obtained from the plurality of mass spectra acquired from a sample comprising melittin and substance P. The average mass spectrum illustrated in Figure 5 was obtained from a plurality of mass spectra. The plurality of mass spectra were each obtained by electrospraying a mixture comprising melittin, and substance P dissolved in a mixture of water and methanol having a concentration of 10' 5mol / l. For example, the mixture was activated by 20eV ultraviolet radiation for 500ms to provide the ionized sample.
[0108] The processing unit 120 can then determine ranges of m / z ratio values from the average mass spectrum, for example by detecting peaks in the average mass spectrum and calculating for each peak (22, 23, 24), an interval of m / z ratios comprising at least the peak considered.
[0109] For example, Figure 6 shows a schematic representation of an average mass spectrum obtained from the N mass spectra of the plurality of mass spectra. In particular, in this example, each interval W1, W2, W3 covers all the abundance values of the peak considered 22, 23, 24 and can also cover m / z ratio values around the peak considered. Thus, a range of m / z ratio values corresponds to a part of the m / z ratios of the average mass spectrum. In this example, three ranges of m / z ratio values W1, W2 and W3 are detected by the processing unit. In this variant, a range of m / z ratio values W is associated with a peak of the average mass spectrum. These ranges of m / z ratio values are then recorded. The processing unit 120 can create a series of ranges of m / z ratio values of size NxM, with N corresponding to the number of mass spectra of the plurality of mass spectra (here an integer greater than to 2) and M corresponding to the number of ranges of m / z ratio values (here an integer greater than 2) and which includes, in the example of figure 6, three elements (the ranges W1, W2, W3). In practice, the M ranges of m / z ratio values are similar for each mass spectrum N, which makes it possible to integrate all the mass spectra in a similar way.
[0110] Subsequently, the processing unit 120 can, for each mass spectrum, integrate abundance values over given parts of each mass spectrum, each given part corresponding to a previously determined range of m / z ratio values. Typically, for a mass spectrum N, the mass spectrum considered is integrated over different ranges of values defined by the M ranges of values of the series of ranges of m / z ratio values at the coordinate N of the series of ranges of m / z ratio values.
[0111] An integrated abundance value is therefore obtained for each determined m / z ratio value range. The integrated abundance values over the same m / z ratio value range across the plurality of mass spectra makes it possible to obtain an abundance series for an ion or fragment ion. Thus, from all the m / z ratio value ranges, series of abundance values are obtained.
[0112] In this embodiment, each series of abundance values is subsequently associated with a specific m / z ratio included in the range of m / z ratio values considered. The specific m / z ratio may correspond to the median value of the m / z ratio of the range of m / z ratio values, or preferably to the m / z ratio associated with a maximum abundance value of the average mass spectrum over the given range of m / z ratio values. Thus, according to this variant, each series of abundance is associated with a peak of the average spectrum. Fewer series of abundance values are therefore calculated in this variant, which improves the speed of the method 1000. Furthermore, since each series of abundance values is calculated over a range of m / z ratio defined on the basis of a peak of the average spectrum, it can be considered that each series of abundance values corresponds to an ion of the ionized sample.Furthermore, integrating the sample abundance values allows ions from the ionized sample to be associated that belong to the same isotopic distribution. When the resolution allows it, the isotopes are separated but we want to group them as a single ion. Of course, this second variant, which may correspond to a step of grouping the isotopes (de-isotoping), can be done by other algorithms known to those skilled in the art.
[0113] In practice, the series of abundance values determined in determination step 1020 can be recorded and grouped in a single file.
[0114] The method 1000 optionally comprises a step 1025 of processing the series of abundance values comprising at least one of the following calculations: - filtering the series of abundance values given by a Kalman filter; - filtering the series of abundance values given by a Fourier filter, - a filtering by moving average of the given series of abundance values.
[0115] Such a step 1025 makes it possible to remove noise from the abundance values of the abundance value series. For example, in the integration over the intervals (second variant of the determination step 1020), the processing unit 120 can apply a time filter to each interval. This has the advantage of attenuating the noise throughout the acquisition, and of highlighting temporal variations. Typically, the Kalman filter may be preferred to remove the integration noise. Typically, Figure 8 illustrates an example comprising a series of abundance values 32 of an ion and a filtered series of abundance values 33 of this same ion here filtered by a Kalman filter. It can be seen that the abundance values of the filtered series 33 fluctuate less. The signal is more stable.
[0116] Subsequently, the method 1000 comprises a step 1030 of arranging the plurality of series of abundance values into pairs of series of abundance values, each pair of series of abundance values comprising a first series of abundance values (denoted X) and a second series of abundance values (denoted Y).
[0117] In practice, the processing unit 120 retrieves a series of abundance values from the plurality of series of values and combines it with another series of abundance values from the plurality of series of abundance values. Preferably, there may be as many combinations as there are series of abundance values. Thus, for P series of abundance values determined (with P an integer greater than 2, P 2 of pairs of series of abundance values are determined by the processing unit 120 considering that a series of abundance values can form a pair of series of abundance values with itself).
[0118] Figure 7 illustrates a pair of abundance value series consisting of a first series of abundance values 30 associated in this example with an m / z ratio of 1365 and a second series of abundance values 31 associated in this example with an m / z ratio of 1424. Here, N corresponds to 5532.
[0119] The method 1000 also comprises, for each pair of series of abundance values, a step 1040 of calculating a correspondence score between the first series of abundance values and the second series of abundance values. This correspondence score is therefore calculated for all pairs of series of abundance values. The set of correspondence scores obtained for all pairs of series of abundance values provides a score map, which is typically a symmetric matrix of size PxP.
[0120] In practice, the match score is based on at least one of the following: - mutual information between the first series of abundance values and the second series of abundance values, said mutual information being based on an entropy of the first series of abundance values and on an entropy of the second series of abundance values; - a correlation between the first series of abundance values and the second series of abundance values.
[0121] The match score determined in calculation step 1040 is based on one of the above cases or a mixture of the above cases. However, the scorecard is composed of match scores of the same type, i.e. calculated in the same way using the same calculation method.
[0122] Of course, the processing unit 120 can calculate different correspondence scores, for example a score based on mutual information, a score based on a correlation and a score based on mutual information and on the previously calculated correlation.
[0123] The different correspondence scores that can be determined at calculation step 1040 will be explained below.
[0124] These matching score examples are determined for each pair of previously recorded abundance value series. For simplicity, the calculations described below describe an application of calculation step 1040 to a single pair of abundance value series.
[0125] Mutual information
[0126] In this variant, the calculation step 1040 comprises a step of determining the mutual information 1041. Different embodiments of the step of calculating the mutual information will be described below.
[0127] Figure 9, Figure 10, Figure 11 and Figure 12 illustrate an embodiment in which the matching score is based on mutual information (mathematics). In this embodiment, the mutual information is determined by an information theory method to be described below that uses entropy calculations.
[0128] Mutual information between two variables
[0129] In a first embodiment of step 1041, the mutual information between two variables is defined by the following formula: [Math. 1] I(X, T) = h(X) + h(Y) - h(X, T) with X corresponding to a first variable (here, for example, it can correspond to the first series of abundance values defined for an ion of the ionized mixture, the ion being able to be a precursor ion or a parent ion) and Y a second variable (here, for example, it can correspond to the second series of abundance values defined for an ion of the ionized mixture).
[0130] Here, hQ corresponds to the entropy (also called intrinsic entropy) defined for a given random variable expressed by the mathematical function: [Math. 2]
[0131] where p x (x) corresponds to the probability density function of the continuous random variable X (corresponding to the probability that a signal appears in the variable X) (here the abundance values of each acquisition considered).
[0132] The entropy h() can also be written for non-continuous variables by the following formula: h(X = - E"= i Pi * log2(Pi), with Pi corresponding to the probability that a signal appears in the variable X.
[0133] The density law g x cannot be directly calculated with a finite number of points as in the present example. Indeed, it is not possible to know the density law g x . It is therefore necessary to approximate the quantity h() by passing through an estimation. The two most used methods for the estimation of h() are - the histogram method, which consists of dividing a series of points into different equal intervals. Each interval contains a number of points ni , i being the ith interval and we associate a probability pi = ni / N where N is the total number of points. Thus a continuous variable is approximated as a discrete variable and we use the formula stated in the paragraph above. This method is invariant, but requires a large number of points to approach the theoretical limit by analogy with a Riemann integral. In addition, the number of intervals chosen offers a bias on the final estimate. - the k-nearest neighbor method, which originates from the midpoint method for numerically calculating an integral. This method is not invariant, but the entropy estimate converges more quickly than the histogram method and therefore requires fewer points. The final estimate using this method is given below.
[0134] Thus, to determine the correspondence score of each pair of series of abundance values, the processing unit 120 estimates (in an estimation step A1 of the determination step 1041 of the mutual information), for each pair, the entropy of the first series of abundance values and the entropy of the second series of abundance values using an estimation algorithm. These estimated entropies are denoted H.
[0135] Estimation step A1 will be described according to the present disclosure.
[0136] In the present disclosure, an equivalence is sought between the entropy h() and the estimated entropy H().
[0137] In practice, the entropy estimated for a variable X (corresponding here for example to the first series of abundance values) is expressed according to the formula: [Math. 3] H( ) = d * mean( log( distance (L)) + log(V) + psi(ri) — psi( ) with d corresponding to the dimension of the variable X (here equal to 1), log (distance corresponding to the set of values of the distances determined by the k-nearest neighbors estimation algorithm (the calculation of this distance is expressed below), n corresponding to the number of abundance values in the series of abundance values (here n is equal to N), L corresponding to the number of neighbors considered in the estimation algorithm (here equal to 1 for example when the distance is determined for a single neighbor), psi() corresponding to a Digamma function which can be expressed by psi(x) = r(x) - 1 which can satisfy a recursion psi(x+1) = psi (x) + 1 / x with psi(1) = -C for C = 0.5777 being the Euler-Mascheroni Constant, V corresponding to a volume of a unit hypersphere expressed according to the following formula: [Math. 4] with r the radius of the hypersphere (equal to 1), a gamma function which can be expressed by
[0138] The k-nearest neighbors algorithm calculates all distances between all points. Thus distance^) is a series of N values where the i ème value of the series corresponds to the L ième smallest distance from i ième point.
[0139] Different methods may be used to estimate the entropy of the first series of abundance values and the second series of abundance values. In particular, the processing unit 120 may use at least one of the following methods: a method based on histograms, a method based on a k-nearest neighbors algorithm, a kernel density estimation method, etc.
[0140] In practice, in the present disclosure, the processing unit 120 may use a k-nearest neighbors (kNN) algorithm to estimate the entropy H of a given series of abundance values used in calculating the mutual information. The kNN algorithm is chosen because this method is robust, compared to an estimation method using a histogram, and operates on small sample sizes.
[0141] Typically, the kNN algorithm uses a distance calculation between abundance values of the same series in order to approximate the probability density of the series.
[0142] In practice, the k-nearest neighbor method uses a Euclidean distance to estimate the joint entropy between the data (abundance values) of the first series of abundance values and the data of the second series of abundance.
[0143] In this example, the Euclidean distance used depends on the first series of abundance values and the second series of abundance values considered, in particular on the processing that may have been carried out at processing step 1015.
[0144] In the following, two consecutive abundance values of the first series of abundance values (associated with an ion) are denoted by A = (xi, Xj+i) and two consecutive abundance values of the second series of abundance values (associated with an ion) are denoted by B = (yi, y i+ i), with xi an abundance value from the first series of abundance values for an acquisition i, with Xj+i an abundance value from the first series of abundance values for an acquisition i+1, with yi an abundance value from the second series of abundance values for an acquisition i and with y i+i an abundance value from the second series of abundance values for an acquisition i+1.
[0145] In practice, when no normalization is calculated in the processing step 1015 or when the method 1000 does not include a processing step 1015, the Euclidean distance between the abundance values of a first series of abundance values and the abundance values of the second series of abundance values is expressed by the formula: [Math. 5] with A a variable (i.e. data corresponding here to the first series of abundance values (A=X)), B a variable (i.e. data corresponding here to the second series of abundance values (B=X)), Xj a variable corresponding to a signal from A (hence here to an abundance value included in the first series of abundance values at acquisition i, i.e. an abundance value of an ion of the ionized sample), Xj+i to the signal following Xj (hence the abundance value of the ion at acquisition i+1), yi a variable corresponding to a signal from B (hence here to an abundance value included in the second series of abundance values at acquisition i, i.e. an abundance value of another ion of the ionized sample), y i+i to the signal following yi (therefore corresponding to an abundance value of the other ion at acquisition i+1).
[0146] In the case of mass spectra normalization at processing step 1015, the Euclidean distance defined by the formula Math. 5 does not apply. Therefore, the entropy h() is not directly equivalent to the estimated entropy H() by the estimation algorithm.
[0147] Indeed, the entropy of a variable X (here the variable X corresponds to a series of abundance values of an ion included in the pair of abundance series) has two major flaws in contrast to discrete entropy. The first flaw is that it is not invariant after a change of variable, as described by the following formula: [Math. 6] h(kX) = h(X) + log(\k[) with k a multiplicative constant, X a variable corresponding here to a series of abundance of an ion (for example a precursor ion), h corresponding to the entropy of the variable X. And the second defect is that the entropy according to the definition [Math2.] can be negative.
[0148] Thus, invariant entropy corrects the possible negativity of entropy and the dimension problem after a change of variable. This entropy is expressed by the mathematical function: h(x) = - (x) / m(x))dx and is estimated in the same way as the equation [Maths.2] by the knn method on the series of values X divided by the invariant measure m(X).
[0149] In this variant, when estimating the entropy of the first series of abundance values and the second series of abundance values from data that have been normalized by the same normalization constant, denoted k (for example when a = k), the entropy is also invariant under multiplication by a multiplicative constant k (here k = a) as described by the following formula: [Math. 7] h(kX, kY) = h(X, Y) + log(k 2) with k is a multiplicative constant, X a variable corresponding in this case to the first series of abundance values, Y a variable corresponding in this case to the second series of abundance values.
[0150] Similarly, Math. 7 does not apply directly when the multiplicative constant of the first set of abundance values is different from that of the second set of abundance values, for example, when the normalization constant that was used to determine the data for the first set of abundance values is different from that used to determine the data for the second set of abundance values.
[0151] Thus, in the case of normalization at processing step 1015, to have an equivalence between the entropy h() and the estimated entropy H, it is necessary to work from invariant entropy and invariant estimated entropy.
[0152] By invariant, we mean that the entropy considered, for example the entropy estimated on a given series of abundance values (H(X), with X the given series of abundance values), is equal to the estimated entropy calculated on the same series of abundance values multiplied by a constant (H(kX) with X the series of abundance values given and k the multiplicative constant, here being able to be equal to the normalization constant used in processing step 1015).
[0153] To achieve such equivalence, the entropy estimated from the first series of abundance values used in the calculation of two-variable mutual information is normalized by a first invariance parameter and the entropy estimated from the second series of abundance values used in the calculation of two-variable mutual information is normalized by a second invariance parameter.
[0154] In practice, the processing unit 120 uses, for a given series of abundance values, the invariant parameter m(X) (also called invariant measure) defined according to the following three properties: i) m(X) = r x >0 ; ii) m kX) = km(X) = krx iii) m[X + b) = m( ) = r%. with X the series of abundance values considered, m(X) the invariant measure of the series of abundance values considered, r x > 0 being a median value of the set of distances between each pair of abundance closest to the series considered, k a chosen multiplicative constant (which can depend on the normalization constant a of the processing step) and b an additive constant whose value can for example depend on noise (for example on noise included in the abundance values of the series of abundance values). Using the properties defined above, we can obtain an equivalence between the entropy h() and the estimated entropy H(), so that: [Math. 8] with X a series of given abundance values, h the intrinsic entropy, H the estimated entropy determined in the calculation step, m(x) the invariant measure. (It should be noted that the invariance property also allows us to have the equivalence H( ) )■
[0155] The Math. 8 formula can be generalized to an entropy between several variables, typically when several series of abundance values are extracted from the plurality of mass spectra. In other words, this is particularly used when several series of abundance values of ionized ions and / or ionized ion fragments are to be taken into account. In this case, the equivalence of the Math. 8 formula generalizes in n dimensions (n = N): [Math. 9] XYZ h X,Y, ...,Z) = H m(X) 'm(Y) ' ' m(Z) with X, Y.... Z corresponding to the series of abundance values of the ion X or Y or Z considered.
[0156] Thus, the equivalence sought for an entropy calculated between two variables is written as follows: [Math. 10] with X and Y two series of abundance values where X is associated in the sequence with the first series of abundance values belonging to a given pair of series of abundance values (here corresponding to a series of abundance values of an ion of the ionized sample) and Y with the second series of abundance values belonging to the same pair (here corresponding to a series of abundance values of another ion of the ionized sample), m(X) with the first invariance parameter of the series X (i.e. invariance parameter of the first series of abundance values) and m(Y) with the second invariance parameter of the series Y (i.e. invariance parameter of the second series of abundance values), h the entropy and H the estimated entropy.
[0157] Based on the above properties, when the multiplicative constant of the first series of abundance values (here A=X) is identical to the multiplicative constant of the second series of abundance values (here B=Y), the Euclidean distance used in the step of determining the mutual information 1041 can also be expressed by the following formula: [Math. 11]
[0158] Indeed, by multiplying the first series of abundance values and the second series of abundance values by the same multiplicative constant (k), the abundances at each acquisition of the first series of abundance are written kA = (kxi, kxi+i) and the abundances at each acquisition of the second series of abundance are written kB = (kyi, kyi+i). Typically, k can be a negative or positive number.
[0159] When the multiplicative constant of the first series of abundance values (denoted here ai) is not identical to the multiplicative constant of the second series of abundance values (denoted here 82), the Euclidean distance used in the determination step 1041 can also be expressed by the following formula: [Math. 12] with ai being the multiplicative constant of the first set of abundance values and a2 being the multiplicative constant of the second set of abundance values. Similarly, a1 and a2 are two constants that can be negative or positive.
[0160] To calculate such a distance d(aiA, a2B), the processing unit 120 can perform a change of variable by setting: BB' = — ry mules Math. 13 and Math. 14, we obtain the following equations:
[0163] The processing unit 120 thus uses the formulas described above to define in its estimation algorithm the Euclidean distance according to the formula described below: [Math. 19]
[0164] It will now be described how the processing unit 120 determines the invariance parameter of the two series of abundance values of each pair of series of abundance values (step of determining the invariance parameter A11 of each series of the pair of series considered in the estimation step A1).
[0165] As explained above, each invariant measurement is associated with a series of abundance values. The processing unit 120 therefore determines for each series of abundance values determined previously, an invariant measurement (step A11). This invariant measurement, as explained above, depends on the values of the abundances of the series of abundance values.
[0166] In practice, the invariant measure can be determined upstream of the estimation of determination 1041, for example following determination step 1020.
[0167] An example of a step of determining an invariance parameter on a first series of abundance values and an invariance parameter on a second series of abundance values (step A11) will therefore be described.
[0168] For a given series of abundance values (e.g. the first series of abundance values), the processing unit 120 calculates (step P1) all the differences between the abundance values included in the series of abundance values. For this, the processing unit 120 scans the series of abundance values and calculates, for each abundance value, a difference with all the other abundance values in the series of abundance values. Thus, for each abundance value in the series of abundance values, N-1 differences are determined. A list of difference values is, for example, recorded for each abundance value.
[0169] The processing unit 120 then searches, for each abundance value in the series of abundance values, for the smallest difference between this abundance value and another abundance value included in the series of abundance values (step P2). To do this, the processing unit scans the list of differences associated with the given abundance value and searches for the smallest difference obtained in the list of differences associated with this abundance value. As a result, each abundance value is associated with a minimum distance. Thus, N minimum distances are determined for this series of abundance values considered. These minimum distances are then arranged in ascending order.
[0170] The processing unit 120 then determines a median from these N minimum distances (step P3) and then multiplies this median value by the number N of abundance values of the series of abundance values considered (step P4). This last calculated value is then assigned to the invariance parameter (here the first invariance parameter) of the series considered.
[0171] These steps are also performed on the other series of the mass spectrum pair (e.g. on the second series) to obtain the second invariance parameter.
[0172] Thus, the step of determining the invariance parameter A11 provides for each series of abundance values, an invariance parameter. The invariance parameter respects the three properties i, ii) and ii) explained above.
[0173] The estimation step A1 comprises a step of normalization (step A12) of each series of abundance values by its invariance parameter. Thus, for example, in this step, the processing unit 120 divides each abundance value of the series of abundance values considered by the invariance parameter determined on this series. In the following, the first series of normalized abundance values is noted for the first set of abundance values and the second set of abundance values normalized is noted for the second series of abundance values. m(Y r
[0174] Therefore, after the normalization step, each series of abundance values is normalized.
[0175] The entropy estimation in the estimation step performed on pairs of mass spectra uses these series of normalized abundance values.
[0176] Thus, in the method 1000, in the estimation step A1 (included in the determination step 1041 of the mutual information), the processing unit 120 determines (step A13) an estimated invariant entropy using the estimation algorithm (for example, the kNN algorithm in this example) as explained below.
[0177] In practice, in this example, the estimation algorithm determines for - the invariant entropy Æ(X) for the ion associated with the first mass spectrum, - the invariant entropy h(Y) for the ion associated with the second mass spectrum, And - the invariant entropy in two dimensions MX, Y)} as defined in Math. 10 equation using Math. 3 equation as defined above to determine each estimated entropy H. The distance formula used in the kNN algorithm is the Euclidean distance defined in Math. 12 formula with k preferably being 1.
[0178] In the step of determining the mutual information 1041, the processing unit determines (step A2) the conditional information l(X,Y) between the first series of abundance and the second series of abundance on the basis of the formula Math. 1 now written thanks to the previous calculations using the estimated entropies H: [Math. 20] / (X, 7) = H(X) + H (7) - H(X, 7)
[0179] The mutual information determination step 1041 is calculated for each pair of series in the mass spectrum series. Thus, for each pair of abundance value series, the calculated mutual information I corresponds to a matching score between the first abundance series and the second abundance series of a given abundance series pair.
[0180] Thus, the set of correspondence scores calculated in the calculation step 1040 (in a step 1042) can be represented in the form of a two-dimensional matrix, with for example along the rows, all the first series of abundance values of the set of pairs and along the columns all the second series of abundance values of the pairs. From this each “pixel” of the matrix corresponds to a correspondence score calculated from a given pair of ions, the ions being able to be identified with the row and column of the matrix. These data, illustrated in two dimensions, correspond to the scorecard (which is a two-dimensional matrix).
[0181] Figure 10 illustrates different scatterplots of mutual information calculated from non-invariant estimated entropies H (i.e. from a first series of abundance values and a second series of abundance values not normalized by their invariance parameter explained above). These curves were plotted by taking on the abscissa the abundance values (intensity) of an ion (here in particular the ion associated with a first series of abundance values which has an m / z ratio of 1222) and on the ordinate the abundance values of the other ion of the pair considered (here the ion associated with the second series of abundance values which has an m / z ratio of 1712). In this example, a mixture of proteins comprising ubiquitin, cytochrome C and lysozyme in solution in a mixture of water and methanol having a concentration of 10' 5mol / l. A region centered on m / z 1760 and with a width of 100 m / z was selected for a tandem mass spectrometry step. The selected ions were then subjected to activation of 20eV ultraviolet radiation for 500ms. Typically, here, the multiplicative constant k was applied to the first series of abundance values and estimated for both series was calculated with the formula Math. 3 with a given multiplicative constant k, namely with k equal to 1 for curve 50, k equal to 5 for curve 51, k equal to 10 for curve 52.
[0182] Figure 11 illustrates different scatterplots of mutual information calculated from an estimated invariant entropy H (i.e. from a first series of abundance values and a second series of abundance values normalized by their invariance parameter as explained above). Typically, the estimated entropy H for both series was calculated with the formula Math. 3 with a given multiplicative constant k, with k equal to 1 for curve 53, k equal to 5 for curve 54, k equal to 10 for curve 55.
[0183] It can be seen that for curve 50, curve 51, curve 52, the mutual information determined with the different values of the multiplicative constant are different. Thus, the multiplicative constant used in the calculation of the entropy of the first series of abundance values affects the calculation of the mutual information. Conversely, for curve 53, curve 54, curve 55, the mutual information determined with different values of the multiplicative constant are similar. Thus, the multiplicative constant used in the calculation of the entropy of the first series of abundance values does not affect the calculation of the mutual information.
[0184] Therefore, thanks to the invention, the calculated mutual information is invariant, that is to say it is constant thanks to the estimated invariant entropy and therefore does not depend on the abundances of the ions of the spectra used. The invariant entropy estimated according to the This disclosure therefore makes it possible to express any estimated entropy value in a single frame of reference.
[0185] In the method 1000, here in particular in step 1041, the mutual information is calculated for all pairs of series of abundance values. All of this mutual information can be concatenated into a score map (step 1042). Thus, in this example, all of the mutual information determined in step 1041 makes it possible to obtain the score map.
[0186] Figure 12 illustrates a score map obtained from a matching score based on a mutual information calculation. The mutual information calculation was obtained from a series of abundance values obtained on an ionized sample, the series of abundance values comprising values integrated over 10 ranges of mass-to-charge ratio values (i.e., according to the second variant of the determination step). The ionized sample is the same as that of Figure 10 and Figure 12 comprising a protein mixture comprising ubiquitin, cytochrome C and lysozyme. In this figure, the 10 ranges of m / z ratio values of the ions of the first series of abundance values of the pairs are illustrated at the column level and the 10 ranges of m / z ratio values of the ions of the second series of abundance values of the pairs are illustrated at the row level.As illustrated in this figure, each value range is associated with an m / z ratio determined as described above.
[0187] We can see that the resulting score map is a symmetric matrix, which has the advantage of only performing calculations on half of the score map. In addition, since fewer pairs of abundance value series were used, the implementation time of the 1000 method is improved.
[0188] A second embodiment of step 1041 will now be described. In the following, the calculated mutual information is directly written from the estimated mutual information H thanks to the equivalence between the entropy h and the entropy H discussed above.
[0189] Conditional mutual information.
[0190] According to this embodiment, the mutual information used is conditional mutual information which is defined by the following formula: [Math. 21] / (X, Y|Z) = H(X,Z) + H(Y,Z) - H(X, Y,Z) ~ H(Z) with X, Y, and Z an additional variable, called the conditional variable, H() the estimated entropy.
[0191] In practice, the conditional variable Z includes at least one of the following: the total number of ions, known as the Total Ions Current (TIC) Typically, the total number of ions is the result of a sum of all the ions in a spectrum mass (of the plurality of mass spectra), a series of abundance values of another ion of the mixture, a series of constant abundance values, etc. In the following, the TIC is considered. In this case of the TIC, each mass spectrum of the plurality of mass spectra has a TIC. All the TICs of the plurality of mass spectra form a series of TICs which is used as a partial or conditional variable.
[0192] Therefore, from the series of mass spectra, the processing unit 120 can calculate a TIC (the conditional variable) for each mass spectrum of the plurality of mass spectra. This can for example be done upstream, for example in the determination step 1020. Thus, from the plurality of mass spectra, the processing unit 120 obtains a series of conditional variables composed of p conditional variables with p equal to the total number of mass spectra of the plurality of mass spectra (here P = N).
[0193] Conditional mutual information is determined using a process similar to that explained above for the example of mutual information between two variables.
[0194] Thus, in the calculation step 1041, the processing unit uses the kNN algorithm to estimate the entropies H(X, Z), H(Y, Z), H(X, Y, Z) and H(Z) of the equation Math. 21. Similarly, each estimated entropy term is an invariant entropy as explained above.
[0195] In practice, the step of estimating the conditional mutual information comprises steps A1 to A2 as described above included in the step of calculating the mutual information 1041.
[0196] Thus, for each pair of abundance value series, an invariant measure is determined for each variable X (corresponding to the first abundance value series), Y (corresponding to the second abundance value series), Z (corresponding to the TIC value series) in step A11 of the entropy estimation step H of the calculation step 1041.
[0197] In practice, to estimate the entropy H(Z) of the conditional variable, the processing unit 120 calculates directly from the series of conditional variables Z, the invariant measure m(Z) associated with this series. To obtain the estimated entropy H(Z) of this series, the processing unit 120 normalizes the series of conditional values by dividing each conditional variable of the series of conditional variables Z by the measure invariant m(z) (calculation of ). Similarly, the processing unit 120 determines the invariant entropy H(Z) using the formula Math. 3 with L = 1 (a neighbor) (step A1). In this same step A1, the estimated entropies H(X, Z), H(Y, Z) and H(X, Y, Z) are determined in this step A1.
[0198] In practice, in this example, the estimation algorithm also determines in this step (in addition to the previous example), the following estimated entropies: With X the first series of abundance values and Y the second series of abundance values and Z the series of conditional variables. For this, the processing unit 120 uses the equation Math. 3 as defined above to determine each estimated entropy H. The distance formula used in the kNN algorithm is the Euclidean distance defined in the formula Math. 11 .
[0199] Subsequently, the processing unit determines the conditional mutual information by applying the formula Math. 20.
[0200] A second embodiment of step 1041 will now be described.
[0201] Mutual information between three variables
[0202] According to this embodiment, mutual information is conditional mutual information between three variables, called information score, and which is defined by the following formula: [Math. 22] / (X, Y,Z) = H(X) + H (Y) + H (Z) - H(X, Y) - H(X,Z) — H(Y,Z) + H(X, Y,Z)
[0203] In practice, the conditional variable Z is also a series of abundance values, which can be, like the second example described above, a series of TIC values (allows to free oneself from variations internal to the system) or a series of abundance values of a third ion (allows to highlight relationships related to this third ion) selected from the plurality of mass spectra (in a similar manner to the ion associated with the first and second series of abundance values).
[0204] The conditional mutual information between three variables is determined using a process similar to that explained above for the example of conditional mutual information.
[0205] Thus, the processing unit uses the kNN algorithm to estimate the terms H(X), H(Y), H(Z), H(X, Y) H(X, Z), H(Y, Z), and H(X, Y, Z) of the equation Math. 22. Similarly, each estimated entropy term is an invariant entropy as explained above.
[0206] In practice, the step 1041 of calculating the mutual information between three variables comprises steps A1 to A2 as described above.
[0207] The processing unit 120 estimates the invariant entropy of each term H(X), H(Y), H(Z), H(X, Y) H(X, Z), H(Y, Z), and H(X, Y, Z) of the equation Math. 22, in a manner similar to the first and second example described above.
[0208] Subsequently, the processing unit 120 determines the conditional mutual information by applying the formula Math. 22.
[0209] As explained above, the matching score can also be based on a correlation between the first series of abundance values and the second series of abundance values. A variant of the calculation step 1040 will now be described in which the matching score is based on a correlation.
[0210] Correlation
[0211] According to a variant of the method 1000, the correspondence score may be based on a statistical method. In practice, in this embodiment, the method 1000 is based on a correlation coefficient calculation but is devoid of entropy calculation. In this embodiment, the processing unit 120 can calculate the correlation coefficient r xy (in a step of calculating a correlation coefficient 1044) on each pair of abundance series from the covariance according to the following formula: with Cov(X, Y) denoting the covariance of the variables X (corresponding for example to the first series of abundance values) and Y (corresponding for example to the second series of abundance values), s x corresponding to the standard deviation of the variable X and s y corresponding to the standard deviation of the variable Y. Typically, the standard deviation of the variable considered corresponds to the standard deviation determined from all the abundance values of the abundance series considered.
[0212] In practice, according to this embodiment, the processing unit 120 calculates, for each pair of series of abundance values, the covariance between the first series of abundance values and the second series of abundance values according to the formula for the covariance of the random variable defined as follows: [Math. 24] Cov(X, Y) = E(XY) - EX) *E(Y) with E() corresponding to the mathematical expectation of the given variable, here the mathematical expectation of the set of abundance values of the abundance series considered.
[0213] Then, the processing unit 120 determines a standard deviation of the variable X and a standard deviation of the variable Y, the standard deviation considered being defined according to the following formula: [Math. 25] with X the variable considered (here the first series of abundance values or second series of abundance values), E(X) the expectation of the random variable considered, Var(X) the variance of the random variable considered.
[0214] Thus, in this embodiment, the processing unit 120 calculates for each pair of series of abundance values, the covariance as defined in the equation Math. 24 as well as the standard deviation s x associated with the first abundance series and the standard deviation sy associated with the second series of abundance values. It then determines the correlation coefficient Cor xy for each pair of abundance value series.
[0215] As the calculation step 1040 is performed iteratively by repeating the calculation step on each pair of abundance value series or in parallel on all pairs of abundance series, a multitude of correlations Cor xy are provided at the end of step 1040.
[0216] Similar to the embodiment using a correspondence score based on the calculation of mutual information, the set of correlation coefficients calculated on the set of pairs of series of abundance values are grouped (step 1042) on a two-dimensional matrix (corresponding to a score map) of a shape similar to that obtained in the context of mutual information. Thus, with respect to Figure 12, the matrix is made of correlation coefficients Cor xy The scale of such a scorecard is, however, different from the example with mutual information, since the calculated correlation coefficients can be negative or positive (unlike mutual information where the score values are positive).
[0217] Furthermore, although not necessary, it is also possible to determine correlation coefficients from series of normalized abundance values. For example, the normalization coefficient can correspond to the invariance parameter of the series under consideration and calculated in a similar way to the example of mutual information.
[0218] A second embodiment of step 1044 will now be described.
[0219] Partial correlation
[0220] According to another variant of the method 1000, the processing unit 120 can calculate the partial correlation coefficient Cor xyzbetween two variables X, Y and a third variable Z. This third variable is similar to the variable Z used in the calculation of mutual information. In other words, the variable Z is a series of TIC values or a series of abundance values of another ion, different from the ion associated with the first and second series of abundance values. It can also be a series made of constant values or a series of values relating to the intensity of a laser beam, a gas pressure, etc. Using a third variable Z makes it possible to avoid a quantity affecting the measurements.
[0221] In this case, the step of calculating the correlation coefficient 1044 can be implemented as follows.
[0222] In practice, the partial correlation coefficient of each pair of abundance value series is defined according to the following equation: with Cor(), the correlation coefficient between the variables considered.
[0223] For this purpose, the processing unit 120 calculates for each pair of series of abundance values, the correlation coefficients Cor yx , Cor xz and Cor yz as previously determined using the formula Math. 23 then applies the formula Math. 26 to determine the partial correlation coefficient between three variables.
[0224] The matching score obtained by this embodiment is more accurate than the matching score of the embodiment based on a calculation of the correlation coefficient between two variables as explained above.
[0225] Typically, the numerator of equation 23 is equivalent to a standard deviation normalization. Such standard deviation normalization in the correlation calculation has the same effect as the invariant measure normalization in the entropy calculation.
[0226] Figure 13 illustrates a score map obtained from a correspondence score based on a correlation coefficient (here partial correlation coefficient). The results displayed correspond to the ionized sample comprising Melittin and substance P. The correspondence scores were determined from a series of abundance values obtained on an ionized sample, the series of abundance values comprising non-integrated values (i.e. according to the first variant of the determination step 1020). In this figure, the values of the m / z ratios of the ions of the first series of abundance values of the pairs are illustrated at the column level (on the abscissa) and the values of the m / z ratios of the ions of the second series of abundance values of the pairs are illustrated at the row level (on the ordinate).
[0227] It can be seen that the resulting score map is a symmetric matrix. In addition, it can be seen that more pairs of abundance value series were used, so the implementation time of the 1000 method may be longer. Matching scores using this method can be positive or negative.
[0228] Figure 14 illustrates a match score map obtained with the previously illustrated series of Kalman-filtered abundance values (i.e., obtained from the mixture of Melittin and substance P). It can be seen that the score map exhibits better contrast. Contrast can be assessed by taking the average of the match scores of all pixels in the match map. In a preferred variant, contrast can be assessed by calculating a variance over all match scores (i.e., all pixels) in the score map. In In this case, the higher the variance, the greater the contrast.
[0229] A combination of the statistical and probabilistic method for determining the match score calculated in calculation step 1040 will now be described.
[0230] As explained above, the matching score can also be based on a correlation between the first set of abundance values and the second set of abundance values and on mutual information between the first set of abundance values and the second set of abundance values.
[0231] Combination of the statistical method and the probabilistic method described above.
[0232] In this embodiment, the matching score is based on the probabilistic method based on an entropy calculation applied to the first series of abundance values and the second series of abundance values and on the statistical method based on a calculation of a correlation between the first series of abundance values and the second series of abundance values.
[0233] Thus, in practice, the processing unit 120 determines, for each pair of series of abundance values, two preliminary correspondence scores by calculating for each pair of series of abundance values: - at least one of the mutual information (step 1041) described above l(X, Y), l(X, Y|Z), l(X, Y, Z), i.e. by calculating the mutual information defined by one of the following formulas: Math. 20, Math. 21, Math. 22; and - at least one of the correlation scores (step 1044) described above Cor xy , Cor xyz, that is, by calculating the correlation coefficient defined by one of the following formulas: Math. 23, Math. 26.
[0234] Preferably, in this embodiment, the processing unit 120 combines the mutual information between three variables l(X, Y, Z) with the partial correlation coefficient Cor xyz .
[0235] In this case, the processing unit 120 can calculate a mixed coefficient Sa(X, Y, Z) (step 1045) on each pair of series of abundance values which is defined according to the following formula: [Math. 27] Sa(X, Y,Z) = I(X, Y,Z) * sign Cor XY Z ) with sign(Cor xyz ) corresponding to the sign of the partial correlation coefficient and I (X, Y, Z) corresponding to the mutual information between three variables.
[0236] In a variant of the method 1000, the processing unit 120 may weight at least one of the following terms, namely the mutual information between three variables, the partial correlation coefficient, in order to obtain a more accurate matching score. In this case, the processing unit 120 calculates a weighted mixed coefficient Sa p (X, Y, Z) on each pair of abundance series can be defined according to the following formula: [Math. 28] Its p (X, Y, Z = alpha * Ï(X, Y, Z * beta * Cor XYZ )
[0237] with alpha a first weighting coefficient greater than or equal to 1 and beta a second positive weighting coefficient greater than or equal to 1.
[0238] In practice, the first weighting coefficient depends on the correspondence score based on the mutual information considered and the second weighting coefficient depends on the information score based on the correlation coefficient considered.
[0239] For example, the matching score can be calculated according to the formula Math. 28 in which the conditional mutual information (i.e. the matching score based on the conditional mutual information) is multiplied with the partial correlation coefficient and in which the first weight coefficient is equal to 1 and the second term of the equation Math. 28 is written as sign(beta*Cor xyz -0.1) with beta, second correlation coefficient equal to 1. In this example, the value 0.1 is selected because it allows to reduce a correlation intrinsic to the fragmentation process from which the TIC term cannot free itself.
[0240] Alternatively, the equation Math. 28 can also be written: Its p (X, Y, Z) = h(f(I(X, Y,Z)), g(Cor XYZ )) with f() and g() two functions which apply a transformation on each coefficient (Here l(X, Y, Z) and Cor xyz . These functions can be for example, a square function, an affine function, a sign function, etc. The function h(a,b) is a function which takes two variables and applies a transformation such as addition, multiplication.
[0241] The other steps of the method 1000 will now be described after obtaining one or more correlation maps from the set of correspondence scores described above.
[0242] In certain cases, the correlation map(s) obtained in step 1042 are not usable because they have too low a contrast. In this case, the method 1000 optionally comprises a step 1050 of selecting the correlation maps as a function of a contrast of the score map considered.
[0243] Such a step makes it possible to eliminate unusable correlation maps so as to limit errors in the use of the correlation maps in the remainder of the method 1000.
[0244] In practice, this selection step 1050 can be directly calculated in the calculation step 1040 after having determined the score card associated with a given correspondence score. In this case, the processing unit 120 calculates, after having determined the score card, a contrast of the given score card and then compares this contrast to a given threshold to keep or eliminate this correspondence card. For example, the threshold can be a value greater than or equal to 0.30, for example 0.5.
[0245] Typically, in one example, the score map resulting from the two-variable mutual information-based matching score may have a contrast of less than 0.3. Therefore, following the selection step 1050, this score map may be deleted from the memory of the processing unit or the storage module 140.
[0246] If only one type of score map is calculated in the calculation step, only one correspondence map is obtained following step 1040. In this case, if this map has a contrast lower than the threshold, the method can repeat the calculation step by determining correspondence scores based on another calculation method (i.e., moving, for example, from mutual information between two variables to mutual information between three variables, or moving to a score based on a correlation coefficient or using a mixed correlation score, etc.).
[0247] The correlation map(s) passing the requirement of step 1050 are stored in the memory of the processing unit and / or the storage module 140. Thus, in the remainder of the method 1000, only the correlation map(s) stored following the selection step 1050 are used in the remainder of the method 1000. Of course, if the method 1000 does not include a selection step 1050 as described above, all the correlation maps determined in the calculation step 1040 are used in the method 1000.
[0248] In the illustrated example, following the selection step 1050, the method 1000 comprises a determination step 1060, as a function of a mass to charge ratio, of at least one precursor ion and the associated fragment ions. For this, the method 1000 uses a classification method which is applied to the correlation map(s) calculated at the end of the calculation step 1040.
[0249] Thus, in this step 1060, the processing unit 120 applies to each score card a grouping method to determine at least one precursor ion as well as the fragment ions associated with this precursor ion. This precursor ion is determined according to an m / z ratio.
[0250] In the case where several precursor ions are determined in this step, each determined precursor ion is associated with its fragment ions.
[0251] Optionally, before applying the clustering method to the correlation maps, the method 1000 may comprise a step 1062 of processing the correlation maps. In practice, this processing step 1062 is carried out in the determination step 1060 but before the application of the classification method.
[0252] Typically, processing step 1062 uses at least one of the following processes: - an algorithm for processing the diagonal of the given scorecard. In this case, the processing unit 120 can assign the same score value (for example an integer) to each element present on the diagonal of the score map. For example, all the elements on the diagonal can be set to zero so that these elements do not significantly influence the classification algorithm. In the case of a score map based on correlation coefficients, the integer can be negative; - an algorithm for standardizing the correlation maps. In this case, the standardization algorithm can be applied to each line of the correlation maps. This standardization algorithm uses an average of a considered line as well as a standard deviation. In practice, the processing unit 120 calculates the average of each line and the standard deviation of each line. Then, for each element of a considered line, the processing unit subtracts from this considered element the average value of the considered line and then divides this value by the standard deviation of the considered line. Thus, with this processing, each line of the score map has a similar weight in the processed score map; - an application of a square function in order to accentuate the gap between the small and large coefficients (correspondence scores) of the scorecard considered.
[0253] In this example, the method 1000 then uses a clustering (or classification) method on the processed correlation maps (step 1064).
[0254] Typically, the clustering algorithm uses at least one unsupervised learning algorithm which may include at least one of the following algorithms: - a hierarchical ascending classification algorithm, for example using a Ward method (as an aggregation method) on the variance, and / or the Complete Link method, - a k-nearest neighbors algorithm, - a k-means algorithm, - a dimensionality reduction algorithm, also known in English as “t-distributed stochastic neighbor embedding algorithm” (t-SNE).
[0255] In the clustering method, each classification algorithm takes as input a score map. Thus, if several correlation maps are calculated in step 1040 or retained following the selection step, the processing unit 120 will implement several clustering algorithms.
[0256] The clustering method may also define other elements in the input data, including at least one of the following parameters: the method used (here for example the Ward method in the case of the unsupervised learning algorithm); the distance used in the (aggregation) method used; etc. These data depend on the type of algorithm used and are therefore known to those skilled in the art.
[0257] An embodiment of the determination step 1060 will now be described with the aid of FIG. 15, FIG. 16, FIG. 17, FIG. 18 and FIG. 19. used in process 1000.
[0258] In the example of method 1000, the determination step 1060 uses a hierarchical clustering algorithm. In practice, in this example, the algorithm used uses the Ward method and takes a Euclidean distance by default. The Ward method is more suitable because it minimizes the variance by cluster fusion and is recommended for continuous values, unlike the other methods mentioned above, such as the Complete link method which takes into account the maximum values and not a global set of the score map considered.
[0259] Figure 15 illustrates a dendrogram obtained by the hierarchical classification method used and obtained from the data calculated on the ionized sample from a mixture of Melittin and substance P. Here, the dendrogram includes on the abscissa the different labels corresponding in this example to the m / z ratios of precursor ions and fragment ions and on the ordinate a distance between different classes. In this example, each class is characterized by a horizontal line on the dendrogram. Vertical branches / links connect the classes to illustrate the links between the different classes.
[0260] As illustrated in Figure 15, the dendrogram comprises a main class C1 which is divided into two main vertical branches C11, C12, each main vertical branch being again connected to a new class C2, which allows to obtain a tree structure in which the vertical branches illustrate the links / proximities between the different classes connected to these branches.
[0261] Thus, we can deduce from this dendrogram that the initial mixture includes two precursor ions (visible by the two links of the main branch). Each main branch is connected to a new class (called secondary class) which is connected to secondary links which are themselves connected to a new class. Each branch connecting two classes positioned at two different distances illustrates the distance separating these two classes.
[0262] Thus, we visualize by the two main branches connected to the main class two large clusters / groupings of precursor ions and by the secondary branches, the fragment ions of this precursor ion. We can deduce that each secondary branch connected to a secondary class corresponds to a fragment of the precursor ion and that the arborescences (other branches connected to classes positioned at a distance lower than that of the fragment considered) connected to this secondary class define the chemical compounds of this fragment. Therefore, the final vertical branches (the lowest branches, connected to no lower class) illustrate the m / z ratios of each chemical compound of the fragment ions of the precursor ion considered. Thus, thanks to this dendrogram, we can directly visualize the m / z ratios of each chemical compound of the fragment ions of the precursor ion of the class.
[0263] Thus, after having obtained this dendrogram, the processing unit 120 determines the precursor ions and the fragments associated with each precursor ion by exploiting the different branches and classes of the dendrograms, the distances separating the classes as well as the m / z ratios given at the distance equal to zero (zero ordinate).
[0264] The method 1000 then comprises a step 1070 of reconstructing the individual mass spectrum of each precursor ion determined and the associated fragments following the determination step 1060. In practice, the processing unit 120 uses the data extracted from the classification method to reconstruct the mass spectrum of each precursor ion, in particular the m / z ratios grouped in the same class. Thus, for each precursor ion determined in the previous step, the processing unit 120 recovers each m / z ratio positioned at the minimum distance (equal to zero), and associates an intensity (i.e. an abundance value) with this ratio. This intensity can be calculated from the series of abundance values corresponding to this ratio, for example by calculating the average abundance value of this series of abundance values) or directly from the average spectrum by selecting the m / z ratio concerned on the average spectrum.The processing unit can therefore construct a peak on the mass spectrum of the precursor ion at each determined m / z ratio and assign it a corresponding intensity.
[0265] Such a reconstructed mass spectrum corresponds to a tandem mass spectrum of the determined precursor ion. Thus, with the method 1000, it is possible to reconstruct the tandem mass spectrum of each determined precursor ion in the sample used using mass spectra acquired from a simple mass spectroscopy device, i.e. without using a tandem mass spectroscopy device.
[0266] Optionally in the example considered, the method 1000 comprises a verification of the reconstruction 1080 comprising a comparison between each individual reconstructed mass spectrum and the average mass spectrum obtained from the plurality of mass spectra (as described above). Such an action makes it possible to verify that the fragments associated with this or these precursor ions exist, which improves the reliability and accuracy of the reconstruction step 1070.
[0267] Figure 16 and Figure 17 each illustrate an individual mass spectrum of one of the two precursor ions determined in the previous step.
[0268] This figure 16 includes a first mass spectrum 701 of a first determined precursor ion (here that corresponding to the Melittin ion), and figure 17 a second mass spectrum 702 of a second determined precursor ion (here that corresponding to the ion of substance P).
[0269] Optionally, these individual mass spectra 701, 702 may compared with the average mass spectrum shown in Figure 5 to verify that the combination of the individual mass spectra 701 and 702 includes all the peaks of the previously determined average mass spectrum. This is the case in this example. This can be done following the reconstruction, in reconstruction step 1070 (step 1071).
[0270] Alternatively, the determining step 1060 may use another clustering algorithm (called the first classification algorithm). Similar to the previous reconstruction step, this other classification algorithm is applied individually to each correlation map.
[0271] In this example, the processing unit 120 uses the T-SNE algorithm as the classification algorithm. Typically, the input data used by this algorithm are similar to those used by the hierarchical classification algorithm (first classification algorithm) described above.
[0272] Figure 18 illustrates a two-dimensional map provided by the T-SNE algorithm. This map includes a multitude of points represented on the map (a plane). Each point is associated with an m / z ratio.
[0273] This algorithm is often used when the ionized sample includes a number of precursors greater than or equal to 2 and / or when there is a high probability that a fragment ion is common to at least two precursor ions in the ionized mixture. The output data is a list of coordinate pairs (x,y) for each label (ion). These coordinates allow for a spatial visualization of the different groups highlighted.
[0274] Thus, the T-SNE algorithm allows to obtain a spatial visualization of the precursor ions and fragment ions of the ionized mixture. It is also commonly combined with another clustering algorithm which in this case takes as input the map obtained by the t-SNE algorithm. In one example, the t-SNE algorithm is combined with a k-means algorithm. Figure 19 illustrates the use of these two algorithms successively. We see that on the map obtained by the t-SNR algorithm, two groups G1 and G2 (visible by the dotted lines) each include points of the map associated with an m / z ratio. All the points included in the G1 group correspond to the fragment ions of one of the precursor ions of the ionized mixture (here the precursor ion of Melittin) and all the points included in the G2 group correspond to the fragment ions of another precursor ion of the ionized mixture (here the precursor ion of substance P).These two groups G1 and G2 can also be separated by a straight line D1 determined by the k-means algorithm. It is then possible to associate with each point (a given m / z ratio) an abundance value (i.e. intensity). This abundance value is determined as previously (in the case of the hierarchical classification algorithm) from the average mass spectrum or. of the abundance value series associated with this m / z ratio).
[0275] The method 1000 may also combine several clustering methods to improve the accuracy of the reconstruction. For example, the method 1000 may use a hierarchical ascending clustering algorithm and then a t-SNE algorithm combined with a k-means algorithm. Combining these methods may allow verification and / or completion of the reconstructed mass spectrum(s), particularly when several precursor ions include fragment ions that are close in m / z ratio.
[0276] After reconstructing the individual mass spectrum of each detected precursor ion, the method may optionally comprise a step 1090 of analyzing the individual mass spectrum of the determined precursor ion and the associated fragment ions to characterize the chemical composition of the determined precursor ion and the chemical composition of the fragment ions associated with the precursor ion considered. Indeed, after the reconstruction step (and after the verification step if present in the method 1000), the chemical composition of the precursor ion and of the different fragments of the precursor ion are unknown (the reconstruction step provides a spectrum with m / z ratios and abundance values associated with each m / z ratio). For this purpose, the analysis step 1090 makes it possible to find the chemical composition of the fragments and of the precursor ion of each determined individual mass spectrum.This can typically be accomplished by known techniques, using a database comprising (known) reference mass spectra and comparing each mass spectrum determined by the method 1000 to the reference mass spectra.
Claims
CLAIMS 1. Method for reconstructing a mass spectrum, said method comprising the following steps: - acquisition (1010) of a plurality of mass spectra measured on an ionized sample comprising at least one precursor ion and fragment ions, - determining (1020) a plurality of series of abundance values of the ionized sample from the plurality of mass spectra, each series of abundance values of the ionized sample being determined from the plurality of mass spectra, each series of abundance values comprising N abundance values where N is an integer greater than or equal to two; - arranging (1030) the plurality of abundance value series into pairs of abundance value series, each pair of abundance value series comprising a first abundance value series and a second abundance value series, - for each pair of abundance value series, calculating (1040) a correspondence score between the first abundance value series and the second abundance value series, so as to obtain correspondence scores for all pairs of abundance value series, the correspondence scores providing a correspondence score map between all pairs of abundance value series, - determination (1060), as a function of a mass to charge ratio, of at least one precursor ion and the fragment ions associated with each determined precursor ion, by applying a grouping method to the calculated correspondence score card, - reconstruction (1070) of an individual mass spectrum of each determined precursor ion and of each of the associated fragment ions from the correspondence score card.
2. The method of claim 1, wherein the step of determining a series of abundance values comprises determining a series of ranges of mass-to-charge ratio values for the plurality of measured mass spectra of size NxM with M an integer greater than 1, and for each mass spectrum of the plurality of mass spectra, a step of integrating abundance values belonging to the same range of mass-to-charge ratio values of the mass spectrum considered, to provide, for each range of mass-to-charge ratio values of the spectrum considered, an integrated abundance value associated with a mass-to-charge ratio, the integrated abundance values over the series of ranges of mass-to-charge ratio values forming a series of abundance values of the ionized sample.
3. A method according to any one of claims 1 to 2, wherein the match score is based on at least one of the following: - mutual information between the first series of abundance values and the second series of abundance values, said mutual information being based on an entropy of the first series of abundance values and on an entropy of the second series of abundance values; and / or - a correlation between the first series of abundance values and the second series of abundance values.
4. Method according to claim 3, in which, for the mutual information, the entropy of the first series of abundance values is an entropy normalized by a first invariance parameter as a function of a first median value determined on this first series of abundance values and the entropy of the second series of abundance values is an entropy normalized by a second invariance parameter as a function of a second median value determined on this second series of abundance values.
5. Method according to any one of claims 3 to 4, wherein, the correspondence score being based on mutual information, the method comprises a step of estimating the entropy of the first series of abundance values and a step of estimating the entropy of the second series of abundance values using a k nearest neighbors algorithm.
6. The method of claim 5, wherein the k-nearest neighbors algorithm is applied to a single neighbor.
7. Method according to any one of claims 3 to 6, wherein when at least two correspondence scores based on two different elements are calculated in the calculating step, several score cards are provided following the calculating step, said method further comprising, for each score card, a step of filtering (1050) said score card according to a contrast of the score card considered.
8. Method according to any one of claims 1 to 7, further comprising a step of processing (1062) the score card before applying the grouping method, said processing step using at least one of the following algorithms: - standardization algorithm, - algorithm for processing diagonal elements of the scorecard, - application of a square function.
9. A method according to any one of claims 1 to 8, wherein the score map is a two-dimensional matrix, the clustering method determining the at least one precursor ion and its associated fragment ions from the matching scores located outside a diagonal of the score map.
10. Method according to any one of claims 1 to 9, wherein the clustering method uses at least one unsupervised learning algorithm.
11. Method according to any one of claims 1 to 10, wherein the grouping method is based on at least one of the following algorithms: - a hierarchical ascending classification algorithm, - a k-nearest neighbors algorithm, - a k-means algorithm, - a stochastically distributed neighbor incorporation algorithm.
12. Method according to any one of claims 1 to 11, further comprising, after the reconstruction step (1070), a step of analyzing (1090) the individual mass spectrum of the determined precursor ion and the associated fragment ions to characterize the chemical composition of the determined precursor ion and the chemical composition of the fragment ions associated with the precursor ion considered.
13. Method according to any one of claims 1 to 12, in which the method comprises a step of processing (1015) the acquired mass spectra comprising at least one of the following calculations: - filtering of the mass spectrum considered by a Gaussian or normal kernel, - a normalization of the mass spectrum considered, - a smoothing of the mass spectrum considered, - a return to the baseline.
14. Method according to any one of claims 1 to 13, in which the method comprises a step of processing (1025) the series of abundance values comprising at least one of the following calculations: - filtering the series of abundance values given by a Kalman filter; - filtering the series of abundance values given by a Fourier filter, - a filtering by moving average of the given series of abundance values.
15. Method according to any one of claims 1 to 14 in which an external variation by chromatographic separation, separation by ion mobility and / or by discriminating transmission of an ion optic is applied during the acquisition step so as to modify the abundance of the ions and wherein the number of mass spectra of the plurality of mass spectra is less than or equal to 10.