Method of reconstructing a mass spectrum
The method for reconstructing individual mass spectra from a series of mass spectra addresses the limitations of existing mass spectrometry techniques by enabling the identification of precursor ions and their fragments without tandem mass spectrometry, improving analysis speed and reducing instrumental complexity.
Patent Information
- Application Number
- FR2023005634
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-06-05
AI Technical Summary
Existing mass spectrometry techniques, such as tandem mass spectrometry, are limited by instrumental constraints, requiring sophisticated and costly instruments, and are prone to errors due to the inability to differentiate ions with identical m/z ratios and the need for prior knowledge of the sample.
A method for reconstructing an individual mass spectrum of a precursor ion from a series of mass spectra using a grouping method based on calculated correspondence scores, which allows for the identification of precursor ions and their associated fragment ions without the need for tandem mass spectrometry devices.
This method enables the reconstruction of individual mass spectra from simple mass spectra, reducing the complexity and cost of instrumentation, improving analysis speed, and allowing for the identification of isobaric species, thereby overcoming the limitations of existing mass spectrometry techniques.
Smart Images

Figure 00000046_0000 
Figure 00000046_0001 
Figure 00000047_0000
Abstract
Description
Title of the invention: Method for reconstructing a spectrum of mass Technical field of the invention
[0001] The present invention generally relates to a method of reconstructing a mass spectrum.
[0002] It relates in particular to a method for reconstructing an individual mass spectrum of at least one precursor ion from the exploitation of a series of mass spectra.
[0003] It also relates to a device implementing such a method. State of the art
[0004] Mass spectrometry (MS) is an analytical technique for detecting ions from a sample and analyzing these ions based on their ratio (m / z), where m represents the mass of an ion and z its electrical charge. Mass spectrometry is used in many applications to analyze, identify and characterize the chemical structure of ionized molecules.
[0005] A mass spectrometer generally comprises an ionization source to form the ions from a sample to be analyzed, an analyzer that separates the ions according to their m / z ratio and a detector. A mass spectrum is obtained by recording the abundance of the ions as a function of their mass to charge ratio (m / z). However, simple mass spectrometry only allows the measurement of the m / z ratio and does not allow the differentiation of ions having identical m / z ratios, in particular because of the resolution limit of the spectrometer used. In addition, the measurement of the m / z ratio does not provide information on the composition or structure of the ions.
[0006] It is known to use tandem mass spectroscopy to study ions of interest in a sample more precisely. Such a technique makes it possible to provide an individual mass spectrum for a given precursor ion after activation. For this purpose, tandem mass spectrometry comprises, after ionization of the molecules in the sample, a step of selection by a mass filter of an ion of interest. This selected ion is then subjected to an activation step which consists of increasing its internal energy, which can lead to its fragmentation. The ions produced from the fragmentation of the selected ion (also called precursor ion or parent ion) are then analyzed in one or more other mass spectrometry step(s) on the fragment ions thus generated, in which the mass analysis steps can be separated spatially or temporally depending on the instrument used.
[0007] For example, tandem mass spectrometry can be implemented by isolating an ion of interest in an ion trap by providing it with sufficient internal energy through collision with a gas so that it fragments. Detection of the products of this fragmentation can provide information on the nature and / or structure of the precursor ion. Tandem mass spectrometry is the basis for applications of mass spectrometry in structural analysis and in particular the sequencing of proteins and other biopolymers (such as sugars or nucleic acids).
[0008] Tandem mass spectrometry is functional but requires more sophisticated instruments generally comprising two mass analyzers coupled in series which adds complexity to the spectrometers and also increases their costs and maintenance.
[0009] Tandem mass spectrometry is also limited by instrumental constraints, in particular the fineness of selection which in certain cases is not sufficient to separate two ions of interest with very close masses (very close m / z ratio). In this case, two different ions but with close m / z ratios will be selected and activated simultaneously. The tandem mass spectrum obtained will therefore be the sum of the tandem mass spectra of the two ions selected simultaneously, which complicates the interpretation of the measurements.
[0010] Tandem mass spectrometry can also be slower in terms of analysis speed compared to other methods due to the sequential acquisition of data. This is because the ions of interest are selected and fragmented one after the other.
[0011] Furthermore, in tandem mass spectrometry, a criterion for choosing the ions to be analyzed in a sample is necessary. In this case, it is often necessary to have a list of ions of interest that one wishes to fragment, which amounts to having a priori knowledge of the sample. In a variant, it is also possible to choose to fragment all the ions whose intensity exceeds a certain threshold, again set according to the sample. Prior knowledge of the sample is therefore necessary.
[0012] The invention aims to remedy at least one of the aforementioned drawbacks. Presentation of the invention
[0013] In order to overcome the aforementioned drawback of the state of the art, the present invention proposes a method for reconstructing a mass spectrum, said method comprising the following steps: - acquisition of a plurality of mass spectra measured on an ionized sample comprising at least one precursor ion and fragment ions, the plurality of spectra mass spectra comprising at least N mass spectra, - determining a plurality of series of abundance values of the ionized sample from the plurality of mass spectra, each series of abundance values of the ionized sample being determined from the plurality of mass spectra, each series of abundance values comprising N abundance values where N is an integer greater than or equal to two; - arranging the plurality of abundance value series into pairs of abundance value series, each pair of abundance value series comprising a first abundance value series and a second abundance value series, - for each pair of abundance value series, calculating a match score between the first abundance value series and the second abundance value series, so as to obtain match scores for all pairs of abundance value series, the match scores providing a match score map between all pairs of abundance value series, - determination, based on a mass to charge ratio, of at least one precursor ion and the fragment ions associated with each determined precursor ion, by applying a grouping method to the calculated correspondence score card, - reconstruction of an individual mass spectrum of each determined precursor ion and each of the associated fragment ions from the correspondence score card.
[0014] Thus, thanks to the invention, in particular to the different steps of the method according to the invention, it is possible to go back to the mass spectrum of an individual ion present in a mixture (i.e. a sample which may contain several ions and fragments) without necessarily using a tandem mass spectroscopy device. As a result, the present invention can carry out tandem mass spectrometry analyses from a series of simple mass spectra having been acquired by a mass spectroscopy device which is not provided with a second mass filter (here an ion trap) but in which either the desorption / ionization step causes the fragmentation of the species, or by the implementation of an in-source activation step by collision-induced dissociation (or "in-source Collision-induced dissociation also known in English as in-source CID") making it possible to activate all of the ionized species.As a result, the process is easier to implement and less expensive compared to a process requiring working from tandem mass spectrometers.
[0015] Furthermore, although this method works from a series of mass spectra, the calculation step and the use of a classification algorithm makes it possible to precisely group at least one precursor ion and the fragments associated with this precursor ion. Such a method is therefore functional on mass spectra acquired with devices whose selection width is sometimes too large to isolate an ion precisely. The invention can also make it possible to identify the presence of isobaric species in a tandem mass spectrometry experiment. As a result, the method according to the invention also makes it possible to work on all types of acquired mass spectra. In other words, the method makes it possible to overcome the technical limitations of mass spectrum devices, which may be linked to the measurement resolution or to constraints linked to the ionization sources used.
[0016] Other advantageous and non-limiting characteristics of the method according to the invention, taken individually or in all technically possible combinations.
[0017] In one embodiment, the step of determining a series of abundance values comprises determining a series of ranges of mass-to-charge ratio values for the plurality of measured mass spectra of size NxM with M an integer greater than 1, and for each mass spectrum of the plurality of mass spectra, a step of integrating abundance values belonging to the same range of mass-to-charge ratio values of the mass spectrum considered, to provide, for each range of mass-to-charge ratio values of the spectrum considered, an integrated abundance value associated with a mass-to-charge ratio determined in each range of mass-to-charge ratio values, the integrated abundance values over the series of ranges of mass-to-charge ratio values forming a series of abundance values of the ionized sample.
[0018] In one embodiment, the match score is based on at least one of the following: - mutual information between the first series of abundance values and the second series of abundance values, said mutual information being based on an entropy of the first series of abundance values and on an entropy of the second series of abundance values; and / or - a correlation between the first series of abundance values and the second series of abundance values.
[0019] In one embodiment, for the mutual information, the entropy of the first series of abundance values is an entropy normalized by a first invariance parameter as a function of a first median value determined on this first series of abundance values and the entropy of the second series of abundance values is an entropy normalized by a second invariance parameter as a function of a second median value determined on this second series of abundance values.
[0020] In one embodiment, the correspondence score being based on mutual information, the method comprises a step of estimating the entropy of the first series of abundance values and a step of estimating the entropy of the second set of abundance values using a k-nearest neighbors algorithm.
[0021] In one embodiment, when at least two correspondence scores based on two different elements are calculated in the calculating step, several score cards are provided following the calculating step, said method further comprising, for each score card, a step of filtering said score card according to a contrast of the considered score card.
[0022] In one embodiment, the method further comprises a step of processing the score card before applying the clustering method, said processing step using at least one of the following algorithms: - standardization algorithm, - algorithm for processing diagonal elements of the scorecard, - application of a square function.
[0023] In one embodiment, the score map is a two-dimensional matrix, wherein the clustering method determines the at least one precursor ion and its associated fragment ions from match scores located outside a diagonal of the score map.
[0024] In one embodiment, the clustering method uses at least one unsupervised learning algorithm.
[0025] In one embodiment, the clustering method is based on at least one of the following algorithms: - a hierarchical ascending classification algorithm, - a k-nearest neighbors algorithm, - a k-means algorithm, - a stochastically distributed neighbor incorporation algorithm.
[0026] In one embodiment, the method further comprises, after the reconstruction step, a step of analyzing the individual mass spectrum of the determined precursor ion and the associated fragment ions to characterize the chemical composition of the determined precursor ion and the chemical composition of the fragment ions associated with the precursor ion considered.
[0027] In one embodiment, the method comprises a step of processing the acquired mass spectra comprising at least one of the following calculations: - filtering of the mass spectrum considered by a Gaussian or normal kernel, - a normalization of the mass spectrum considered, - a smoothing of the mass spectrum considered, - a return to the baseline.
[0028] In one embodiment, the method comprises a step of processing the series of abundance values comprising at least one of the following calculations: - filtering the series of abundance values given by a Kalman filter; - filtering of the given series of abundance values by a Fourier filter, - filtering by moving average of the given series of abundance values.
[0029] In one embodiment, the method comprises a compressed acquisition step to reduce the number of acquisitions.
[0030] In one embodiment, the plurality of mass spectra comprises at least 1000 mass spectra, preferably between 2000 and 8000 mass spectra.
[0031] Of course, the various features, variants and embodiments of the invention may be combined with each other in various combinations to the extent that they are not incompatible or mutually exclusive. Detailed description of the invention
[0032] The description which follows with reference to the appended drawings, given as non-limiting examples, will make it clear what the invention consists of and how it can be implemented.
[0033] In the attached drawings:
[0034] [Fig. 1] is a schematic representation of a device for reconstructing an individual mass spectrum of at least one precursor ion according to the present disclosure;
[0035] [Fig.2] is a schematic representation of a mass spectroscopy device used in the device according to the present disclosure;
[0036] [Fig.3] is a schematic representation of an example of a method according to the present disclosure;
[0037] [Fig.4] is a schematic representation of an embodiment of an acquisition step of a method according to the present disclosure.
[0038] [Fig.5] is an example of an average mass spectrum determined in an acquisition step of a method according to the present disclosure;
[0039] [Fig.6] is a schematic representation of ranges of m / z ratio values determined from an average mass spectrum obtained in an acquisition step of a method according to the present disclosure;
[0040] [Fig.7] is a representation of an example of a first series of abundance values and a second series of abundance values obtained in an arranging step of a method according to the present disclosure;
[0041] [Fig.8] is a representation of a first series of abundance values and a second series of abundance values illustrated in [Fig.7] for which the abundance values of each of the series have been filtered by a Kalman filter in a method according to the present disclosure;
[0042] [Fig.9] is a schematic representation of a determination step included in a step of calculating a matching score in a method according to the present disclosure, said score being based on a calculation of mutual information;
[0043] [Fig. 10] is a succession of curves each illustrating a cloud of points of a plurality of mutual information determined from non-invariant entropies;
[0044] [Fig. 11] is a succession of curves each illustrating a cloud of points of a plurality of mutual information determined from invariant entropies according to the present disclosure;
[0045] [Fig. 12] is a score map obtained at a step of calculating a matching score based on mutual information by applying a method according to the present disclosure;
[0046] [Fig. 13] is a score map obtained at a step of calculating a correspondence score based on partial correlation by applying a method according to the present disclosure;
[0047] [Fig. 14] is a score map obtained at a step of calculating a correspondence score based on partial correlation by applying a method according to the present disclosure, the correlation scores being determined from pairs of series of abundance values filtered by a Kalman filter;
[0048] [Fig. 15] is a schematic representation of a dendrogram obtained by a clustering method used in a method according to the present disclosure;
[0049] [Fig. 16] is an exemplary representation of an individual mass spectrum of a precursor ion of an ionized sample, said individual mass spectrum being obtained according to a method according to the present disclosure;
[0050] [Fig. 17] is an exemplary representation of another individual mass spectrum of another precursor ion of ionized sample, said individual mass spectrum being obtained according to a method according to the present disclosure;
[0051] [Fig. 18] is a schematic representation of a two-dimensional map obtained by a clustering method used in a method according to the present disclosure;
[0052] [Fig. 19] is a schematic representation of the two-dimensional map obtained from [Fig. 18] to which another clustering algorithm has been applied.
[0053] There will be described with the aid of [Fig.l] and [Fig.2] a device 100 for reconstructing a mass spectrum of an individual ion present in a sample according to the present disclosure.
[0054] The device illustrated in [Fig.l] comprises a measurement module 110 configured to determine at least one mass spectrum 2, preferably a plurality of mass spectra.
[0055] For this purpose, the measurement module 110 comprises a mass spectrometry device 10 as illustrated in [Fig.2]. The mass spectroscopy device 10 can include a single-stage mass spectrometer device (as opposed to a tandem mass spectrometer).
[0056] In the example illustrated in [Fig.2], the mass spectroscopy device 10 is a single-stage device. It comprises an ionization source 11 for forming ions from the sample 1 to be analyzed and an activation means 12 configured to fragment the ions. In the present disclosure, the chemical composition of the sample is unknown. Therefore, it comprises a mixture comprising one or more ions suspended in a solution. In other words, the sample may comprise different molecular species.
[0057] In the present disclosure, sample 1 is for example a solution containing molecules. This sample 1 is then placed in the gas phase and is then ionized by the ionization source 11.
[0058] The ionization source 11 comprises at least one of the following ionization sources, namely: an electron impact ionization source, an electro-nebulization source, a chemical ionization source, a photoionization source, a laser ionization source, a matrix-assisted laser induced desorption ionization source (known by the acronym MALDI source), a surface-assisted desorption ionization (SELDI) source, an atmospheric pressure MALDI source, a matrix-assisted fast evaporation ionization source, any other method involving a matrix or a surface, any method involving a laser, a plasma desorption ionization source, an atmospheric pressure chemical ionization source or an atmospheric pressure photoionization source, an atmospheric pressure ionization source.
[0059] In a preferred embodiment of the invention, the ionization source 11 comprises an electrospray source. These ions are then activated by the activation means 12 configured to fragment the ionized ions in the gas phase in an in-source activation step (CID). All the ions are activated simultaneously and can produce electrically charged fragments or fragment ions, also called fragments in the remainder of the present disclosure.
[0060] In these different variants, the activation means 12 comprises at least one of the following means: activation by collision, activation by absorption of electromagnetic radiation, thermal activation, etc.
[0061] In another embodiment, the mass spectrometry device 10 is devoid of activation means 12. In this case, the ionization source can also be configured to fragment the ions of the sample. In other words, this type of ionization source 11 provides sufficient energy during ionization to produce fragments. In practice, such an ionization source 11 comprises for example an ionization source 11 by electron impact or an ionization source 11 using laser desorption ionization (LDI) technology.
[0062] In another embodiment (not shown), the mass spectrometry device 10 comprises a first mass filter for selecting a range of m / z ratios of interest, an activation means 12 and a second mass filter for analyzing fragmentation products resulting from the activation. In this embodiment, one or more precursors are selected and activated simultaneously. These ions will each produce one or more fragment ions. The recorded tandem mass spectrum therefore includes all of the fragments generated and possibly the precursor ions. Such a mass spectroscopy device is also called a two-stage mass spectrometry device.
[0063] The ions obtained at the output of the activation means 12 are called ionized samples. They include, in particular, at least one ion, called a precursor ion.
[0064] In the present disclosure, a precursor ion or parent ion refers to an ion that can be decomposed into different fragments. Therefore, a fragment ion (or fragment) is a part or chemical structure of the parent ion.
[0065] The mass spectroscopy device 10 also comprises a mass filter 13 (also called analyzer) which separates at least one precursor ion as well as the different fragments of the ionized sample according to their m / z ratio.
[0066] In the present disclosure, the abundance of each ion can be used to obtain quantitative information on the concentration of the molecular species in the sample. The mass spectroscopy device 10 also comprises a detector 14 configured to detect the ionized species (ionized ions). The mass spectrum 2 is obtained by recording the abundance of the ions as a function of their mass to charge ratio (m / z).
[0067] A mass spectrum 2 thus provides information on the chemical composition of the sample, in particular on the number of species present, typically ions and fragments, and on their respective mass.
[0068] As illustrated in [Fig.2], the mass spectroscopy device lacks a second mass filter necessary for performing a tandem mass spectrometry measurement.
[0069] In the present disclosure, the mass spectrometry device 10 is also configured to acquire a time series of mass spectra 2. By time series, it is meant that the mass spectrometry device 10 successively acquires (over time) mass spectra. In practice, the sample 1 used to acquire these mass spectra 2 is identical for each acquisition of mass spectrum 2 by the mass spectrometry device 10. However, the ionized sample of each acquisition (of mass spectrum) may vary during the acquisitions.
[0070] Optionally, the mass spectrometry device 10 may also comprise a computer (not shown) comprising a memory for storing the mass spectra acquired at the output of the detector.
[0071] The device 100 illustrated in [Fig.l] also comprises a processing unit 120.
[0072] By processing unit 120 is meant a computer or a processor or a central processing unit (CPU) or any electronic device making it possible to implement a succession of commands and / or calculations. Typically, the processing unit 120 comprises a processor, a memory and different input and output interfaces. For this purpose, the processing unit 120 is configured to control the operations and functions performed by the device 100. The input and output interface may be the connection point allowing the device 100 to receive data and send commands or data. The processing unit 120 may be connected to other processing units to exchange information with them.
[0073] The processing unit 120 comprises for this purpose one or more (calculation) modules configured to carry out a succession of commands and / or calculations which will be explained below with [Fig.3]. Typically, the processing unit is configured to implement the method explained below.
[0074] The device 100 may also comprise a communication module 130 configured to exchange data with elements connected to the device 100 and / or elements internal to the device 100. By exchanging data, we mean sending and / or receiving data, the received data corresponding in this example to input data which may comprise a multitude of mass spectra, for example a part of the series of mass spectra or the entire series of mass spectra acquired by the measurement module 110.
[0075] In this example, the communication module 130 receives the mass spectra determined by the measurement module 110 and stores them in the storage module 140.
[0076] The communication module 130 is, optionally in this example, connected to a database 4, external to the device 100. The database 4 typically comprises previously acquired mass spectra.
[0077] The database 4 is used in particular when the device 100 does not comprise a measuring device 110. Therefore, in this case, the device 100 does not directly measure mass spectra 2 but extracts, from the database 4, mass spectra 2 which have been previously recorded. Therefore, it is possible to work with previous mass spectra, no measurement then being necessary. In this case, these received mass spectra (from the database 4) have preferably all been acquired by the same mass spectroscopy device 10.
[0078] In the example considered, the communication module 130 receives mass spectra 2 from the measuring device 110 (for example a series of mass spectra) and from the database 4. Similarly, these received mass spectra were acquired by the same mass spectroscopy device 10. In other words, the mass spectra of the database were previously acquired by the measuring device 110 and recorded in the database, external to the device 100. In a variant, this database is included in the measuring device 110 (memory of the device 10) but remains different from the memory of the measuring device 110.
[0079] To receive the data from the database 4, the communication module 130 is also configured to be connected to at least one communication network 5 which may be a wireless communication network or a wired communication network. In practice, the wireless communication network uses at least one of the following technologies: WiFi technology, Bluetooth technology, 2G / 3G / 4G / 5G technology, etc. The wired communication network uses ADSL technology, etc. Thus, in the device 100, the mass spectra measured by the measuring device 110 or extracted from the database 4 are received by the communication module 130 which transfers this data to the storage module 140 of the device 100.
[0080] The storage module 140 comprises, for example, a memory and / or a hard disk.
[0081] The storage module 140 stores in particular computer program instructions designed so that the processing unit 120 implements at least some of the steps of the method described below with reference to [Fig. 3] when these instructions are executed by the processing unit 120.
[0082] All mass spectra received by the communication module 130 are concatenated to form a single series of mass spectra. This single series of mass spectra comprises at least 1000 mass spectra, preferably between 2000 and 8000 mass spectra.
[0083] Thus, the mass spectra measured by the measuring device and / or the mass spectra extracted from the database 4 are concatenated in the storage module 140 to form a plurality of mass spectra.
[0084] The storage module 140 thus stores this plurality of mass spectra.
[0085] Preferably, the plurality of mass spectra is a time series of mass spectra. In other words, all the mass spectra of the series of mass spectra are measured by the same mass spectroscopy device 10 for the same sample 1. Such a configuration has the advantage of obtaining a more precise series of mass spectra because it is directly acquired with the same mass spectroscopy device 10. In a variant, the series of mass spectra can be acquired by several devices of similar mass spectrum 10 and under similar measurement conditions. This has the advantage of being faster to obtain the mass spectrum series.
[0086] [Fig.3] represents an example of a method implemented by the device 100.
[0087] This method 1000 is a method of reconstructing a mass spectrum of at least one precursor ion from a plurality of mass spectra measured on the ionized sample comprising at least one precursor ion and fragment ions.
[0088] In practice, the method 1000 is implemented by computer, for example by being implemented by the device 100 illustrated in [Fig.l].
[0089] The method 1000 can also be implemented in microcontrollers or programmable gate arrays (e.g. FPGAs).
[0090] This method 1000 comprises a step 1010 of acquiring a series of mass spectra on the ionized sample comprising at least one precursor ion and fragment ions. In the present disclosure, by acquisition is meant the direct measurement of mass spectra and / or the reception of mass spectra.
[0091] For this purpose, the acquisition step 1010 comprises a preliminary step (optional) 1001 for recovering the plurality of mass spectra. In practice, the preliminary step 1001 begins with a step 1002 of measuring mass spectra by the measuring unit 110 of the device 100 and / or a step 1004 of receiving mass spectra from the database 4. An example of a preliminary step is illustrated in [Fig.4]. In this example, the preliminary step comprises the step 1002 of measuring the mass spectrum by the measuring unit 110 and the step 1004 of receiving mass spectra from the database 4. In this example, all of these mass spectra (from the measuring module 110 and from the database 4) are concatenated in a concatenation step 1006 in the memory module 4 to form the plurality of mass spectra of the at least one precursor ion and the fragment ions.
[0092] This plurality of mass spectra (or series of mass spectra) is then detected and acquired 1008, by the processing unit 120 of the device 100, to reconstruct a mass spectrum of at least one precursor ion from this series of mass spectra.
[0093] The method 1000 illustrated in [Fig.3] then optionally comprises a step 1015 of processing the plurality of acquired mass spectra. Typically the processing unit 120 processes all the mass spectra of the plurality of mass spectra.
[0094] Different treatments can be implemented on the mass spectra of the mass spectrum series. The treatments include at least one of the treatments below: - smoothing, for example, using a Savitzky-Golay algorithm with an al- baseline correction algorithm, etc.; - normalization, for example by dividing each acquired mass spectrum by a constant a which is a function of the acquired mass spectrum. In practice, the constant a corresponds to the abundance (intensity) of a peak in the mass spectrum, in particular to the abundance of the most intense peak (i.e. the one with the highest intensity). Thus, for normalization, each peak in the mass spectrum considered is divided by the abundance of the most intense peak in the mass spectrum; - extraction of peaks known in English by peak picking algorithm; - elimination or filtering, for example using a signal-to-noise ratio to remove irrelevant mass spectra; - deisotoping, - grouping of data, for example using techniques known as binning; - filtering, for example by convolving each spectrum to a particular kernel.
[0095] In a non-limiting manner, the processing unit 120 can filter each mass spectrum of the plurality of mass spectra in order to attenuate noise on each spectrum linked to the measurement of these mass spectra which can impact the abundances of the m / z ratios. Here, typically, the processing unit 120 determines the signal-to-noise ratio of each mass spectrum of the plurality of mass spectra and eliminates, for example, the mass spectra comprising a signal-to-noise ratio less than or equal to 0.5.
[0096] The processing step 1015 may also comprise compressed acquisition processing, also known as “compressed sensing”. This processing may be carried out as an alternative to the processing listed below or in combination thereof.
[0097] After filtering, the processing unit 120 normalizes each mass spectrum of the plurality of acquired mass spectra. For this, for each mass spectrum, the processing unit 120 determines the peak having a maximum abundance (for example via the peak picking algorithm) and then divides the abundances of the other peaks of the given mass spectrum by the previously determined maximum value. The normalization constant a is made to vary from one mass spectrum to another.
[0098] After normalization, the processing unit 120 can perform peak alignment processing to select ranges of given m / z ratios. This has the advantage of subsequently working on only a portion of the mass spectra of the mass spectrum series.
[0099] Following the processing step 1015 (if present in the method 1000), the method 1000 then comprises a step 1020 of determining a plurality of series of abundance values of the ionized sample. In practice, in the determining step 1020, a series of abundance values of the at least one precursor ion and a series of abundance values of each fragment ion are determined. In the present disclosure, each series of abundance values comprises N abundance values where N is an integer greater than or equal to two. In practice, each integer assigned from 1 to N corresponds to one mass spectrum of the plurality of mass spectra.
[0100] In a first variant, the determination step 1020 may be implemented by the processing unit 120 as follows. For example, the processing unit 120 scans each mass spectrum used as input data of the determination step 1020 and for each mass spectrum, retrieves and records an abundance value for each m / z ratio, thus creating a series of abundance values (i.e. equivalent to a list of abundance values) for each m / z ratio of the mass spectra of the plurality of mass spectra. Here, by series of abundance values, we mean a succession of abundance values obtained from the same given m / z ratio. Thus, typically, for a given m / z ratio, each value of the series of abundance values corresponds to an acquisition (i.e. a measurement), i.e. to a given mass spectrum N.There are therefore, in a series of abundance values, as many abundance values as there are mass spectra used in the determination step 1020. In this example, there are therefore as many series of abundance values as there are m / z ratios detected in the mass spectra of the mass spectrum series. Each series of abundance values is typically recorded in the memory of the processing unit 120 and / or in the storage module 140.
[0101] In a second variant, the values of the abundance series correspond to abundance values integrated over a range of m / z ratio values.
[0102] To obtain this type of abundance value series, the processing unit 120 can first calculate an average mass spectrum from the plurality of mass spectra by averaging, for each m / z ratio of the mass spectra of the plurality of mass spectra, the abundance values. An average spectrum is therefore obtained.
[0103] For example, [Fig.5] illustrates an average mass spectrum obtained from the plurality of mass spectra acquired from a sample comprising melittin and substance P. The average mass spectrum illustrated in [Fig.5] was obtained from a plurality of mass spectra. The plurality of mass spectra were each obtained by electrospraying a mixture comprising melittin, and substance P in solution in a mixture of water and methanol having a concentration of 105 mol / l. The mixture was for example activated by 20eV ultraviolet radiation for 500ms to provide the ionized sample.
[0104] The processing unit 120 can then determine ranges of m / z ratio values from the average mass spectrum, for example by detecting peaks in the average mass spectrum and calculating for each peak (22, 23, 24), an interval of m / z ratios comprising at least the peak considered.
[0105] For example, [Fig.6] shows a schematic representation of an average mass spectrum obtained from the N mass spectra of the plurality of mass spectra. In particular in this example, each interval W1, W2, W3 covers all of the abundance values of the peak considered 22, 23, 24 and can also cover m / z ratio values around the peak considered. Thus, a range of m / z ratio values corresponds to a part of the m / z ratios of the average mass spectrum. In this example, three ranges of m / z ratio values W1, W2 and W3 are detected by the processing unit. In this variant, an m / z ratio range W is associated with a peak of the average mass spectrum. These ranges of m / z ratio values are then recorded.The processing unit 120 can create a series of m / z ratio value ranges of size NxM, with N corresponding to the number of mass spectra of the plurality of mass spectra (here an integer greater than 2) and M corresponding to the number of m / z ratio value ranges (here an integer greater than 2) and which comprises, in the example of [Fig.6], three elements (the ranges W1, W2, W3). In practice, the M m / z ratio value ranges are similar for each mass spectrum N, which makes it possible to integrate all the mass spectra in a similar manner.
[0106] Subsequently, the processing unit 120 can, for each mass spectrum, integrate abundance values over given parts of each mass spectrum, each given part corresponding to a range of m / z ratio values determined previously. Typically, for a mass spectrum N, the mass spectrum considered is integrated over different ranges of values defined by the M ranges of values of the series of ranges of m / z ratio values at the coordinate N of the series of ranges of m / z ratio values.
[0107] An integrated abundance value is therefore obtained for each determined range of m / z ratio values. The integrated abundance values over the same range of m / z ratio values across the entire plurality of mass spectra make it possible to obtain a series of abundances of an ion or fragment ion. Thus, from all the ranges of m / z ratio values, series of abundance values are obtained.
[0108] In this embodiment, each series of abundance values is subsequently associated with a specific m / z ratio included in the range of m / z ratio values considered. The specific m / z ratio may correspond to the median value of the m / z ratio of the range of m / z ratio values, or preferably to the m / z ratio associated with a maximum abundance value of the average mass spectrum over the given range of m / z ratio values. Thus, according to this variant, each series of abundance is associated with a peak of the average spectrum. Fewer series of abundance values are therefore calculated in this variant, which improves the speed of the method 1000. Furthermore, since each series of abundance values is calculated over a range of m / z ratio defined on the basis of a peak of the average spectrum, it can be considered that each series of abundance values corresponds to an ion of the ionized sample. Furthermore, integrating the abundance values of the sample allows to associate ions of the ionized sample belonging to the same isotopic distribution. When the resolution allows it, the isotopes are separated but we wish to group them as a single ion. Of course, this second variant, which can correspond to a step of grouping the isotopes (de-isotoping), can be done by other algorithms known to those skilled in the art.
[0109] In practice, the series of abundance values determined in the determination step 1020 can be recorded and grouped in the same file.
[0110] The method 1000 optionally comprises a step 1025 of processing the series of abundance values comprising at least one of the following calculations: - filtering the given series of abundance values by a Kalman filter; - filtering the given series of abundance values by a Fourier filter, - filtering by moving average of the given series of abundance values.
[0111] Such a step 1025 makes it possible to remove noise from the abundance values of the abundance value series. For example, in the integration over the intervals (second variant of the determination step 1020), the processing unit 120 can apply a time filter to each interval. This has the advantage of attenuating the noise throughout the acquisition, and of highlighting temporal variations. Typically, the Kalman filter may be preferred to remove the integration noise. Typically, [Fig.8] illustrates an example comprising a series of abundance values 32 of an ion and a filtered series of abundance values 33 of this same ion here filtered by a Kalman filter. It can be seen that the abundance values of the filtered series 33 fluctuate less. The signal is more stable.
[0112] Subsequently, the method 1000 comprises a step 1030 of arranging the plurality of series of abundance values into pairs of series of abundance values, each pair of series of abundance values comprising a first series of abundance values (denoted X) and a second series of abundance values (denoted Y).
[0113] In practice, the processing unit 120 retrieves a series of abundance values from the plurality of series of values and combines it with another series of abundance values from the plurality of series of abundance values. Preferably, there may be as many combinations as there are series of abundance values. Thus, for P series of abundance values determined (with P an integer greater than 2, P2 pairs of series of abundance values are determined by the processing unit 120 considering that a series of abundance values can form a pair of series of abundance values with itself).
[0114] [Fig.7] illustrates a pair of abundance value series consisting of a first abundance value series 30 associated in this example with an m / z ratio of 1365 and a second abundance value series 31 associated in this example with an m / z ratio of 1424. Here, N corresponds to 5532.
[0115] The method 1000 also comprises, for each pair of series of abundance values, a step 1040 of calculating a correspondence score between the first series of abundance values and the second series of abundance values. This correspondence score is therefore calculated for all the pairs of series of abundance values. The set of correspondence scores obtained for all the pairs of series of abundance values provides a score map, which is typically a symmetric matrix of size PxP.
[0116] In practice, the match score is based on at least one of the following: - mutual information between the first series of abundance values and the second series of abundance values, said mutual information being based on an entropy of the first series of abundance values and on an entropy of the second series of abundance values; - a correlation between the first series of abundance values and the second series of abundance values.
[0117] The match score determined in calculation step 1040 is based on one of the above-mentioned cases or on a mixture of the above-mentioned cases. However, the score card is composed of match scores of the same type, i.e. calculated in the same way using the same calculation method.
[0118] Of course, the processing unit 120 can calculate different correspondence scores, for example a score based on mutual information, a score based on a correlation and a score based on mutual information and on the previously calculated correlation.
[0119] The different correspondence scores that can be determined at calculation step 1040 will be explained below.
[0120] These matching score examples are determined for each pair of previously recorded abundance value series. For the sake of simplification, the calculations described below describe an application of calculation step 1040 to a single pair of abundance value series.
[0121] Mutual information
[0122] In this variant, the calculation step 1040 comprises a step of determining the mutual information 1041. Different embodiments of the step of calculating the mutual information will be described below.
[0123] [Fig.9], [Fig.10], [Fig.11] and [Fig.12] illustrate an embodiment in which the matching score is based on mutual information (mathematics). In this embodiment, the mutual information is determined by an information theory method to be described below that uses entropy calculations.
[0124] Mutual information between two variables
[0125] In a first embodiment of step 1041, the mutual information between two variables is defined by the following formula:
[0126] [Math.l] / (XY)=h(X)+h(Y)-h(X, Y)
[0127] with X corresponding to a first variable (here, for example, it can correspond to the first series of abundance values defined for an ion of the ionized mixture, the ion being able to be a precursor ion or a parent ion) and Y a second variable (here, for example, it can correspond to the second series of abundance values defined for an ion of the ionized mixture).
[0128] Here, h() corresponds to the entropy (also called intrinsic entropy) defined for a given random variable expressed by the mathematical function:
[0129] [Math.2] h[x^ = J
[0130] where px(x) corresponds to the probability density function of the continuous random variable X (corresponding to the probability that a signal appears in the variable X) (here the abundance values of each acquisition considered).
[0131] The entropy h() can also be written for non-continuous variables by the following formula: p *log (pÿ with corresponding to the probability that a signal appears in the variable X.
[0132] The entropy h() cannot be directly calculated in this example. It is therefore necessary to approximate this quantity h() by using an estimate.
[0133] Thus, to determine the correspondence score of each pair of series of abundance values, the processing unit 120 estimates (in an estimation step A1 of the determination step 1041 of the mutual information), for each pair, the entropy of the first series of abundance values and the entropy of the second series of abundance values using an estimation algorithm. These estimated entropies are denoted H.
[0134] The estimation step A1 according to the present disclosure will be described.
[0135] In the present disclosure, an equivalence is sought between the entropy h() and the estimated entropy H().
[0136] In practice, the entropy estimated for a variable X (corresponding here for example to the first series of abundance values) is expressed according to the formula:
[0137] [Math.3] H(X] = d^average^^ distance (L)) + log(y) + psi(n)~ psi(L)
[0138] with d corresponding to the dimension of the variable X (here equal to 1), Y distanced} corresponding to the set of values of the distances determined by the estimation algorithm (the calculation of this distance is expressed below), n corresponding to the number of abundance values in the series of abundance values (here n is equal to N), L corresponding to the number of neighbors considered in the estimation algorithm (here equal to 1 for example when the distance is determined for a single neighbor), psi() corresponding to a Digamma function which can be expressed by psi(x^ = 9ui can satisfy a recursion psi(x+l) = psi (x) + 1 / x with psi(l) = -C for C = 0.5777 being the Euler-Mascheroni Constant, V corresponding to a volume of a unit hypersphere expressed according to the following formula:
[0139] [Math.4] ^[rl - v L 1 r(*+i)
[0140] with r the radius of the hypersphere (equal to 1), a gamma function F which can be expressed by r.
[0141] Different methods can be used to estimate the entropy of the first series of abundance values and the second series of abundance values. In particular, the processing unit 120 can use at least one of the following methods: a method based on histograms, a method based on a k-nearest neighbors algorithm, a kernel density estimation method, etc.
[0142] In practice, in the present disclosure, the processing unit 120 may use a k-nearest neighbors (kNN) algorithm to estimate the entropy H of a given series of abundance values used in calculating the mutual information. The kNN algorithm is chosen because this method is robust, compared to an estimation method using a histogram, and operates on small sample sizes.
[0143] Typically, the kNN algorithm uses a distance calculation between the abundance values of the same series in order to approximate the probability density of the series.
[0144] In practice, the k-nearest neighbor method uses a Euclidean distance to estimate the joint entropy between the data (abundance values) of the first series of abundance values and the data of the second series of abundance.
[0145] In this example, the Euclidean distance used depends on the first series of abundance values and the second series of abundance values considered, in particular on the processing that may have been carried out in processing step 1015.
[0146] In the following, two consecutive abundance values of the first series of abundance values (associated with an ion) are denoted by A = (x;, xi+i) and two values consecutive abundance values of the second series of abundance values (associated with an ion) are denoted by B = (yi5 yi+1), with x; an abundance value of the first series of abundance values for an acquisition i, with xi+i an abundance value of the first series of abundance values for an acquisition i+1, with y; an abundance value of the second series of abundance values for an acquisition i and with yi+i an abundance value of the second series of abundance values for an acquisition i+1.
[0147] In practice, when no normalization is calculated in the processing step 1015 or when the method 1000 does not include a processing step 1015, the Euclidean distance between the abundance values of a first series of abundance values and the abundance values of the second series of abundance values is expressed by the formula:
[0148] [Math.5]
[0149] with A a variable (i.e. data corresponding here to the first series of abundance values (A=X)), B a variable (i.e. data corresponding here to the second series of abundance values (B=X)), x; a variable corresponding to a signal from A (therefore here to an abundance value included in the first series of abundance values at acquisition i, i.e. an abundance value of an ion of the ionized sample), xi+i to the signal following x; (therefore the abundance value of the ion at acquisition i+1), y; a variable corresponding to a signal from B (therefore here to an abundance value included in the second series of abundance values at acquisition i, i.e. an abundance value of another ion of the ionized sample), yi+i to the signal following y; (therefore corresponding to an abundance value of the other ion at acquisition i+1).
[0150] In the case of normalization of the mass spectra at processing step 1015, the Euclidean distance defined by the formula Math. 5 does not apply. Therefore, the entropy h() is not directly equivalent to the estimated entropy H() by the estimation algorithm.
[0151] Indeed, the entropy of a variable X (here the variable X corresponds to a series of abundance values of an ion included in the pair of abundance series) is not invariant by multiplication, as described by the following formula:
[0152] [Math.6] hkkX) =h(X) + log(\k\)
[0153] with k is a multiplicative constant, X a variable corresponding here to a series of abundance of an ion (for example a precursor ion), h corresponding to the intrinsic entropy of the variable X.
[0154] In this variant, when the entropy of the first series of abundance values and the second series of abundance values is estimated from data that have been normalized by the same normalization constant, noted k (for example when a = k), the entropy is also invariant by multiplication by a multiplicative constant k (here k = a) as described by the following formula:
[0155] [Math.7] h(kX, kY) — h(X, K) + log(k2)
[0156] with k is a multiplicative constant, X a variable corresponding in this case to the first series of abundance values, Y a variable corresponding in this case to the second series of abundance values.
[0157] Similarly, Math. 7 does not apply directly when the multiplicative constant of the first set of abundance values is different from that of the second set of abundance values, for example when the normalization constant that was used to determine the data of the first set of abundance values is different from that which was used to determine the data of the second set of abundance values.
[0158] Thus, in the case of a normalization at the processing step 1015, to have an equivalence between the entropy h() and the estimated entropy H, it is necessary to work from invariant entropy and invariant estimated entropy.
[0159] By invariant, we mean that the entropy considered, for example the entropy estimated on a series of given abundance values (H(X), with X the series of given abundance values), is equal to the estimated entropy calculated on the same series of abundance values multiplied by a constant (H(kX) with X the series of given abundance values and k the multiplicative constant, here possibly being equal to the normalization constant used in the processing step 1015).
[0160] To obtain such an equivalence, the entropy estimated from the first series of abundance values used in the calculation of the two-variable mutual information is normalized by a first invariance parameter and the entropy estimated from the second series of abundance values used in the calculation of the two-variable mutual information is normalized by a second invariance parameter.
[0161] In practice, the processing unit 120 uses, for a given series of abundance values, the invariant parameter m(X) (also called invariant measurement) defined according to the following three properties: i) m(X) = rx >0 ; ii) m(kX) = km(X) = krx iii) m(X + b) = m(X) = rx. with X the series of abundance values considered, m(X) the invariant measure of the series of abundance values considered, rx> 0 being a median value of the set of distances between each pair of abundances closest to the series considered, k a chosen multiplicative constant (which can depend on the normalization constant a of the processing step) and b an additive constant whose value can for example depend on noise (for example on noise included in the abundance values of the series of abundance values).
[0162] Using the properties defined above, we can obtain an equivalence between the entropy h() and the estimated entropy H(), so that:
[0163] [Math. 8] h(x) = / / (7^)
[0164] with X a series of given abundance values, h the intrinsic entropy, H the estimated entropy determined in the calculation step, m(x) the invariant measure. (It should be noted that the invariance property also allows us to have the equivalence H(X) = 1
[0165] The Math. 8 formula can be generalized to an entropy between several variables, typically when several series of abundance values are extracted from the plurality of mass spectra. In other words, this is notably used when several series of abundance values of ionized ions and / or fragments of ionized ions are to be taken into account. In this case, the equivalence of the Math. 8 formula is generalized in n dimensions (n = N):
[0166] [Math.9] h(xyz) -2- -2-1
[0167] with X, Y.. ..Z corresponding to the series of abundance values of the ion X or Y or Z considered.
[0168] Thus, the equivalence sought for an entropy calculated between two variables is written as follows:
[0169] [Math. 10] htx y1 -2- 1 Ij n\m(xy m{Y} J
[0170] with X and Y two series of abundance values where X is associated in the following with the first series of abundance values belonging to a given pair of series of abundance values (here corresponding to a series of abundance values of an ion of the ionized sample) and Y with the second series of abundance values belonging to the same pair (here corresponding to a series of abundance values of another ion of the ionized sample), m(X) with the first invariance parameter of the series X (i.e. invariance parameter of the first series of abundance values) and m(Y) with the second invariance parameter of the series Y (i.e. invariance parameter of the second series of abundance values), h the entropy and H the estimated entropy.
[0171] Based on the above properties, when the multiplicative constant of the first series of abundance values (here A=X) is identical to the multiplicative constant of the second series of abundance values (here B=Y), the Euclidean distance used in the step of determining the mutual information 1041 can also be expressed by the following formula:
[0172] [Math. 11] d[kA, kBj = (kXj-kxi+ 3 )" + ( / cy- ky.+]) - kd^
[0173] Indeed, by multiplying the first series of abundance values and the second series of abundance values by the same multiplicative constant (k), the abundances at each acquisition of the first series of abundance are written kA = (kx, kxi+i) and the abundances at each acquisition of the second series of abundance are written kB = (ky i, kyi+i). Typically, k can be a negative or positive number.
[0174] When the multiplicative constant of the first series of abundance values (denoted here aj) is not identical to the multiplicative constant of the second series of abundance values (denoted here a2), the Euclidean distance used in the determination step 1041 can also be expressed by the following formula:
[0175] [Math. 12] _ __ _ _ OjB] = J
[0176] with ai multiplicative constant of the first series of abundance values and a2 multiplicative constant of the second series of abundance values. Similarly al and a2 are two constants which can be negative or positive.
[0177] To calculate such a distance d(aiA, a2B), the processing unit 120 can carry out a change of variable by setting:
[0178] [Math. 13]
[0179] [Math. 14]
[0180] With the formulas Math. 13 and Math. 14, we obtain the following equations:
[0181] [Math. 15] Û|A — ~ a (G- , And
[0182] [Math. 16] rtjy aB / 7 W --- ■■■»■■■ --- ....-7..., , And
[0183] [Math. 17] 4=(¾¾1) , And
[0184] [Math. 18] 5 = (¾¾)
[0185] The processing unit 120 thus uses the formulas described above to define in its estimation algorithm the Euclidean distance according to the formula described below:
[0186] [Math. 19] / > / aix \2 / thy «2y d^A, a2B ) + \^77 - tte- J
[0187] , , I y ” 77 y 75 z , rf(aiA,^B) (^-^) =d(A,B )
[0188] It will now be described how the processing unit 120 determines the invariance parameter of the two series of abundance values of each pair of series of abundance values (step of determining the invariance parameter A1 of each series of the pair of series considered in the estimation step A1).
[0189] As explained above, each invariant measurement is associated with a series of abundance values. The processing unit 120 therefore determines for each series of abundance values determined previously, an invariant measurement (step Ail). This invariant measurement, as explained above, depends on the values of the abundances of the series of abundance values.
[0190] In practice, the invariant measure can be determined upstream of the determination estimate 1041, for example following the determination step 1020.
[0191] An example of a step of determining an invariance parameter on a first series of abundance values and an invariance parameter on a second series of abundance values (step Ail) will therefore be described.
[0192] For a given series of abundance values (for example the first series of abundance values), the processing unit 120 calculates (step PI) all the differences between the abundance values included in the series of abundance values. For this, the processing unit 120 scans the series of abundance values and calculates, for each abundance value, a difference with all the other abundance values in the series of abundance values. Thus for each abundance value in the series of abundance values, Nl differences are determined. A list of difference values is for example recorded for each value of abundance.
[0193] The processing unit 120 then searches, for each abundance value in the series of abundance values, for the smallest difference between this abundance value and another abundance value included in the series of abundance values (step P2). To do this, the processing unit scans the list of differences associated with the given abundance value and searches for the smallest difference obtained in the list of differences associated with this abundance value. As a result, each abundance value is associated with a minimum distance. Thus, N minimum distances are determined for this series of abundance values considered. These minimum distances are then arranged in ascending order.
[0194] The processing unit 120 then determines a median from these N minimum distances (step P3) and then multiplies this median value by the number N of abundance values of the series of abundance values considered (step P4). This last calculated value is then assigned to the invariance parameter (here the first invariance parameter) of the series considered.
[0195] These steps are also performed on the other series of the mass spectrum pair (e.g. on the second series) to obtain the second invariance parameter.
[0196] Thus, the step of determining the invariance parameter A1 1 provides for each series of abundance values, an invariance parameter. The invariance parameter respects the three properties i, ii) and ii) explained above.
[0197] The estimation step A1 comprises a step of normalization (step A12) of each series of abundance values by its invariance parameter. Thus, for example, in this step, the processing unit 120 divides each abundance value of the series of abundance values considered by the invariance parameter determined on this series. In the following, the first series of normalized abundance values is noted — for the first series of abundance values and the second series of normalized abundance values is noted -Xr for the second series of abundance values.
[0198] Therefore, after the normalization step, each series of abundance values is normalized.
[0199] The estimation of the entropy in the estimation step carried out on the pairs of mass spectra uses these series of normalized abundance values.
[0200] Thus, in the method 1000, in the estimation step A1 (included in the determination step 1041 of the mutual information), the processing unit 120 determines (step A13) an estimated invariant entropy using the estimation algorithm (for example, the kNN algorithm in this example) as explained below.
[0201] In practice, in this example, the estimation algorithm determines for - the invariant entropy / 4%) — x 1 for the ion associated with the first spectrum of \ 7 \ m( X ) / mass, - the invariant entropy ht y ] — H { y 1 for the ion associated with the second spectrum of ■ / \ «413 / mass, and - the invariant entropy in two dimensions y\ — fil xyj as defined \ ' / \ m(X) ' m(Y) / in equation Math. 10 using equation Math. 3 as defined above to determine each estimated entropy H. The distance formula used in the kNN algorithm is the Euclidean distance defined in formula Math. 12 with k preferably being equal to 1.
[0202] In the step of determining the mutual information 1041, the processing unit determines (step A2) the conditional information I(X,Y) between the first series of abundance and the second series of abundance on the basis of the formula Math. 1 now written thanks to the previous calculations using the estimated entropies H:
[0203] [Math.20] I(X, Y) = H(X) + H(Y) - H(X, Y)
[0204] The step of determining the mutual information 1041 is calculated for each pair of series of the mass spectrum series. Thus, for each pair of abundance value series, the calculated mutual information I corresponds to a correspondence score between the first abundance series and the second abundance series of a given pair of abundance series.
[0205] Thus, the set of correspondence scores calculated in the calculation step 1040 (in a step 1042) can be represented in the form of a two-dimensional matrix, with for example along the rows, all the first series of abundance values of the set of pairs and along the columns all the second series of abundance values of the pairs. From this, each “pixel” of the matrix corresponds to a correspondence score calculated from a given pair of ions, the ions being able to be identified with the row and column of the matrix. This data, illustrated in two dimensions, corresponds to the score map (which is a two-dimensional matrix).
[0206] [Fig. 10] illustrates different point clouds of mutual information calculated from non-invariant estimated entropies H (i.e. from a first series of abundance values and a second series of abundance values not normalized by their invariance parameter explained above). These curves were plotted by taking on the abscissa the abundance values (intensity) of an ion (here in particular the ion associated with a first series of abundance values which has an m / z ratio of 1222) and on the ordinate the abundance values of the other ion of the pair considered (here the ion associated with the second series of abundance values which has an m / z ratio of 1712). In this example, a mixture of proteins comprising ubiquitin, cytochrome C and lysozyme in solution in a mixture of water and methanol having a concentration of 105 mol / 1. A region centered on m / z 1760 and with a width of 100 m / z was selected for a tandem mass spectrometry step. The selected ions were then subjected to activation of 20eV ultraviolet radiation for 500ms. Typically, here, the multiplicative constant k was applied to the first series of abundance values and estimated for both series was calculated with the formula Math. 3 with a given multiplicative constant k, in particular with k equal to 1 for curve 50, k equal to 5 for curve 51, k equal to 10 for curve 52.
[0207] [Fig. 11] illustrates different point clouds of mutual information calculated from an invariant estimated entropy H (i.e. from a first series of abundance values and a second series of abundance values normalized by their invariance parameter as explained above). Typically, the estimated entropy H for both series was calculated with the formula Math. 3 with a given multiplicative constant k, with k equal to 1 for curve 53, k equal to 5 for curve 54, k equal to 10 for curve 55.
[0208] It can be seen that for curve 50, curve 51, curve 52, the mutual information determined with the different values of the multiplicative constant are different. Thus, the multiplicative constant used in calculating the entropy of the first series of abundance values affects the calculation of the mutual information. Conversely, for curve 53, curve 54, curve 55, the mutual information determined with different values of the multiplicative constant are similar. Thus, the multiplicative constant used in calculating the entropy of the first series of abundance values does not affect the calculation of the mutual information.
[0209] Therefore, thanks to the invention, the calculated mutual information is invariant, that is to say it is constant thanks to the estimated invariant entropy and therefore does not depend on the abundances of the ions of the spectra used. The invariant entropy estimated according to the present disclosure therefore makes it possible to express any estimated entropy value in a single frame of reference.
[0210] In the method 1000, here in particular in step 1041, the mutual information is calculated for all pairs of series of abundance values. All of this mutual information can be concatenated in a score map (step 1042). Thus, in this example, all of the mutual information determined in step 1041 makes it possible to obtain the score map.
[0211] [Fig. 12] illustrates a score map obtained from a matching score based on a calculation of mutual information. The calculation of mutual information was obtained from a series of abundance values obtained on an ionized sample, the series of abundance values comprising values integrated over 10 ranges of mass-to-charge ratio values (i.e. according to the second variant of the de termination). The ionized sample is the same as that of [Fig. 10] and [Fig. 12] comprising a protein mixture including ubiquitin, cytochrome C and lysozyme. In this figure, the 10 ranges of m / z ratio values of the ions of the first series of pair abundance values are illustrated at the column level and the 10 ranges of m / z ratio values of the ions of the second series of pair abundance values are illustrated at the row level. As illustrated in this figure, each range of value is associated with an m / z ratio determined as described above.
[0212] It can be seen that the score map obtained is a symmetric matrix, which has the advantage of only performing calculations on half of the score map. Furthermore, as fewer pairs of abundance value series were used, the implementation time of the method 1000 is improved.
[0213] A second embodiment of step 1041 will now be described. In the following, the calculated mutual information is directly written from the estimated mutual information H thanks to the equivalence between the entropy h and the entropy H discussed above.
[0214] Conditional mutual information.
[0215] According to this embodiment, the mutual information used is conditional mutual information which is defined by the following formula:
[0216] [Math.21] I(X, Y\Z ) = H(X, Z) + H(Y, Z) - H(X, Y,Z)-H(Z)
[0217] with X, Y, and Z an additional variable, called conditional variable, H() the estimated entropy.
[0218] In practice, the conditional variable Z comprises at least one of the following elements: the total number of ions, known as the "Total Ions Current" (TIC). Typically, the total number of ions is the result of a sum of all the ions of a mass spectrum (of the plurality of mass spectra), a series of abundance values of another ion of the mixture, a series of constant abundance values, etc. In the following, the TIC is considered. In this case of the TIC, each mass spectrum of the plurality of mass spectra has a TIC. All the TICs of the plurality of mass spectra form a series of TICs which is used as a partial or conditional variable.
[0219] Therefore, from the series of mass spectra, the processing unit 120 can calculate a TIC (the conditional variable) for each mass spectrum of the plurality of mass spectra. This can for example be done upstream, for example in the determination step 1020. Thus, from the plurality of mass spectra, the processing unit 120 obtains a series of conditional variables composed of p conditional variables with p equal to the total number of mass spectra of the plurality of mass spectra (here p = N).
[0220] The conditional mutual information is determined according to a method similar to that explained above for the example of mutual information between two variables.
[0221] Thus, in the calculation step 1041, the processing unit uses the kNN algorithm to estimate the entropies H(X, Z), H(Y, Z), H(X, X, Z) and H(Z) of the equation Math. 21. Similarly, each estimated entropy term is an invariant entropy as explained above.
[0222] In practice, the step of estimating the conditional mutual information comprises the steps A1 to A2 as described above included in the step of calculating the mutual information 1041.
[0223] Thus, for each pair of series of abundance values, an invariant measure is determined for each variable X (corresponding to the first series of abundance values), Y (corresponding to the second series of abundance values), Z (corresponding to the series of TIC values) in step A1 1 of the entropy estimation step H of the calculation step 1041.
[0224] In practice, to estimate the entropy H(Z) of the conditional variable, the processing unit 120 calculates directly from the series of conditional variables Z, the invariant measure m(Z) associated with this series. To obtain the estimated entropy H(Z) of this series, the processing unit 120 normalizes the series of conditional values by dividing each conditional variable of the series of conditional variables Z by the invariant measure m(z) (calculation of ). Similarly, the processing unit 120 determines the invariant entropy H(Z) using the formula Math. 3 with L = 1 (a neighbor) (step Al). In this same step Al, the estimated entropies H(X, Z), H(Y, Z) and H(X, Y, Z) are determined in this step Al.
[0225] In practice, in this example, the estimation algorithm also determines in this step (in addition to the previous example), the following estimated entropies: n\m(Xy m(Z) J h(yz) = Vet \ ' / \ m(Y)' m(Z) / H(X, Y, Z] — h( ) ■ \ / \ m(X) ' m(Y) ' m(Z) / With X as the first set of abundance values and Y as the second set of abundance values and Z as the set of conditional variables. For this, the processing unit 120 uses the equation Math. 3 as defined above to determine each estimated entropy H. The distance formula used in the kNN algorithm is the Euclidean distance defined in the formula Math. 11.
[0226] Subsequently, the processing unit determines the conditional mutual information by applying the formula Math. 20.
[0227] A second embodiment of step 1041 will now be described.
[0228] Mutual information between three variables
[0229] According to this embodiment, the mutual information is conditional mutual information between three variables, called information score, and which is defined by the following formula:
[0230] [Math.22] I(X, Y, Z) = H(X) + H(Y) + H(Z) - H(X, Y) - H(X, Z) -H(Y, Z) + H(X, Y, Z)
[0231] In practice, the conditional variable Z is also a series of abundance values, which can be, like the second example described above, a series of TIC values (allows to overcome internal variations in the system) or a series of abundance values of a third ion (allows to highlight relationships related to this third ion) selected from the plurality of mass spectra (in a similar manner to the ion associated with the first and second series of abundance values).
[0232] The conditional mutual information between three variables is determined according to a method similar to that explained above for the example of conditional mutual information.
[0233] Thus, the processing unit uses the kNN algorithm to estimate the terms H(X), H(Y), H(Z), H(X, Y) H(X, Z), H(Y, Z), and H(X, X, Z) of the equation Math. 22. Similarly, each estimated entropy term is an invariant entropy as explained above.
[0234] In practice, the step 1041 of calculating the mutual information between three variables comprises the steps A1 to A2 as described above.
[0235] The processing unit 120 estimates the invariant entropy of each term H(X), H(Y), H(Z), H(X, Y) H(X, Z), H(Y, Z), and H(X, X, Z) of the equation Math. 22, in a manner similar to the first and second examples described above.
[0236] Subsequently, the processing unit 120 determines the conditional mutual information by applying the formula Math. 22.
[0237] As explained above, the correspondence score can also be based on a correlation between the first series of abundance values and the second series of abundance values. A variant of the calculation step 1040 will now be described in which the correspondence score is based on a correlation.
[0238] Correlation
[0239] According to a variant of the method 1000, the correspondence score may be based on a statistical method. In practice, in this embodiment, the method 1000 is based on a calculation of correlation coefficient but is devoid of entropy calculation.
[0240] In this embodiment, the processing unit 120 can calculate the correlation coefficient rxy (in a step of calculating a correlation coefficient 1044) on each pair of abundance series from the covariance according to the following formula:
[0241] [Math.23] Horn -
[0242] with Cov(X, Y) designating the covariance of the variables X (corresponding for example to the first series of abundance values) and Y (corresponding for example to the second series of abundance values), sx corresponding to the standard deviation of the variable X and sy corresponding to the standard deviation of the variable Y. Typically, the standard deviation of the variable considered corresponds to the standard deviation determined from all the abundance values of the abundance series considered.
[0243] In practice, according to this embodiment, the processing unit 120 calculates, for each pair of series of abundance values, the covariance between the first series of abundance values and the second series of abundance values according to the formula for the covariance of a random variable defined as follows:
[0244] [Math.24] Cov(X, Y) = E(XY) - E(X)*E(Y)
[0245] with E() corresponding to the mathematical expectation of the given variable, here the mathematical expectation of the set of abundance values of the abundance series considered.
[0246] Then, the processing unit 120 determines a standard deviation of the variable X and a standard deviation of the variable Y, the standard deviation considered being defined according to the following formula:
[0247] [Math.25] J = TÏMxy = Je(x 2 )-(ex)) 2
[0248] with X the variable considered (here the first series of abundance values or second series of abundance values), E(X) the expectation of the random variable considered, Var(X) the variance of the random variable considered.
[0249] Thus, in this embodiment, the processing unit 120 calculates for each pair of series of abundance values, the covariance as defined in the equation Math. 24 as well as the standard deviation sx associated with the first series of abundance and the standard deviation sy associated with the second series of abundance values. It then determines the correlation coefficient Corxy for each pair of series of abundance values.
[0250] As the calculation step 1040 is performed iteratively by repeating the calculation step on each pair of abundance value series or in parallel on all pairs of abundance series, a multitude of Corxy correlations are provided at the end of the step 1040.
[0251] Similar to the embodiment using a correspondence score based on the calculation of mutual information, the set of correlation coefficients calculated on the set of pairs of series of abundance values are grouped (step 1042) on a two-dimensional matrix (corresponding to a score map) of a similar shape to that obtained in the context of mutual information. Thus, compared to [Fig. 12], the matrix is made of Corxy correlation coefficients. The scale of such a score map is, however, different from the example with mutual information, since the calculated correlation coefficients can be negative or positive (unlike mutual information where the score values are positive).
[0252] Furthermore, although not necessary, it is also possible to determine the correlation coefficients from series of normalized abundance values. For example, the normalization coefficient may correspond to the invariance parameter of the series considered and calculated in a similar manner to the example of mutual information.
[0253] A second embodiment of step 1044 will now be described.
[0254] Partial correlation
[0255] According to another variant of the method 1000, the processing unit 120 can calculate the partial correlation coefficient Corxyz between two variables X, Y and a third variable Z. This third variable is similar to the variable Z used in the calculation of the mutual information. In other words, the variable Z is a series of TIC values or a series of abundance values of another ion, different from the ion associated with the first and second series of abundance values. This can also be a series made of constant values or a series of values relating to the intensity of a laser beam, a gas pressure, etc. Using a third variable Z makes it possible to avoid a quantity affecting the measurements.
[0256] In this case, the step of calculating the correlation coefficient 1044 can be implemented as follows.
[0257] In practice, the partial correlation coefficient of each pair of series of abundance values is defined according to the following equation:
[0258] [Math.26] _ CorXY-Corxz*Cori2 ^orXYZ~ r ^ / 777-7-^1-Cin xz ^l-Coryz
[0259] with Cor(), the correlation coefficient between the variables considered.
[0260] For this purpose, the processing unit 120 calculates for each pair of series of abundance values, the correlation coefficients Coryx, Corxz and Coryz as previously determined using the formula Math. 23 then applies the formula Math. 26 to determine the partial correlation coefficient between three variables.
[0261] The matching score obtained by this embodiment is more accurate than the matching score of the embodiment based on a calculation of the correlation coefficient between two variables as explained above.
[0262] Typically, the numerator of equation 23 is equivalent to a standard deviation normalization. Such standard deviation normalization in the correlation calculation has the same effect as the invariant measure normalization in the entropy calculation.
[0263] [Fig. 13] illustrates a score map obtained from a correspondence score based on a correlation coefficient (here partial correlation coefficient). The results displayed correspond to the ionized sample comprising Melittin and substance P. The correspondence scores were determined from a series of abundance values obtained on an ionized sample, the series of abundance values comprising non-integrated values (i.e. according to the first variant of the determination step 1020). In this figure, the values of the m / z ratios of the ions of the first series of abundance values of the pairs are illustrated at the column level (on the abscissa) and the values of the m / z ratios of the ions of the second series of abundance values of the pairs are illustrated at the row level (on the ordinate).
[0264] It can be seen that the obtained score map is a symmetric matrix. Furthermore, it can be seen that more pairs of abundance value series were used, so the implementation time of the method 1000 may be longer. The matching scores according to this method can be positive or negative.
[0265] [Fig. 14] illustrates a correspondence score map obtained with series of Kalman-filtered abundance values illustrated previously (i.e. obtained from the mixture of Melittin and substance P). It can be seen that the score map has a better contrast. The contrast can be evaluated by taking the average of the correspondence scores of all the pixels in the correspondence map. In a preferred variant, the contrast can be evaluated by calculating a variance over all the correspondence scores (i.e. all the pixels) in the score map. In this case, the higher the variance, the greater the contrast.
[0266] A combination of the statistical and probabilistic method for determining the match score calculated in calculation step 1040 will now be described.
[0267] As explained above, the matching score may also be based on a correlation between the first series of abundance values and the second series of abundance values and on mutual information between the first series of abundance values and the second series of abundance values.
[0268] Combination of the statistical method and the probabilistic method described above.
[0269] In this embodiment, the correspondence score is based on the probabilistic method based on an entropy calculation applied to the first series of abundance values and the second series of abundance values and on the statistical method based on a calculation of a correlation between the first series of abundance values and the second series of abundance values.
[0270] Thus, in practice, the processing unit 120 determines, for each pair of series of abundance values, two preliminary correspondence scores by calculating for each pair of series of abundance values: - at least one of the mutual information (step 1041) described above I(X, Y), I(X, YIZ), I(X, Y, Z), i.e. by calculating the mutual information defined by one of the following formulas: Math. 20, Math. 21, Math. 22; and - at least one of the correlation scores (step 1044) described above Corxy, Corxyz, i.e. by calculating the correlation coefficient defined by one of the following formulas: Math. 23, Math. 26.
[0271] Preferably, in this embodiment, the processing unit 120 combines the mutual information between three variables I(X, Y, Z) with the partial correlation coefficient Corxyz.
[0272] In this case, the processing unit 120 can calculate a mixed coefficient Sa(X, Y, Z) (step 1045) on each pair of series of abundance values which is defined according to the following formula:
[0273] [Math.27] Y, Z) = I(X, Y, Z) * sign(CorXYX)
[0274] with sign(Corxyz) corresponding to the sign of the partial correlation coefficient and I(X, Y, Z) corresponding to the mutual information between three variables.
[0275] In a variant of the method 1000, the processing unit 120 may weight at least one of the following terms, namely the mutual information between three variables, the partial correlation coefficient, in order to obtain a more precise matching score. In this case, the processing unit 120 calculates a weighted mixed coefficient Sap(X, Y, Z) on each pair of abundance series which can be defined according to the following formula:
[0276] [Math.28] Sa^X, Y, Z) = (alpha*I(X, Y, Z\)\beta* CorXYZ)
[0277] with alpha a first weighting coefficient greater than or equal to 1 and beta a second positive weighting coefficient greater than or equal to 1.
[0278] In practice, the first weighting coefficient depends on the correspondence score based on the mutual information considered and the second weighting coefficient depends on the information score based on the correlation coefficient considered.
[0279] For example, the correspondence score can be calculated according to the formula Math. 28 in which the conditional mutual information (i.e. the correspondence score based on the conditional mutual information) is multiplied with the partial correlation coefficient and in which the first weighting coefficient is equal to 1 and the second term of the equation Math. 28 is written sign(beta*Corxyz-0.1) with beta, second correlation coefficient equal to 1. In this example, the value 0.1 is selected because it allows to reduce a correlation intrinsic to the fragmentation process from which the TIC term cannot free itself.
[0280] In a variant, the equation Math. 28 can also be written: Sa^X, Y, Z) = Y, Z)\ g(CorYrz))with f() and g() two functions which apply a transformation on each coefficient (Here I(X, Y, Z) and Corxyz. These functions can be for example, a square function, an affine function, sign function, etc. The function h(a,b) a function which takes two variables and which applies a transformation like an addition, multiplication.
[0281] The other steps of the method 1000 will now be described after having obtained one or more correlation maps from the set of correspondence scores described above.
[0282] In certain cases, the correlation map(s) obtained in step 1042 are not usable because they have too low a contrast. In this case, the method 1000 optionally comprises a step 1050 of selecting the correlation maps as a function of a contrast of the score map considered.
[0283] Such a step makes it possible to eliminate the unusable correlation maps so as to limit errors in the exploitation of the correlation maps in the remainder of the method 1000.
[0284] In practice, this selection step 1050 can be directly calculated in the calculation step 1040 after having determined the score card associated with a given correspondence score. In this case, the processing unit 120 calculates, after having determined the score card, a contrast of the given score card and then compares this contrast to a given threshold to keep or eliminate this correspondence card. For example, the threshold can be a value greater than or equal to 0.30, for example 0.5.
[0285] Typically, in one example, the score map resulting from the two-variable mutual information-based matching score may have a contrast of less than 0.3. Therefore, following the selection step 1050, this score map may be deleted from the memory of the processing unit or the storage module 140.
[0286] If only one type of score map is calculated in the calculation step, only one correspondence map is obtained following step 1040. In this case, if this map has a contrast lower than the threshold, the method can repeat the calculation step by determining correspondence scores based on another calculation method (i.e., moving, for example, from mutual information between two variables to mutual information between three variables, or moving to a score based on a correlation coefficient or using a mixed correlation score, etc.).
[0287] The correlation map(s) passing the requirement of step 1050 are retained. in memory of the processing unit and / or of the storage module 140. Thus, in the remainder of the method 1000, only the correlation map(s) retained following the selection step 1050 are used in the remainder of the method 1000. Of course, if the method 1000 does not include a selection step 1050 as described above, all the correlation maps determined in the calculation step 1040 are used in the method 1000.
[0288] In the illustrated example, following the selection step 1050, the method 1000 comprises a determination step 1060, as a function of a mass to charge ratio, of at least one precursor ion and the associated fragment ions. For this, the method 1000 uses a classification method which is applied to the correlation map(s) calculated at the end of the calculation step 1040.
[0289] Thus, in this step 1060, the processing unit 120 applies to each score card a grouping method to determine at least one precursor ion as well as the fragment ions which are associated with this precursor ion. This precursor ion is determined according to an m / z ratio.
[0290] In the case where several precursor ions are determined in this step, each determined precursor ion is associated with its fragment ions.
[0291] Optionally, before applying the clustering method to the correlation maps, the method 1000 may comprise a step 1062 of processing the correlation maps. In practice, this processing step 1062 is carried out in the determination step 1060 but before the application of the classification method.
[0292] Typically, processing step 1062 uses at least one of the following processes: - an algorithm for processing the diagonal of the given score card. In this case, the processing unit 120 can assign the same score value (for example an integer) to each element present on the diagonal of the score card. For example, all the elements on the diagonal can be set to zero so that these elements do not significantly influence the classification algorithm. In the case of a score card based on correlation coefficients, the integer can be negative; - an algorithm for standardizing the correlation maps. In this case, the standardization algorithm can be applied to each line of the correlation maps. This standardization algorithm uses an average of a considered line as well as a standard deviation. In practice, the processing unit 120 calculates the average of each line and the standard deviation of each line. Then, for each element of a considered line, the processing unit subtracts from this considered element the average value of the considered line and then divides this value by the standard deviation of the considered line. Thus, with this processing, each line of the score card has a similar weight in the processed scorecard; - an application of a square function in order to accentuate the gap between the small and large coefficients (correspondence scores) of the scorecard considered.
[0293] In this example, the method 1000 then uses a clustering (or classification) method on the processed correlation maps (step 1064).
[0294] Typically, the clustering algorithm uses at least one unsupervised learning algorithm which may include at least one of the following algorithms: - a hierarchical ascending classification algorithm, for example using a Ward method (as an aggregation method) on the variance, and / or the Complete Link method, - a k-nearest neighbors algorithm, - a k-means algorithm, - a dimensionality reduction algorithm, also known in English as “t-distributed stochastic neighbor embedding algorithm” (t-SNE).
[0295] In the clustering method, each classification algorithm takes as input a score map. Thus, if several correlation maps are calculated in step 1040 or retained following the selection step, the processing unit 120 will implement several clustering algorithms.
[0296] The grouping method can also define in the input data other elements, comprising at least one of the following parameters: the method used (here for example the Ward method in the case of the unsupervised learning algorithm); the distance used in the (aggregation) method used; etc. These data depend on the type of algorithm used and are therefore known to those skilled in the art.
[0297] An embodiment of the determination step 1060 used in the method 1000 will now be described with the aid of [Fig.15], [Fig.16], [Fig.17], [Fig.18] and [Fig.19].
[0298] In the example of the method 1000, the determination step 1060 uses a hierarchical classification algorithm. In practice, in this example, the algorithm used uses the Ward method and takes by default a Euclidean distance. The Ward method is more suitable because it minimizes the variance by cluster fusion and is recommended for continuous values, unlike the other methods mentioned above, such as the Complete link method which takes into account the maximum values and not a global set of the score map considered.
[0299] [Fig. 15] illustrates a dendrogram obtained by the hierarchical classification method used and obtained from the data calculated on the ionized sample from a mixture of Melittin and substance P. Here, the dendrogram includes in The abscissa represents the different labels corresponding in this example to the m / z ratios of precursor ions and fragment ions and the ordinate represents a distance between different classes. In this example, each class is characterized by a horizontal line on the dendrogram. Vertical branches / links connect the classes to illustrate the links between the different classes.
[0300] As illustrated in [Fig. 15], the dendrogram comprises a main class Cl which is divided into two main vertical branches Cil, C12, each main vertical branch being again connected to a new class C2, which makes it possible to obtain a tree structure in which the vertical branches illustrate the links / proximities between the different classes connected to these branches.
[0301] Thus, it can be deduced from this dendrogram that the initial mixture comprises two precursor ions (visible by the two links of the main branch). Each main branch is connected to a new class (called secondary class) which is connected to secondary links which are themselves connected to a new class. Each branch connecting two classes positioned at two different distances illustrates the distance separating these two classes.
[0302] Thus, we visualize by the two main branches connected to the main class two large clusters / groupings of precursor ions and by the secondary branches, the fragment ions of this precursor ion. We can deduce that each secondary branch connected to a secondary class corresponds to a fragment of the precursor ion and that the arborescences (other branches connected to classes positioned at a distance less than that of the fragment considered) connected to this secondary class define the chemical compounds of this fragment. Consequently, the final vertical branches (the lowest branches, connected to no lower class) illustrate the m / z ratios of each chemical compound of the fragment ions of the precursor ion considered. Thus, thanks to this dendrogram, we can directly visualize the m / z ratios of each chemical compound of the fragment ions of the precursor ion of the class.
[0303] Thus, after having obtained this dendrogram, the processing unit 120 determines the precursor ions and the fragments associated with each precursor ion by exploiting the different branches and classes of the dendrograms, the distances separating the classes as well as the m / z ratios given at the distance equal to zero (zero ordinate).
[0304] The method 1000 then comprises a step 1070 of reconstructing the individual mass spectrum of each precursor ion determined and the associated fragments following the determination step 1060. In practice, the processing unit 120 uses the data extracted from the classification method to reconstruct the mass spectrum of each precursor ion, in particular the m / z ratios grouped in the same class. Thus, for each precursor ion determined in the previous step, the processing unit 120 retrieves each m / z ratio positioned at the minimum distance (equal to zero), and associates an intensity (i.e. an abundance value) with this ratio. This intensity can be calculated from the series of abundance values corresponding to this ratio, for example by calculating the average abundance value of this series of abundance values) or directly from the average spectrum by selecting the m / z ratio concerned on the average spectrum. The processing unit can therefore construct a peak on the mass spectrum of the precursor ion at each determined m / z ratio and assign it a corresponding intensity.
[0305] Such a reconstructed mass spectrum corresponds to a tandem mass spectrum of the determined precursor ion. Thus, with the method 1000, it is possible to reconstruct the tandem mass spectrum of each determined precursor ion in the sample used using mass spectra acquired from a simple mass spectroscopy device, i.e. without using a tandem mass spectroscopy device.
[0306] Optionally in the example considered, the method 1000 comprises a verification of the reconstruction 1080 comprising a comparison between each individual reconstructed mass spectrum and the average mass spectrum obtained from the plurality of mass spectra (as described above). Such an action makes it possible to verify that the fragments associated with this or these precursor ions exist, which improves the reliability and accuracy of the reconstruction step 1070.
[0307] [Fig.16] and [Fig.17] each illustrate an individual mass spectrum of one of the two precursor ions determined in the previous step.
[0308] This [Fig.16] comprises a first mass spectrum 701 of a first determined precursor ion (here that corresponding to the Melittin ion), and [Fig. 17] a second mass spectrum 702 of a second determined precursor ion (here that corresponding to the ion of substance P).
[0309] Optionally, these individual mass spectra 701, 702 can be compared with the average mass spectrum illustrated in [Fig.5] to verify that the combination of the individual mass spectra 701 and 702 includes all the peaks of the average mass spectrum determined previously. This is the case in this example. This can be done following the reconstruction, in the reconstruction step 1070 (step 1071).
[0310] Alternatively, the determining step 1060 may use another clustering algorithm (called a first classification algorithm). Similar to the previous reconstruction step, this other classification algorithm is applied individually to each correlation map.
[0311] In this example, the processing unit 120 uses the T-SNE algorithm as classification algorithm. Typically, the input data used by this algorithm are similar to those used by the hierarchical classification algorithm. archique (first classification algorithm) described above.
[0312] [Fig.18] illustrates a two-dimensional map provided by the T-SNE algorithm. This map comprises a multitude of points represented on the map (a plane). Each point is associated with an m / z ratio.
[0313] This algorithm is often used when the ionized sample includes a number of precursors greater than or equal to 2 and / or when there is a high probability that a fragment ion is common to at least two precursor ions of the ionized mixture. The output data is a list of pairs of coordinates (x,y) for each label (ion). These coordinates allow for a spatial visualization of the different groups highlighted.
[0314] Thus, the T-SNE algorithm makes it possible to obtain a spatial visualization of the precursor ions and fragment ions of the ionized mixture. It is also commonly combined with another clustering algorithm which in this case takes as input the map obtained by the t-SNE algorithm. In one example, the t-SNE algorithm is combined with a k-means algorithm. [Fig. 19] illustrates the use of these two algorithms successively. We see that on the map obtained by the t-SNR algorithm, two groups G1 and G2 (visible by the dotted lines) each comprising the points of the map associated with an m / z ratio. All the points included in the G1 group correspond to the fragment ions of one of the precursor ions of the ionized mixture (here the precursor ion of Melittin) and all the points included in the G2 group correspond to the fragment ions of another precursor ion of the ionized mixture (here the precursor ion of substance P).These two groups G1 and G2 can also be separated by a straight line DI determined by the k-means algorithm. It is then possible to associate with each point (a given m / z ratio) an abundance value (i.e. intensity). This abundance value is determined as previously (in the case of the hierarchical classification algorithm) from the average mass spectrum or the series of abundance values associated with this m / z ratio).
[0315] The method 1000 may also combine several clustering methods to improve the accuracy of the reconstruction. For example, the method 1000 may use a hierarchical ascending clustering algorithm and then a t-SNE algorithm combined with a k-means algorithm. Combining these methods may make it possible to verify and / or complete the reconstructed mass spectrum(s), in particular when several precursor ions include fragment ions that are close in m / z ratio.
[0316] After having reconstructed the individual mass spectrum of each detected precursor ion, the method may optionally comprise a step 1090 of analyzing the individual mass spectrum of the determined precursor ion and the associated fragment ions to characterize the chemical composition of the determined precursor ion and the chemical composition of the fragment ions associated with the precursor ion considered. Indeed, after the reconstruction step (and after the verification step if present in the method 1000), the chemical composition of the precursor ion and of the different fragments of the precursor ion are unknown (the reconstruction step provides a spectrum with m / z ratios and abundance values associated with each m / z ratio). For this purpose, the analysis step 1090 makes it possible to find the chemical composition of the fragments and of the precursor ion of each individual mass spectrum determined. This can typically be carried out by known techniques, using a database comprising (known) reference mass spectra and comparing each mass spectrum determined by the method 1000 to the reference mass spectra.
Claims
Claims
1. A method of reconstructing a mass spectrum, said method comprising the following steps: - acquisition (1010) of a plurality of mass spectra measured on an ionized sample comprising at least one precursor ion and fragment ions, - determining (1020) a plurality of series of abundance values of the ionized sample from the plurality of mass spectra, each series of abundance values of the ionized sample being determined from the plurality of mass spectra, each series of abundance values comprising N abundance values where N is an integer greater than or equal to two; - arranging (1030) the plurality of abundance value series into pairs of abundance value series, each pair of abundance value series comprising a first abundance value series and a second abundance value series, - for each pair of abundance value series, calculating (1040) a correspondence score between the first abundance value series and the second abundance value series, so as to obtain correspondence scores for all pairs of abundance value series, the correspondence scores providing a correspondence score map between all pairs of abundance value series, - determination (1060), as a function of a mass to charge ratio, of at least one precursor ion and the fragment ions associated with each determined precursor ion, by applying a grouping method to the calculated correspondence score card, - reconstruction (1070) of an individual mass spectrum of each determined precursor ion and of each of the associated fragment ions from the correspondence score card.
2. The method of claim 2, wherein the step of determining a series of abundance values comprises determining a series of ranges of mass-to-charge ratio values for the plurality of measured mass spectra of size NxM with M an integer greater than 1, and for each mass spectrum of the plurality of mass spectra, a step of integrating abundance values belonging to a same range of mass-to-charge ratio values of the mass spectrum considered, to provide, for each range of mass-to-charge ratio values of the spectrum considered, an integrated abundance value associated with a mass-to-charge ratio, the integrated abundance values over the series of ranges of mass-to-charge ratio values forming a series of abundance values of the ionized sample.
3. A method according to any one of claims 1 to 2, wherein the matching score is based on at least one of the following: - mutual information between the first series of abundance values and the second series of abundance values, said mutual information being based on an entropy of the first series of abundance values and on an entropy of the second series of abundance values; and / or - a correlation between the first series of abundance values and the second series of abundance values.
4. Method according to claim 3, in which, for the mutual information, the entropy of the first series of abundance values is an entropy normalized by a first invariance parameter as a function of a first median value determined on this first series of abundance values and the entropy of the second series of abundance values is an entropy normalized by a second invariance parameter as a function of a second median value determined on this second series of abundance values.
5. A method according to any one of claims 3 to 4, wherein, the matching score being based on mutual information, the method comprises a step of estimating the entropy of the first series of abundance values and a step of estimating the entropy of the second series of abundance values using a k-nearest neighbors algorithm.
6. The method of claim 5, wherein the k-nearest neighbor algorithm is applied to a single neighbor.
7. A method according to any one of claims 3 to 6, wherein when at least two matching scores based on two different elements are calculated in the calculating step, several score cards are provided following the calculating step, said method further comprising, for each score card, a step of filtering (1050) said score card according to a contrast of the considered score card.
8. Method according to any one of claims 1 to 7, further comprising a step of processing (1062) the score card before applying the grouping method, said processing step using at least one of the following algorithms: - standardization algorithm, - algorithm for processing the diagonal elements of the score card, - application of a square function.
9. A method according to any one of claims 1 to 8, wherein the score map is a two-dimensional matrix, the clustering method determining the at least one precursor ion and its associated fragment ions from matching scores located outside a diagonal of the score map.
10. A method according to any one of claims 1 to 9, wherein the clustering method uses at least one unsupervised learning algorithm.
11. A method according to any one of claims 1 to 10, wherein the clustering method is based on at least one of the following algorithms: - a hierarchical ascending clustering algorithm, - a k-nearest neighbors algorithm, - a k-means algorithm, - a stochastically distributed neighbor embedding algorithm.
12. Method according to any one of claims 1 to 11, further comprising, after the reconstruction step (1070), a step of analyzing (1090) the individual mass spectrum of the determined precursor ion and the associated fragment ions to characterize the chemical composition of the determined precursor ion and the chemical composition of the fragment ions associated with the precursor ion considered.
13. Method according to any one of claims 1 to 12, in which the method comprises a step of processing (1015) the acquired mass spectra comprising at least one of the following calculations: - filtering the mass spectrum considered by a Gaussian or normal kernel, - normalization of the mass spectrum considered, - smoothing of the mass spectrum considered, - returning to the baseline.
14. A method according to any one of claims 1 to 13, wherein the method comprises a step of processing (1025) the series of values abundance calculation comprising at least one of the following calculations: - filtering the given series of abundance values by a Kalman filter; - filtering the given series of abundance values by a Fourier filter, - filtering by moving average of the given series of abundance values.
15. A method according to any one of claims 1 to 14 comprising a compressed acquisition step to reduce the number of acquisitions.