Method for detecting the presence of polymers in a spectrum obtained by mass spectrometry
Patent Information
- Application Number
- EP2024704519
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2024-01-11
- Publication Date
- 2025-11-19
AI Technical Summary
Polymers in biological samples interfere with mass spectrometry signals, causing spurious signals and false identifications of bacteria, as they are not effectively distinguished from bacterial peaks, leading to suboptimal patient care in microbiology.
A method involving Lomb-Scargle period calculation and peak difference analysis to determine descriptors for polymer presence, assigning scores, and classifying spectra to detect polymers, allowing for contamination identification and potential exclusion or signal refinement to isolate bacterial signals.
Enables rapid detection of polymer contamination, preventing false identifications and ensuring accurate bacterial identification by distinguishing polymer-induced signals from bacterial signals, thus improving diagnostic reliability.
Smart Images

Figure 1.1
Abstract
Description
[0001] Title: Method for detecting the presence of polymers in a spectrum obtained by mass spectrometry
[0002] The subject of the application is a method for detecting the presence of at least one polymer in a spectrum obtained by mass spectrometry.
[0003] Technical field
[0004] This disclosure falls within the field of in vitro diagnostics.
[0005] Prior art
[0006] In microbiology, rapid identification of microorganisms in a given sample is essential for optimal patient care. Over the past decade, a technique known as MALDI-TOF mass spectrometry has greatly accelerated this identification, enabling practitioners to reliably and effectively prescribe targeted antibiotic therapy.
[0007] To properly identify the bacteria in a sample for analysis, it is usually necessary to grow the sample in a nutrient-rich environment. In such a rich environment, bacteria build polymers to store energy and / or dispose of waste. Other sources of polymers include solvents used to clean the sample.
[0008] However, polymers have a significant impact on the spectra obtained by mass spectrometry insofar as they generate parasitic signals, which can prevent the detection of peaks of interest for the identification of bacteria, and can lead to false identifications.
[0009] Summary
[0010] This disclosure improves the situation by at least partially remedying the aforementioned drawbacks.
[0011] To this end, a method is proposed for detecting the presence of a signal from at least one polymer in at least one spectrum obtained by mass spectrometry of a biological sample, comprising: a step of determining a set of at least one parameter, called a descriptor, of the presence of the polymer, a step of assigning a score to each value of each descriptor, and a step of classifying the spectrum on the basis of each descriptor and each score, so as to detect or not the presence of said at least one polymer.
[0012] Thus, thanks to the method according to the present invention, the practitioner is quickly warned if the sample is contaminated by one or more polymers, and, as a result, can act accordingly: either discard the analysis of said sample, or stop the identification process when it is determined that it is impossible to identify the bacteria contained in the sample, or to work on the signal.
[0013] According to another aspect, the step of determining a set of at least one descriptor comprises a step of calculating a period of spectrum obtained by Lomb-Scargle calculation, called Lomb-Scargle period (Ti(LSP)) and / or a period of spectrum obtained by calculating, for at least part of the peaks, the deviations, in an m / z scale, between a peak and all the other peaks, called difference period (Ti(HD)) and / or a parameter for comparing said Lomb-Scargle periods (Ti(LSP)) and difference (Ti(HD)).
[0014] In another aspect, the score associated with the Lomb-Scargle period (Ti(LSP)) depends on a ratio of a maximum value of a spectral density spectrum and an average value of said spectral density, in a given m / z interval.
[0015] In another aspect, the score associated with the period of differences depends on a prominence of a difference in peak height.
[0016] According to another aspect, the method comprises a step of refining an analysis window in an m / z scale.
[0017] According to another aspect, during the refining step, one of the limits of said window is evaluated using a period obtained by Lomb-Scargle calculation.
[0018] According to another aspect, the method comprises a step of identifying at least one bacterium in said at least one spectrum obtained by mass spectrometry if, at the end of the classification step, no polymer has been detected.
[0019] According to another aspect, the method comprises a step of alerting the detection of said at least one polymer if, at the end of the classification step, said at least one polymer has been detected.
[0020] According to another aspect, the method comprises a step of modeling the signal from a polymer.
[0021] According to another aspect, the method comprises a step of removing the signal obtained by said modeling step.
[0022] The invention also relates to a computer program comprising instructions for implementing the method as already described when this program is executed by a processor.
[0023] The invention also relates to a non-transitory recording medium readable by a computer on which is recorded a program for implementing the method as already when this program is executed by a processor.
[0024] The invention also relates to a device for detecting bacteria in a biological sample, comprising said recording medium.
[0025] Brief description of the drawings
[0026] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analyzing the attached drawings, in which:
[0027] Figure 1 shows a spectrum of a polymer-free sample of a Dermacoccus nishinomiyaensis strain obtained by mass spectrometry. Figure 2 shows a spectrum of a polymer-containing sample of a Dermacoccus nishinomiyaensis strain obtained by mass spectrometry.
[0028] Figure 3 shows a flowchart of a method for detecting the presence of polymers in a biological sample, according to the present invention.
[0029] Figure 4 shows a periodogram of a spectral density of the signal of a sample obtained by mass spectrometry as a function of a period for the implementation of the method of Figure 3, according to a first variant (abscissa: dalton (Da) - ordinate: power spectral density (Da 2 / Hz)).
[0030] Figure 5 shows the periodogram of Figure 4 in the implementation of a score calculation step.
[0031] Figure 6 shows a histogram of the peak differences for the implementation of another embodiment variant.
[0032] Figure 7 shows the histogram of Figure 6 in the implementation of a prominence calculation step.
[0033] Figure 8 shows a spectrum after a centering step of the analysis window.
[0034] Figure 9 shows a classification to distinguish between spectra with polymers on the one hand and spectra without polymer on the other.
[0035] Figure 10 shows a selection of a window for polymer signal modeling.
[0036] Figure 11 shows a repeat of the window of Figure 10.
[0037] Figure 12 shows a theoretical spectrum obtained after applying a Gaussian to the spectrum of Figure 11.
[0038] Detailed description
[0039] The method according to the present invention, referenced 1 in the figures, applies to spectra obtained by mass spectrometry.
[0040] In a preliminary step, a signal from a biological sample likely to include bacteria is acquired by mass spectrometry. To do this, in a known manner, a device such as a mass spectrometer is used, comprising a laser ionizing the biological sample. The flight time of the ions makes it possible to obtain a graph representing the intensity of the ions as a function of their m / z ratio, where m is the molecular mass of each ion and z the charge of each ion. Subsequently, we speak of signal or spectrum to designate the overall profile of the intensity as a function of the m / z ratio.
[0041] Preferably, a calibration step precedes the acquisition step. During the calibration step, reference strains are used, such as strains of the bacterium Escherichia coli for example, which produce proteins of a known mass. These reference values make it possible to move from a time-of-flight reference to the m / z reference using a regression. As already explained, polymers may be present in the sample analyzed. Their signal interferes with the signal emitted by the bacteria that one seeks to identify. The present invention applies particularly to homopolymers. Thus, the overall profile includes a signal from the bacteria and a signal due to the polymers.
[0042] Figure 1 illustrates a spectrum of a polymer-free sample of a Dermacoccus nishinomiyaensis strain obtained by mass spectrometry. Figure 2 illustrates a spectrum from a sample of the same species contaminated with polymers. As can be seen from Figure 2, the polymers generate a succession of identical patterns (peaks), the intensity of which follows a Gaussian pattern, making it more difficult to detect the signal of interest.
[0043] The repetition of the pattern is particularly related to the degree of polymerization. A polymer of degree of polymerization n and a polymer of degree of polymerization n+1 emit peaks with a distance on the m / z scale equal to monomer_mass / charge, if they have the same charge. Since the chains exist with n degrees of polymerization, they are found on the spectrum in the form of n peaks spaced by the constant distance monomer_mass / charge. If they have the same charge, this distance is constant.
[0044] As already indicated, the distribution follows a Gaussian distribution of the intensity of the polymer peaks. Depending on the culture medium, one degree of polymerization is more likely than another; this relative abundance explains the Gaussian shape formed by the polymer patterns. It is noted that, in a rich environment, bacteria tend to accumulate more amino acids and therefore form longer chains, which shifts the maximum of the Gaussian towards higher m / z.
[0045] Furthermore, for two polymer chains that are identical except for one adduct, for example sodium, the peaks of the two polymers are spaced apart by a distance equal to adduct_mass / charge. This is why a degree of polymerization is characterized by a pattern of peaks rather than an isolated peak.
[0046] The Holder has empirically shown that the majority of the signal from polymers in the m / z portion between 3000 and 17000 Da is composed of single-charged peaks.
[0047] It is noted that the method according to the present invention uses this regular pattern due to the repeated monomers and the periodicity of the signal (peaks) of the polymers, unlike spectra not contaminated by polymers, which are not periodic.
[0048] One of the objectives of the present invention is to distinguish the spectra of samples contaminated by at least one polymer. The method according to the present invention, referenced 100, comprises a step (denoted DESCR) 101 of developing at least one descriptor. Each descriptor is a parameter whose behavior makes it possible to determine the presence of polymers in this sample. A good descriptor strongly differentiates the samples containing at least one polymer in a class, called positive, from those which do not contain any, in another class, called negative. Thus, step 101 advantageously comprises a step 102 of determining the period of the signal emitted by the polymer, denoted DET-Ti. The period of interest Ti in the signal S of the polymer is denoted by the DET-Ti step.
[0049] According to a first variant, the DET-Ti step is based on a Lomb-Scargle (LSP) method, known to those skilled in the art in that it makes it possible to detect and characterize periodic signals in irregularly sampled data.
[0050] To do this, during a PER step, a periodogram P of the spectral density of the signal S is established as a function of the period T. The period of interest Ti corresponds to a maximum of the periodogram, noted Max, as illustrated in figure 4.
[0051] During the PER step, a so-called scipy LSP process is used on a regular frequency grid extending between 40 Da and 400 Da, preferably.
[0052] The method 100 also comprises a step 103 (SCORE) of defining a criterion of a score, Sc(LSP) of the period of interest Ti. The score makes it possible to quantify to what extent the spectral density of the period of interest Ti differs from those calculated for other periods.
[0053] Preferably, the score Sc(LSP) is calculated according to the following equation: Sc(LSP)=Max / Avg, where Avg is the average of the spectral density over all periods, as shown in Figure 5.
[0054] We note that this first variant can also be optimized by keeping in mind that the data are not strictly sinusoidal due in particular to the fact that:
[0055] - The signal is not perfectly periodic because of peaks that do not correspond to the polymers: the method can advantageously include a step of suppressing peaks of intensity higher than a threshold in order to avoid them disturbing the underlying periodic signal. It should be noted that this noise is in fact the signal of interest (the proteins of the bacteria) and that the periodic signal of interest is that of the polymers. We are here in the context of the search for polymers in the signal, the roles are therefore reversed;
[0056] - The intensity of the polymer patterns follows a Gaussian, and the signal is therefore only periodic up to a correction factor: the method can advantageously include a step of normalizing the intensity of each of the peaks.
[0057] According to a second variant, called the HD variant, the DET-Ti step is based on the use of inter-peak differences.
[0058] To do this, we calculate a difference E in m / z between a peak and all the following peaks. Then, as illustrated in Figure 6, we plot a histogram of the differences. Then, we select Ti as the period of interest according to the class sizes.
[0059] For example, the period of interest Ti is the class of the histogram which has the largest sample size.
[0060] Preferably, we replace the size of each class by a linear combination of the sizes of the close classes. The closer the class, the higher the associated coefficient. In other words, we can write: let En be the size associated with class n and a1, a2, a3 and R such that 0 < a1 < a2 < a3 < 1.
[0061] We re-evaluate En in the following way, in order to take into account the numbers of nearby classes: En = a1 En-2 + a2En-1 + a3En + a2En+1 + a1 En+2.
[0062] These steps make it possible to determine the period of interest Ti even when it is between two classes.
[0063] Then, a peak search is performed to isolate the peak in the histogram that has the highest prominence, where the prominence of a peak is defined as the distance between the peak's top and the surrounding minimum, as shown in Figure 7.
[0064] Alternatively, one could use absolute peak size but prominence allows one to ignore the fact that longer inter-peak differences are less common.
[0065] The period of interest, denoted Ti(HD), is calculated as Ti(HD)=(Cinf+Csup) / 2, where Cinf and Csup are the lower and upper bounds, respectively, of the class which has the most prominent signal size.
[0066] According to this variant, the score Sc(HD) is written Sc(HD)=prominence(C1). Preferably, this is the regularized prominence, as will be detailed later.
[0067] Advantageously, the method according to the present invention also comprises the determination of a descriptor for comparing the periods of interest Ti(LSP) and Ti(HD).
[0068] This descriptor is the ratio Ra = Ti(HD) / Ti(LSP) to which we associate a score of 1 if the difference Ti(LSP)-Ti(HD)>a, with a> 0 fixed empirically. The score of 1 is the worst considered. The threshold noted a makes it possible to eliminate cases where the result of one of the variants is not physically correct, which is more likely when the spectrum does not include a polymer. More generally, we carry out a test to determine which natural number the descriptor Ra is close to, using the fact that the period Ti(HD) is a multiple of the true period, while Ti(LSP) is a submultiple of the true period.
[0069] It is noted that the period comparison descriptor is not the only possible descriptor according to the present invention. In particular, the maximum intensity of the periodogram for the LSP method or the prominence for the HD method are other possible descriptors.
[0070] It is also noted that the notions of descriptor and score can be confused or distinct, the descriptor corresponding to the raw value and the score to the normalized value.
[0071] Advantageously, the method comprises a preliminary step 104, called centering, noted CENT, to center the analyzed spectrum. In other words, it is a question of correctly selecting the m / z interval of the spectrum for analysis, because the signals corresponding to the polymers are not present on the entire m / z axis.
[0072] According to a first variant, a given m / z interval is selected at the start of the signal, for example between 2000 Da and 7000 Da. According to another variant, one of the limits of the interval is optimized with respect to the Sc(LSP) score of the period of interest found by the LSP variant. By varying this limit, the period and its score are recalculated so as to obtain the highest possible score, advantageously via a method known as the "golden search".
[0073] This method is a variant of the dichotomy with an intermediate point chosen to respect the proportions of the golden ratio rather than chosen in the middle. To use it, it is necessary to have three points x1, x2, x3 which verify x1 < x2 < x3 and f(x1) > x2 < f(x3) (if we are looking for a minimum).
[0074] We implement a bracket method based on the golden search that reduces the interval until we obtain three points that verify the previous property. Once the upper bound is found, we solve a second optimization problem for the lower bound, taking care to keep a minimal interval between the lower and upper bound to ensure that we have a sufficient m / z interval to be exploitable.
[0075] Figure 8 shows a spectrum after centering. Thus, the selected m / z range is between 2000 and 12000 Da. And, the signal that corresponds to the polymers is indeed in the selected range.
[0076] The CENT centering step allows to efficiently center the studied area on the polymer peaks and to significantly improve the results and the calculation time. In particular, this step makes it possible to avoid performing the calculations on all the peaks but only on a fraction of them.
[0077] The method comprises a step 105 of classifying the spectra, denoted CLASS, into at least one class or group of positive spectra, i.e. which contain peaks generated by polymers and into a class or group of negative spectra, i.e. considered as not containing peaks generated by polymers.
[0078] According to a first variant, an unsupervised method is used, preferably “K-means clustering”.
[0079] The principle of K-means clustering is to optimize groups so that each member of the group is as close as possible to all the others, in the descriptor space.
[0080] One of the resulting groups contains most of the positive spectra, while the negative spectra are distributed in at least one other group. Classifying the negative spectra into more than two groups allows us to account for the diversity of the group's spectra.
[0081] This first unsupervised classification has the significant advantage of not having any a priori assumptions about the data: it creates groups blindly, without knowing the classes of the different spectra.
[0082] According to a second variant, a supervised method is used, advantageously called support vector machine (SVM). Preferably, a linear SVM method is used. This method consists of making a decision from a linear combination of descriptors.
[0083] In Figure 9, which represents the data in the descriptor space, we find a hyperplane (solid line in the figure) which maximizes a margin between the class of positive spectra (dashed line) and the class of negative spectra (mixed dashed line), that is to say the gap between the solid line and those in dashed line.
[0084] Note that, for classification, the data that is classified can be considered as being the spectrum summarized by the descriptor, and its score.
[0085] Preferably, prior to the classification step, each descriptor is regularized, according to the following equation, for a given descriptor X: regularized = [X - min(Xentrainement)] / [max(Xentrainement) - min(Xentrainement)].
[0086] By subtracting the minimum value taken by this descriptor in a training dataset and dividing by the difference between this minimum value and their maximum, we allow each descriptor to take comparable values. We can thus classify the data without favoring a descriptor solely on the basis of its order of magnitude.
[0087] As already indicated, if the spectrum is considered, at the end of the classification step, as not containing a polymer, the method 100 comprises a step of identifying the bacteria or bacteria present in the sample, if any. This step notably comprises a comparison of the spectrum with spectra belonging to a bank of referenced spectra, each referenced spectrum being that of a given bacterium.
[0088] If the spectrum is considered, at the end of the classification step, as containing one or more polymers, the method 100 optionally comprises a step of indicating, for example by visual and / or audible alert, that the signal comprises at least one polymer. This contamination indication step may induce a stoppage of the method 100, and no bacteria identification is carried out.
[0089] According to another variant, the method 100 according to the present invention optionally comprises a step 106, noted SUPP, of suppressing the signal emitted by the polymers.
[0090] Thus, for a given spectrum, when the classification step has considered that polymers are present, the method 100 makes it possible to eliminate the peaks of the polymers so as to isolate the signal emitted by proteins of the bacteria.
[0091] The suppression step is preceded by a step 107 of modeling MODEL of the signal emitted by the polymers.
[0092] During this step, a pattern of polymer peaks is isolated.
[0093] We start by dividing the signal into successive windows of the size of a period, to isolate one polymer pattern per window (with a potential shift). Preferably, we use the period Ti(HD), because, on the one hand, it is generally more precise, and, on the other hand, it also makes it possible to obtain a multiple of the period, and therefore to isolate one or more polymer patterns.
[0094] On the contrary, the period Ti(LSP) is a sub-multiple of the period, and its use could lead to isolating only part of the pattern per window.
[0095] Note that the precision of the period is very important to avoid inducing a shift along the m / z axis; the windows would then no longer be centered on a polymer pattern.
[0096] Once the best window has been determined (figure 10), it is copied onto the entire spectrum (figure 11) and a Gaussian is applied to the copied spectrum (figure 12).
[0097] We then obtain a theoretical signal from the polymers, which must then be subtracted from the overall signal to eliminate it.
[0098] Thus, the method 100 according to the present invention ensures reliable and rapid classification of spectra, making it possible to distinguish between spectra contaminated by polymers and spectra without polymers.
[0099] It is also possible, using process 100, to isolate the signal emitted by bacteria, by eliminating that of the polymers.
[0100] As already evident from the above description, the invention has many advantages, among which: a / the method is purely digital, it is therefore very inexpensive in time and money compared to a biological or chemical method, b / depending on the use: the detection method can be used in quality control, or serve as a preliminary step before the step of removing the polymer peaks, c / the invention can be directly used to determine which polymer is present in a mass spectrometry sample, it can therefore be applied to other themes than the identification of bacteria.
Claims
Claims 1. Method for detecting the presence of a signal from at least one polymer in at least one spectrum obtained by mass spectrometry of a biological sample, comprising: a step (101) of determining a set of at least one parameter, called descriptor, of the presence of the polymer, a step (103) of assigning a score to each value of each descriptor, and a step (105) of classifying the spectrum on the basis of each descriptor and each score, so as to detect or not the presence of said at least one polymer.
2. Method according to claim 1, in which the step (101) of determining a set of at least one descriptor comprises a step of calculating a period of spectrum obtained by Lomb-Scargle calculation, called Lomb-Scargle period (Ti(LSP)) and / or a period of spectrum obtained by calculating, for at least part of the peaks, the deviations, in an m / z scale, between a peak and all the other peaks, called difference period (Ti(HD)) and / or a parameter for comparing said Lomb-Scargle periods (Ti(LSP)) and difference (Ti(HD)).
3. Method according to claim 2, in which the score associated with the Lomb-Scargle period (Ti(LSP)) depends on a ratio of a maximum value of a spectral density and an average value of said spectral density, in a given m / z interval.
4. Method according to one of claims 2 or 3, in which the score associated with the period of the differences depends on a prominence of a difference in peak height.
5. Method according to one of the preceding claims, comprising a step (104) of refining an analysis window in an m / z scale.
6. Method according to the preceding claim, in which, during the refining step, one of the limits of said window is evaluated using a period obtained by Lomb-Scargle calculation.
7. Method according to one of the preceding claims, comprising a step of identifying at least one bacterium in said at least one spectrum obtained by mass spectrometry if, at the end of the classification step, no polymer has been detected.
8. Method according to one of the preceding claims, comprising a step of alerting the detection of said at least one polymer if, at the end of the classification step, said at least one polymer has been detected.
9. Method according to one of the preceding claims, comprising a step (107) of modeling the signal from a polymer.
10. Method according to the preceding claim, comprising a step (108) of removing the signal obtained by said modeling step.
11. Computer program comprising instructions for implementing the method according to one of the preceding claims when this program is executed by a processor.
12. Non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to one of claims 1 to 10 when this program is executed by a processor.