Method for detecting the presence of polymers in spectra obtained by mass spectrometry - Patent Application 20070122997
The method addresses polymer interference in mass spectrometry by using Lomb-Scargle and differential periods for scoring and classification, enabling reliable bacterial identification by distinguishing between contaminated and clean spectra and removing polymer signals.
Patent Information
- Application Number
- JP2025540730
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2024-01-11
- Publication Date
- 2026-02-03
AI Technical Summary
Polymers in biological samples interfere with mass spectrometry signals, leading to erroneous bacterial identification due to parasitic signals.
A method for detecting polymer presence in mass spectrometry spectra using Lomb-Scargle periods and differential periods, with scoring and classification to alert practitioners of contamination, allowing for sample discard or stopping the identification process, and optionally removing polymer signals.
Enables reliable and rapid differentiation between polymer-contaminated and clean spectra, ensuring accurate bacterial identification by removing polymer interference.
Smart Images

Figure 2026504059000001_ABST
Abstract
Description
[Technical Field]
[0001] The subject of this patent application is a method for detecting the presence of at least one polymer in a spectrum obtained by mass spectrometry.
[0002] The present disclosure relates to the field of in vitro diagnostics.
[0003] Prior art In microbiology, rapid identification of microorganisms in a given sample is essential for optimal patient management. Over the past decade, a technique known as MALDI-TOF mass spectrometry has greatly accelerated this identification, enabling practitioners to reliably and effectively prescribe targeted antibiotic therapy.
[0004] To correctly identify bacteria in a sample to be analyzed, the sample must generally be cultured in a medium where the bacteria build polymers to store energy or process waste products. Polymers can also come from other sources, such as solvents used to wash the sample.
[0005] However, these polymers have a significant impact on the spectra obtained by mass spectrometry insofar as they generate parasitic signals, which can interfere with the detection of peaks of interest for bacterial identification, leading to erroneous identification.
[0006] overview The present disclosure improves the situation by at least partially overcoming the aforementioned drawbacks.
[0007] To this end, a method is provided for detecting the presence of a signal produced by at least one polymer in at least one spectrum obtained by mass spectrometry of a biological sample, the method comprising the steps of determining a set of at least one parameter (called descriptors) for the presence of the polymer, assigning a score to each value of each descriptor, and classifying the spectrum based on each descriptor and each score in order to detect the presence or absence of said at least one polymer.
[0008] Thus, according to the method of the present invention, if a sample is contaminated with one or more polymers, the practitioner is quickly alerted and can act accordingly: the analysis of the sample is discarded, or the identification process is stopped if it is determined that it is not possible to identify the bacteria contained in the sample, or the signal is investigated.
[0009] According to another embodiment, the step of determining the set of at least one descriptor comprises calculating the spectral periods obtained by Lomb-Scargle calculations (called Lomb-Scargle periods (Ti(LSP)) and / or the spectral periods obtained by calculating, for at least some of the peaks, the separation between one peak and all the other peaks on the m / z scale (called differential periods (Ti(HD)) and / or parameters allowing a comparison of said Lomb-Scargle periods (Ti(LSP)) and the differential periods (Ti(HD)).
[0010] According to another embodiment, the score related to the Lomb-Scargle period (Ti(LSP)) depends on the ratio between the maximum value of the spectral density and the mean value of said spectral density in a given m / z interval.
[0011] According to another embodiment, the score associated with the difference period depends on the prominence of the peak height difference.
[0012] According to another embodiment, the method comprises fine-tuning the analysis window on the m / z scale.
[0013] According to another embodiment, in the fine-tuning step, one of the limits of said window is evaluated using the period obtained by the Lomb-Scargle calculation.
[0014] According to another embodiment, the method comprises a step of identifying at least one bacterium in said at least one spectrum obtained by mass spectrometry if no polymer is detected at the end of the classification step.
[0015] According to another embodiment, the method comprises the step of alerting to the detection of said at least one polymer if said at least one polymer is detected at the end of the classification step.
[0016] According to another embodiment, the method includes modeling the signal produced by the polymer.
[0017] According to another embodiment, the method includes the step of removing (108) the signal obtained by the modeling step.
[0018] The invention also provides a computer program comprising instructions for carrying out the above-described method when the program is executed by a processor.
[0019] The present invention also provides a non-transitory computer-readable recording medium having a program recorded thereon for performing the above-described method when the program is executed by a processor.
[0020] The present invention also provides an apparatus for detecting bacteria in a biological sample, comprising the recording medium.
[0021] Other features, details, and advantages will become apparent upon reading the following detailed description and examining the accompanying drawings. [Brief explanation of the drawings]
[0022] [Figure 1]The spectrum of a sample of a polymer-free strain of Dermacoccus nishinomiyaensis obtained by mass spectrometry is shown. [Figure 2] 1 shows the spectrum of a sample of a polymer-containing strain of Dermacoccus nishinomiyaensis obtained by mass spectrometry. [Figure 3] 1 shows a flow chart of a method for detecting the presence of a polymer in a biological sample according to the present invention. [Figure 4] 4 shows a periodogram of the spectral density of the signal of a sample obtained by mass spectrometry as a function of period (x-axis: Daltons (Da) - y-axis: power spectral density (Da2 / Hz)) for carrying out the method of FIG. 3 according to a first variant. [Figure 5] The periodogram of FIG. 4 is shown when performing the score calculation process. [Figure 6] 10 shows a histogram of peak differences for another variation of the embodiment. [Figure 7] 7 shows the histogram of FIG. 6 when performing the prominence calculation step. [Figure 8] 10 shows the spectrum after the step of centering the analysis window. [Figure 9] The classification that allows the polymer-free spectra to be distinguished from the polymer-containing spectra is shown. [Figure 10] Window selection for modeling macromolecular signals is shown. [Figure 11] The window repetition in FIG. 10 is shown. [Figure 12] The theoretical spectrum obtained after applying a Gaussian to the spectrum of FIG. 11 is shown. DETAILED DESCRIPTION OF THE INVENTION
[0023] The method according to the invention, designated in the drawing by reference number 1, is applied to spectra obtained by mass spectrometry.
[0024] In a preliminary step, signals from biological samples that may contain bacteria are obtained by mass spectrometry. To do this, a device such as a mass spectrometer is used, which includes a laser that ionizes the biological sample in a known manner. The flight time of the ions allows obtaining a graph of the intensity of the ions as a function of the m / z ratio, where m is the molecular weight of each ion and z is the charge of each ion. In the following, the terms signal or spectrum are used to refer to this overall profile of intensity as a function of the m / z ratio.
[0025] Preferably, the acquisition step is preceded by a calibration step. During the calibration step, reference strains (e.g., strains such as E. coli) that produce proteins of known mass are used. These reference values allow the use of regression to switch from a time-of-flight reference system to an m / z reference system.
[0026] As explained above, polymers may be present in the analyzed sample. The signal of the polymer adds noise to the signal emitted by the bacteria we are trying to identify. The present invention applies most specifically to homopolymers. The overall profile therefore includes the signal produced by the bacteria and the signal due to the polymer.
[0027] Figure 1 shows the spectrum of a polymer-free strain of Dermococcus nishinomiyaensis obtained by mass spectrometry. Figure 2 shows the spectrum produced by a polymer-contaminated sample of the same species. As shown in Figure 2, the polymer produces a series of identical patterns (peaks), the intensities of which follow a Gaussian distribution, making it more difficult to detect the signal of interest.
[0028] The repetition of the pattern is particularly related to the degree of polymerization: a polymer of degree n and a polymer of degree n+1 will emit peaks equally separated by the monomer_mass / charge on the m / z scale if they have the same charge. As a chain exists at degree of polymerization n, it will be represented in the spectrum by n peaks spaced a constant monomer_mass / charge distance apart. If they have the same charge, this distance is constant.
[0029] As already shown, the intensities of the polymer peaks follow a Gaussian distribution. In some media, some degrees of polymerization are more prevalent than others, and this relative abundance explains the Gaussian shape formed by the polymer pattern. Note that in media, bacteria tend to accumulate more amino acids and therefore tend to form longer chains, which shifts the maximum of the Gaussian distribution to higher m / z.
[0030] Furthermore, for two chains of a polymer that are identical except for one adduct (e.g., sodium), the peaks of the two polymers are separated by a separation equal to the adduct mass / charge. For this reason, the degree of polymerization is characterized by the pattern of peaks, rather than by one isolated peak.
[0031] It has been empirically shown by the owner that in the m / z segment between 3000 and 17000 Da, the polymer signal consists mainly of singly charged peaks.
[0032] It should be noted that the method according to the present invention takes advantage of the regular motifs resulting from the repeating monomers and the periodicity of the polymer signals (peaks) (in contrast to spectra without polymer contamination, which are not periodic).
[0033] One of the objectives of the present invention is to identify the spectra of samples contaminated with at least one polymer. The method according to the present invention, designated by the reference numeral 100, comprises a step 101 of generating at least one descriptor (denoted DESCR). Each descriptor is a parameter whose behavior makes it possible to determine the presence of a polymer in a sample. A good descriptor clearly distinguishes samples containing at least one polymer from those not containing a polymer, with the former falling into one class, called the positive class, and the latter falling into another class, called the negative class.
[0034] Step 101 therefore advantageously comprises a step 102 of determining the period of the signal emitted by the polymer, this step being denoted DET-Ti. The period determined by step DET-Ti is denoted as the period of interest Ti in the signal S of the polymer.
[0035] According to a first variant, the process DET-Ti is based on the LSP method (LSP stands for Lomb-Scargle Periodogram), a method of this type known in the art that allows the detection and characterization of periodic signals in irregularly sampled data.
[0036] For this purpose, a periodogram P of the spectral density of the signal S as a function of the period T is generated in a step PER. The period of interest Ti corresponds to the maximum value (denoted Max) of the periodogram, as shown in FIG. 4.
[0037] In the process PER, the so-called SciPy LSP method is applied to a regular frequency grid, preferably between 40 Da and 400 Da.
[0038] The method 100 also includes a step 103 (SCORE) of defining a score criterion, the score being Sc(LSP) of the period of interest Ti, which makes it possible to quantify how different the spectral density of the period of interest Ti is from the spectral densities calculated for other periods.
[0039] The score Sc(LSP) is preferably calculated using the following formula: Sc(LSP)=Max / Avg where Avg is the average of the spectral density over all periods (see Figure 5).
[0040] It should be noted that this first variant can also be optimized keeping in mind that the data are not strictly sinusoidal, in particular due to the fact that: the signal is not perfectly periodic due to peaks that do not correspond to polymers, the method may advantageously include a step of removing peaks with an intensity higher than a threshold to prevent noise being added to the underlying periodic signal. Note that this noise is actually the signal of interest (bacterial proteins), and the periodic signal of interest is the signal of the polymer. Here, the roles are reversed, since the context is the search for polymers in the signal; Since the intensities of the polymer pattern follow a Gaussian distribution, the signal is periodic only within a correction factor: the method may advantageously comprise a step of normalizing the intensity of each peak.
[0041] According to a second variant (called HD variant), the process DET-Ti is based on using the differences between peaks.
[0042] For this purpose, the separation E(m / z) between the peak and all subsequent peaks is calculated. A histogram of the differences is then plotted, as shown in Figure 6. The periods of interest Ti are then selected according to their class membership.
[0043] For example, the period of interest Ti is the class of the histogram with the highest membership.
[0044] Each class's membership is preferably replaced by a linear combination of the memberships of neighboring classes: the closer the classes, the higher the associated coefficient.
[0045] In other words, you can write: Let En be the membership associated with class n, and a1, a2, a3, ∈ R, be 0 <a1<a2<a3<1とする。
[0046] To take into account neighborhood class membership, En is re-evaluated as follows: En=a1En-2+a2En-1+a3En+a2En+1+a1En+2.
[0047] These procedures make it possible to determine the period of interest Ti even if it lies between two classes.
[0048] Peak detection then makes it possible to identify the peak of the histogram that has the greatest prominence, where the prominence of a peak is defined as the distance between the top of the peak and the minimum value surrounding it (see Figure 7).
[0049] Alternatively, absolute peak size can be used, but prominence allows one to ignore the fact that differences between longer peaks are less common.
[0050] The period of interest (denoted Ti(HD)) is Ti(HD)=(Cinf+Csup) / 2 where Cinf and Csup are the lower and upper bounds, respectively, of the signal class with the highest membership.
[0051] According to this variant, the score Sc(HD) is Sc(HD)=Prominence (C1) Preferably, the degree of prominence is regularized, which will be explained in more detail below.
[0052] The method according to the invention advantageously also comprises determining a descriptor that allows a comparison of the periods Ti(LSP) and Ti(HD) of interest.
[0053] This descriptor is the ratio Ra = Ti(HD) / Ti(LSP), and is associated with a score of 1 if the difference Ti(LSP)-Ti(HD)>a is empirically set to a>0. A score of 1 is the worst possible score. The threshold indicated by "a" allows us to eliminate cases where the results of one of the variations are physically incorrect, which is likely to occur when the spectrum does not contain a polymer. More generally, tests are performed to determine which natural number the descriptor Ra approaches, using the fact that the period Ti(HD) is a multiple of the true period, while Ti(LSP) is a partial multiple of the true period.
[0054] It should be noted that the descriptor that allows the comparison of periods is not the only descriptor possible according to the invention: in particular, the maximum intensity of the periodogram obtained with the LSP method or the salience obtained with the HD method could be other descriptors.
[0055] Also, note that the annotations for the descriptors and scores can be the same or different, with the descriptors corresponding to raw values and the scores corresponding to normalized values.
[0056] The method advantageously includes a preliminary step 104 (called the centering step, denoted CENT) for centering the analyzed spectrum. In other words, the signals corresponding to the polymer are not present across the entire m / z axis, so the question is to correctly select the m / z interval of the spectrum to be analyzed.
[0057] According to a first variant, a predetermined m / z interval is selected at the beginning of the signal, for example an interval lying between 2000 Da and 7000 Da.
[0058] According to another variant, one of the interval limits is optimized with respect to the score Sc(LSP) of the period of the object of interest found by the LSP method. By changing this limit, the period and its score are recalculated, and advantageously, the highest possible score is obtained by the golden-search method.
[0059] This method is a variant of the binary search method in which the midpoint is selected to respect the ratio of the golden number, rather than being selected to be in the middle. To use this, three points x1, x2, x3 are required that respect x1 < x2 < x3 and f(x1) > x2 < f(x3) (when the minimum value is sought).
[0060] Adopt a bracket method based on the golden section search and reduce the interval until three points that respect the above properties are obtained. Once the upper limit is found, a second optimization problem is solved for the lower limit while maintaining a minimum interval between the lower and upper limits to ensure a sufficiently large size to make the m / z interval available.
[0061] FIG. 8 shows the spectrum after centering. Thus, the selected m / z interval is between 2000 and 12000 Da. Furthermore, the signal corresponding to the polymer is actually in the selected interval.
[0062] The centering step CENT can efficiently center the research area on the polymer peak, significantly improving the results and the calculation time. In particular, this step makes it possible to perform calculations only on the peak fractions, rather than on all peaks.
[0063] The method includes step 105 (denoted CLASS) of classifying the spectra into at least one class or group of positive spectra (i.e., spectra that contain peaks produced by the polymer) and a class or group of negative spectra (i.e., spectra that are not believed to contain peaks produced by the polymer).
[0064] According to a first variant, an unsupervised method is used, preferably K-means clustering.
[0065] The principle of K-means clustering is to optimize a group so that each member of the group is as close as possible to all other members in the descriptor space.
[0066] One of the resulting groups contains most of the positive spectra, and the negative spectra are distributed among at least one other group. By dividing the negative spectra into more than two groups, it becomes possible to explain the diversity of the group spectra.
[0067] This unsupervised first classification has the non-negligible advantage that it does not require any prior knowledge of the data: it blindly generates groups without knowing the different spectral classes.
[0068] According to a second variant, a supervised method is used, advantageously an SVM method (SVM stands for Support Vector Machine).
[0069] Preferably, a linear SVM method is used, which consists in making decisions based on a linear combination of descriptors.
[0070] In Figure 9, which shows data in descriptor space, the hyperplane (solid line in the figure) maximizes the margin between the positive spectral class (dotted line) and the negative spectral class (dashed line-dotted line), i.e., the separation between the solid line, dotted line, and dashed line and dashed line.
[0071] Note that for the purposes of classification, the classified data can be viewed as spectra summarized by descriptors and their scores.
[0072] Preferably, before the classification step, each descriptor is regularized for a given descriptor X according to the following formula: Xregularized=[X-min(Xtraining)] / [max(Xtraining)-min(Xtraining)].
[0073] By subtracting the minimum value obtained by this descriptor in the training data set and dividing by the difference between this minimum and their maximum, each descriptor takes on a comparable value, thus making it possible to classify the data based solely on its magnitude, without favoring one descriptor over another.
[0074] As already indicated, if the spectrum is deemed to be free of polymer at the end of the classification step, the method 100 includes a step of identifying one or more bacteria, if any, present in the sample, which step in particular involves comparing the spectrum with spectra belonging to a bank of reference spectra, each reference spectrum being the spectrum of a given bacterium.
[0075] At the end of the classification step, if the spectrum is deemed to contain one or more polymers, method 100 optionally includes indicating, for example by a visual and / or audio warning, that the signal contains at least one polymer. This step indicating contamination can cause method 100 to stop and no bacterial identification is made.
[0076] According to another variant, the method 100 according to the invention optionally comprises a step 106 (denoted RMVL) of removing the signal emitted by the polymer.
[0077] Thus, for a given spectrum, if the classification process indicates that a polymer is present, method 100 allows for removal of the polymer peak in order to differentiate the signal emitted by the bacterial protein.
[0078] The removal step is preceded by a step 107 (MODEL) of modeling the signal emitted by the polymer.
[0079] In this step, the pattern of polymer peaks is resolved.
[0080] First, the signal is divided into successive windows of one period size, and one polymer pattern is resolved per window (possibly shifted). The use of the period Ti(HD) is preferred, as it is generally relatively accurate and multiples of the period can also be obtained, allowing resolution of one or more polymer patterns.
[0081] In contrast, the period Ti(LSP) is a submultiple of the period, and using this only part of the pattern can be resolved per window.
[0082] Note that the accuracy of the period is very important to avoid inducing a shift along the m / z axis, which would otherwise result in the window not being centered on the polymer pattern.
[0083] Once the optimal window is determined (Figure 10), it is copied to the entire spectrum (Figure 11), and a Gaussian is applied to the copied spectrum (Figure 12).
[0084] This gives a theoretical polymer signal which then has to be subtracted out from the total signal.
[0085] The method 100 according to the invention thus ensures a reliable and rapid classification of spectra, making it possible to distinguish between spectra contaminated with polymers and those free of polymers.
[0086] Method 100 also allows for differentiation of signals emitted by bacteria by removing signals due to polymers.
[0087] As is already apparent from the above description, the present invention has many advantages, such as: a. The method is purely numerical and therefore very cheap in time and cost compared to biological or chemical methods. b. Depending on the application, this detection method can be used for quality control or as a preliminary step before removing the polymer peak. c. The present invention is applicable in contexts other than bacterial identification, as it can be used directly to determine the polymers present in mass spectrometry samples.
Claims
1. 1. A method for detecting the presence of a signal produced by at least one polymer in at least one spectrum obtained by mass spectrometry of a biological sample, the method comprising the steps of: determining (101) a set of at least one parameter (called descriptors) for the presence of the polymer; assigning (103) a score to each value of each descriptor; and classifying (105) the spectrum based on each descriptor and each score to detect the presence or absence of the at least one polymer.
2. 2. The method of claim 1, wherein the step (101) of determining the set of at least one descriptor comprises calculating spectral periods obtained by Lomb-Scargle calculations (called Lomb-Scargle periods (Ti(LSP)) and / or spectral periods obtained by calculating, for at least some of the peaks, the separation between one peak and all other peaks on the m / z scale (called difference periods (Ti(HD)) and / or parameters that allow a comparison of the Lomb-Scargle periods (Ti(LSP)) and the difference periods (Ti(HD)).
3. 3. The method of claim 2, wherein the score associated with a Lomb-Scargle period depends on the ratio of the maximum value of the spectral density to the mean value of said spectral density in a given m / z interval.
4. 4. The method of claim 2 or 3, wherein the score associated with the difference period depends on the prominence of the peak height difference.
5. 5. The method of claim 1, further comprising the step of fine-tuning (104) the analysis window in the m / z scale.
6. The method of claim 5 , wherein in the fine-tuning step, one of the window limits is evaluated using a period obtained by a Lomb-Scargle calculation.
7. 7. The method of claim 1, further comprising identifying at least one bacterium in the at least one spectrum obtained by mass spectrometry if no polymer is detected at the end of the classification step.
8. 8. The method of claim 1, further comprising the step of issuing an alert for the detection of said at least one polymer if said at least one polymer is detected at the end of the classification step.
9. 9. The method of any one of claims 1 to 8, comprising the step of modeling (107) the signal produced by the polymer.
10. 10. The method of claim 9, further comprising the step of removing (108) signals obtained by the modeling step.
11. A computer program comprising instructions for carrying out the method according to any one of claims 1 to 10 when the program is executed by a processor.
12. A non-transitory computer-readable recording medium having a program recorded thereon for performing the method according to any one of claims 1 to 10 when the program is executed by a processor.