Method for detecting presence of polymer in spectrum obtained by mass spectrometry
By calculating the presence parameters and spectrum classification of polymers, the problem of polymer signals interfering with bacterial recognition in mass spectrometry is solved, and rapid and accurate bacterial detection and polymer signal removal are achieved.
Patent Information
- Application Number
- CN202480007407.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2024-01-11
- Publication Date
- 2025-08-19
AI Technical Summary
Parasitic signals generated by polymers in mass spectrometry affect bacterial recognition, resulting in difficulty in misidentification and detection.
By determining the presence parameters of the polymer, compute the Lomb-Scargle period and the differential period, use K-mean clustering or support vector machines for spectrum classification, remove the polymer signal, and identify bacteria.
Rapidly distinguish between polymer-contaminated and polymer-free spectrum, ensure accurate bacterial recognition and reduce time and cost.
Smart Images

Figure CN120513484A_ABST
Abstract
Description
[0001] The subject of the present patent application is a method for detecting the presence of at least one polymer in a spectrum obtained by mass spectrometry. Technical Field
[0002] The present disclosure relates to the field of in vitro diagnostics. Background Art
[0003] In microbiology, rapid identification of microorganisms in a given sample is crucial for optimizing patient management. Over the past decade, a technology called MALDI-TOF mass spectrometry has greatly accelerated this identification, enabling practitioners to reliably and effectively prescribe targeted antibiotic therapies.
[0004] To correctly identify bacteria in a sample to be analyzed, it is often necessary to grow the sample in a culture medium. In this culture medium, the bacteria build polymers to store energy and / or dispose of waste. Other sources of polymers exist, such as solvents used to clean the sample.
[0005] However, these polymers have a non-negligible impact on the spectrum obtained by mass spectrometry, since they generate parasitic signals, which may hinder the detection of peaks of interest for identifying bacteria and may lead to false identifications. Summary of the Invention
[0006] The present disclosure improves the situation by at least partially overcoming the above disadvantages.
[0007] To this end, a method is provided for detecting the presence of a signal generated by at least one polymer in at least one spectrum obtained by mass spectrometry of a biological sample, comprising: a step of determining a set of at least one parameter of the presence of a polymer, called descriptors, a step of assigning a score to each value of each descriptor, and a step of classifying the spectrum based on each descriptor and each score in order to detect the presence or absence of said at least one polymer.
[0008] Thus, with the method according to the invention, the practitioner is quickly alerted whether a sample is contaminated with one or more polymers and can act accordingly: abandon the analysis of the sample, or stop the identification process when it is determined that it is impossible to identify the bacteria contained in the sample, or continue processing the signal.
[0009] According to another aspect, the step of determining at least one set of descriptors comprises the step of calculating: a spectral period obtained by a Lomb-Scargle calculation, called the Lomb-Scargle period (Ti(LSP))), and / or a spectral period obtained by calculating, for at least some of the peaks, the interval between a peak and all other peaks on the m / z scale, called the differential period (Ti(HD)), and / or a parameter allowing comparison of said Lomb-Scargle period (Ti(LSP)) and the differential period (Ti(HD)).
[0010] According to another aspect, the fraction associated with the Lomb-Scargle period (Ti(LSP)) depends on the ratio of the maximum value of the spectral density to the average value of said spectral density in a given m / z bin.
[0011] According to another aspect, the score associated with the differential period depends on the prominence of the peak height difference.
[0012] According to another aspect, the method comprises the step of refining the analysis window on the m / z scale.
[0013] According to another aspect, in the refinement step, a limit of the window is evaluated using the period obtained by Lomb-Scargle calculation.
[0014] According to another aspect, if no polymer is detected at the end of the classification step, the method comprises a step of identifying at least one bacterium in said at least one spectrum obtained by mass spectrometry.
[0015] According to another aspect, if said at least one polymer has been detected at the end of the classification step, the method comprises the step of issuing an alarm of the detection of said at least one polymer.
[0016] According to another aspect, the method includes the step of modeling the signal generated by the polymer.
[0017] According to another aspect, the method comprises the step of removing the signal obtained by said modeling step.
[0018] The present invention also provides a computer program comprising instructions for implementing the above method when the program is executed by a processor.
[0019] The present invention also provides a non-transitory computer-readable recording medium having a program recorded thereon, which is used to implement the above method when the program is executed by a processor.
[0020] The present invention also provides a device for detecting bacteria in a biological sample, comprising the recording medium. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Other features, details, and advantages will become apparent upon reading the following detailed description and analyzing the accompanying drawings, in which:
[0022] Figure 1 Shown is the spectrum of a polymer-free strain sample of Dermacoccus nishinomiyaensis obtained by mass spectrometry.
[0023] Figure 2 Shown is the spectrum of a polymer-containing strain sample of Dermatococcus nishinomiya obtained by mass spectrometry.
[0024] Figure 3 A flow chart showing a method according to the present invention for detecting the presence of a polymer in a biological sample is shown.
[0025] Figure 4 According to the first variant embodiment Figure 3 Periodogram of the spectral density of the signal of the sample obtained by mass spectrometry as a function of the period (abscissa: Dalton (Da) - ordinate: power spectral density (Da) 2 / Hz)).
[0026] Figure 5 Shows that when implementing the score calculation step Figure 4 Periodogram of .
[0027] Figure 6 A histogram of peak differences is shown when another variant embodiment is implemented.
[0028] Figure 7 Shows the implementation of the protrusion calculation step Figure 6 Histogram of .
[0029] Figure 8 The spectrum after a step of centering the analysis window is shown.
[0030] Figure 9 A classification is shown which makes it possible to distinguish between polymer-free spectra on the one hand and polymer-containing spectra on the other hand.
[0031] Figure 10 Shown is the selection of the window used to model the polymer signal.
[0032] Figure 11 Shown Figure 10 Duplicate of the window.
[0033] Figure 12 Shown in Figure 11 Theoretical spectrum obtained by applying a Gaussian function to the spectrum of . DETAILED DESCRIPTION
[0034] The method according to the invention is referenced 1 in the figures and is applied to spectra obtained by mass spectrometry.
[0035] In a preliminary step, a signal of a biological sample that may contain bacteria is obtained by mass spectrometry. To this end, an apparatus such as a mass spectrometer (which includes a laser for ionizing the biological sample) is used in a known manner. The flight time of the ions makes it possible to obtain a graph representing the ion intensity as a function of its m / z ratio, where m is the molecular weight of each ion and z is the charge of each ion. hereinafter, the term "signal" or "spectrum" is used to refer to the overall profile of this intensity as a function of the m / z ratio.
[0036] Preferably, a calibration step is performed before the acquisition step. In the calibration step, reference strains are used, such as Escherichia coli strains, which produce proteins of known mass. These reference values make it possible to convert the time-of-flight reference system to an m / z reference system using regression.
[0037] As mentioned previously, polymers may be present in the sample being analyzed. Their signal adds noise to the signal emitted by the bacteria being identified. The present invention is particularly applicable to homopolymers. Therefore, the overall spectrum contains both the signal generated by the bacteria and the signal due to the polymers.
[0038] Figure 1 Illustrated is the spectrum of a polymer-free strain sample of Dermatococcus nishinomiya obtained by mass spectrometry. Figure 2 Spectra generated from samples of the same species contaminated with polymer are shown. Figure 2 As shown, the polymer generates a series of identical patterns (peaks) whose intensities follow a Gaussian distribution, which makes it more difficult to detect the signal of interest.
[0039] The repetition of the pattern is particularly dependent on the degree of polymerization. If a polymer with degree of polymerization n and a polymer with degree of polymerization n+1 have the same charge, they will emit peaks separated by a distance equal to monomer_mass / charge on the m / z scale. Since chains exist with n degrees of polymerization, they are represented in the spectrum by n peaks separated by a constant distance of monomer_mass / charge. If they have the same charge, this distance is constant.
[0040] As mentioned previously, the intensity of the polymer peak follows a Gaussian distribution. Depending on the culture medium, one degree of polymerization is more likely than another, and this relative abundance explains the Gaussian shape formed by the polymer pattern. It is important to note that in culture, bacteria tend to accumulate more amino acids, thus forming longer chains, which shifts the maximum of the Gaussian distribution toward higher m / z values.
[0041] Furthermore, for two polymer chains that are identical except for one adduct (e.g. sodium), the peaks for the two polymers are separated by a distance equal to adduct_mass / charge. For this reason, a degree of polymerization is characterized by a peak pattern rather than an isolated peak.
[0042] The patentee has empirically shown that in the m / z range between 3000 and 17000 Da, the polymer signal is primarily composed of singly charged peaks.
[0043] It is to be noted that the method according to the invention exploits this regular pattern caused by the repeating monomers and the periodicity of the signals (peaks) of the polymer, in contrast to the aperiodic spectrum which is not contaminated by the polymer.
[0044] One of the goals of the present invention is to distinguish spectra of samples contaminated with at least one polymer. The method according to the present invention (referenced 100) comprises a step of generating at least one descriptor (labeled DESCR) 101. Each descriptor is a parameter whose behavior allows the presence of a polymer in a sample to be determined. Good descriptors strongly distinguish between samples containing at least one polymer (classified as a positive class) and samples that do not contain any polymer (classified as a negative class).
[0045] Thus, step 101 advantageously comprises a step 102 of determining the period of the signal emitted by the polymer, this step being denoted DET-Ti. The period determined by step DET-Ti is called the period of interest Ti in the signal S of the polymer.
[0046] According to a first variant, the step DET-Ti is based on the LSP method (LSP stands for Lomb-Scargle Periodogram), such methods known in the art allowing the detection and characterization of periodic signals in irregularly sampled data.
[0047] To this end, a periodogram P of the spectral density of the signal S as a function of the period T is generated in a step PER. The period of interest Ti corresponds to the maximum of the periodogram, called Max, as Figure 4 shown.
[0048] In step PER, the so-called SciPyLSP method is applied to a regular frequency grid, preferably between 40 Da and 400 Da.
[0049] The method 100 further comprises a step 103 (SCORE) of defining a criterion for the score (Sc(LSP) of the period of interest Ti). This score makes it possible to quantify the degree to which the spectral density of the period of interest Ti differs from the spectral density calculated for other periods.
[0050] Preferably, the score Sc(LSP) is calculated using the following equation: Sc(LSP)=Max / Avg, where Avg is the average value of the spectral density over all cycles, such as Figure 5 shown.
[0051] It should be noted that this first variant can also be optimized by taking into account the fact that the data are not strictly sinusoidal, especially due to the following:
[0052] The signal is not perfectly periodic because of the presence of peaks that do not correspond to polymers: the method may advantageously include a step of removing peaks with intensities above a threshold to prevent them from adding noise to the underlying periodic signal. It is important to note that this noise is actually the signal of interest (bacterial protein), and the periodic signal of interest is that of the polymer. Here, the background is looking for the polymer in the signal, so the roles are reversed;
[0053] The intensity of the polymer mode follows a Gaussian distribution, so the signal is only periodic within a correction factor: the method may advantageously comprise a step of normalizing each peak intensity.
[0054] According to a second variant, called HD variant, the step DET-Ti is based on the use of peak-to-peak differences.
[0055] To do this, the interval E in m / z between a peak and all subsequent peaks is calculated. Then, as Figure 6 As shown, a histogram of the differences is plotted. Next, the period Ti of interest is selected based on the number of members in the class.
[0056] For example, the period Ti of interest is the class of the histogram with the highest number of members.
[0057] Preferably, the number of members in each category is replaced by a linear combination of the number of members in nearby categories. The closer the categories are, the higher the correlation coefficient.
[0058] In other words, we can write: Let En be the number of members associated with category n, a1, a2, a3∈R and satisfy 0 <a1<a2<a3<1。
[0059] To take into account the number of members in nearby categories, En is re-evaluated as follows: En = a1En-2+a2En-1+a3En+a2En+1+a1En+2.
[0060] These steps make it possible to determine even when the period Ti of interest lies between two categories.
[0061] Peak detection then makes it possible to isolate the histogram peak with the greatest prominence, where the prominence of a peak is defined as the distance between the peak top and its surrounding minima, as Figure 7 shown.
[0062] Alternatively, absolute peak size can be used, but prominence makes it possible not to take into account the fact that longer peak-to-peak differences are less common.
[0063] The period of interest (denoted as Ti(HD)) is calculated as Ti(HD) = (Cinf + Csup) / 2, where Cinf and Csup are the lower and upper limits, respectively, of the signal class with the highest number of members.
[0064] According to this variant, the score Sc(HD) is written as Sc(HD) = prominence(C1). Preferably, this is a regularized prominence, as will be explained in detail below.
[0065] Advantageously, the method according to the invention also comprises determining a descriptor allowing the periods of interest Ti(LSP) and Ti(HD) to be compared.
[0066] This descriptor is the ratio Ra = Ti(HD) / Ti(LSP), with a score of 1 associated with it if the difference Ti(LSP) - Ti(HD) > a (where a > 0 and is set empirically). A score of 1 is the worst-case scenario assumed. The threshold value, labeled 'a', eliminates cases where the result of one of the variants is physically incorrect, which is more likely to occur when the spectrum does not contain any polymer. More generally, a test is performed to determine which natural number the descriptor Ra is close to, using the fact that Ti(HD) is a multiple of the true period, while Ti(LSP) is a submultiple of the true period.
[0067] It should be noted that descriptors allowing comparison of periods are not the only possible descriptors according to the invention. In particular, the maximum intensity of the periodogram obtained with the LSP method or the prominence obtained with the HD method are other possible descriptors.
[0068] It is also important to note that the concepts of descriptors and scores can be the same or different, with descriptors corresponding to raw values and scores corresponding to normalized values.
[0069] Advantageously, the method comprises a preliminary step 104 for centering the analyzed spectrum, called centering step and denoted CENT. In other words, it is a question of correctly selecting the m / z interval of the spectrum for analysis, since the signal corresponding to the polymer is not present on the entire m / z axis.
[0070] According to a first variant, a given m / z interval is selected at the beginning of the signal, for example the interval between 2000 Da and 7000 Da.
[0071] According to another variant, one limit of the interval is optimized with respect to the fraction Sc(LSP) of the period of interest found by the LSP variant. By varying this limit, the period and its fraction are recalculated to obtain the highest possible fraction, preferably via the golden section search method.
[0072] This method is a variant of the binary search method which selects the midpoint to conform to the golden ratio instead of choosing the one placed in the middle. To use it, three points x1, x2, x3 are required such that x1 < x2 < x3 and f(x1) > x2 < f(x3) (if looking for a minimum).
[0073] A bracketing method based on the golden search is employed which shrinks the interval until three points satisfying the above properties are obtained. Once the upper limit is found, a second optimization problem is solved for the lower limit while ensuring a minimum separation is maintained between the lower and upper limits to guarantee that the m / z interval is large enough to be exploitable.
[0074] Figure 8 The centered spectrum is illustrated. Thus, the selected m / z interval is between 2000 and 12000 Da. In addition, the signal corresponding to the polymer is indeed within the selected interval.
[0075] The centering step CENT allows the region of interest to be effectively centered on the polymer peak and significantly improves the results and the calculation time. In particular, this step makes it possible to calculate only for a part of the peaks instead of all.
[0076] The method includes step 105 (denoted as CLASS) which classifies the spectra into at least one class or group of positive spectra (i.e., spectra containing peaks generated by the polymer) and one class or group of negative spectra (i.e., spectra considered to be free of peaks generated by the polymer).
[0077] According to the first variant, an unsupervised method is used, preferably K-means clustering.
[0078] The principle of K-means clustering is to optimize the groups such that each member in a group is as close as possible to all other members in the descriptor space.
[0079] One group thus obtained contains most of the positive spectra, while the negative spectra are distributed among at least one other group. Classifying the negative spectra into more than two groups makes it possible to reflect the diversity of the spectra in that group.
[0080] This first unsupervised classification has a non-negligible advantage in that no prior information about the data is required: it blindly generates groups without knowing the various spectral classes.
[0081] According to the second variant, a supervised method is used, advantageously the SVM method (SVM stands for support vector machine).
[0082] Preferably a linear SVM method is used. This method consists in making a decision based on a linear combination of descriptors.
[0083] exist Figure 9 (which shows the data in the descriptor space), a hyperplane (solid line in the figure) maximizes the margin between the classes with positive spectrum (dashed line) and the classes with negative spectrum (dash-dot line), that is, the separation between the solid line and the dashed and dash-dot lines.
[0084] It is important to note that for classification purposes, the classified data can be viewed as a spectrum summarized by the descriptors and their scores.
[0085] Preferably, before the classification step, each descriptor is regularized according to the following equation: For a given descriptor X: XRegularized = [X-min(Xtrain)] / [max(Xtrain)-min(Xtrain)].
[0086] By subtracting the minimum value taken by this descriptor in the training dataset and dividing by the difference between this minimum and its maximum, each descriptor is made to take comparable values. This makes it possible to classify data without favoring one descriptor solely based on its magnitude range.
[0087] As previously mentioned, if at the end of the classification step the spectrum is deemed to be free of polymers, the method 100 then comprises a step of identifying the one or more bacteria (if any) present in the sample. This step comprises, in particular, comparing the spectrum with spectra belonging to a library of reference spectra, each reference spectrum being a spectrum of a given bacterium.
[0088] If at the end of the classification step the spectrum is deemed to contain one or more polymers, method 100 optionally includes a step of indicating that the signal contains at least one polymer, for example via a visual and / or audible alarm. This step of indicating contamination may cause method 100 to stop without performing bacterial identification.
[0089] According to another variant, the method 100 according to the invention optionally comprises a step 106 of removing the signal emitted by the polymer (denoted RMVL).
[0090] Thus, for a given spectrum, when the classification step deems a polymer to be present, the method 100 allows the removal of the polymer's peak in order to isolate the signal emitted by the bacterial protein.
[0091] The removal step is preceded by a step 107 (MODEL) of modeling the signal emitted by the polymer.
[0092] In this step, the pattern of polymer peaks is separated.
[0093] First, the signal is divided into successive windows of the size of one period in order to isolate one polymer mode (possibly with an offset) in each window. It is preferred to use the period Ti(HD) because, on the one hand, it is generally more accurate and, on the other hand, it also makes it possible to obtain multiples of the period and thus isolate one or more polymer modes.
[0094] In contrast, the period Ti(LSP) is a submultiple of the period, and its use may result in each window isolating only a portion of the pattern.
[0095] It is important to note that the accuracy of the period is important to avoid introducing an offset along the m / z axis; the window would then no longer be centered on the polymer pattern.
[0096] Once the optimal window ( Figure 10 ), it is copied to the entire spectrum ( Figure 11 ), and applying a Gaussian function to the copied spectrum ( Figure 12 ).
[0097] This yields a theoretical polymer signal which must then be subtracted from the overall signal to remove it.
[0098] The method 100 according to the invention thus ensures a reliable and rapid classification of spectra, making it possible to distinguish between spectra contaminated with polymer and spectra free of polymer.
[0099] By means of the method 100 it is also possible to isolate the signal emitted by the bacteria by removing the signal caused by the polymer.
[0100] As should be clear from the above description, the present invention has many advantages, including: a / the method is purely numerical and therefore very cheap in terms of time and money compared to biological or chemical methods; b / depending on the use case: the detection method can be used for quality control or as a preliminary step before a step to remove polymer peaks; c / the present invention can be used directly to determine which polymers are present in a mass spectrometry sample, so it is applicable in other contexts besides the identification of bacteria.
Claims
1. A method for detecting the presence of a signal generated by at least one polymer in at least one spectrum obtained by mass spectrometry of a biological sample, comprising: a step (101) of determining a set of at least one parameter of the presence of a polymer, said parameters being called descriptors; a step (103) of assigning a score to each value of each descriptor; and a step (105) of classifying the spectrum based on each descriptor and each score in order to detect the presence or absence of said at least one polymer.
2. The method of claim 1, wherein the step (101) of determining at least one set of descriptors comprises the steps of calculating: the spectral period obtained by Lomb-Scargle calculation, said period being called the Lomb-Scargle period (Ti(LSP)); and / or the spectral period obtained by calculating, for at least some of the peaks, the interval between a peak and all other peaks on the m / z scale, said period being called the differential period (Ti(HD)); and / or a parameter allowing comparison of said Lomb-Scargle period (Ti(LSP)) and the differential period (Ti(HD)).
3. The method of claim 2, wherein the score associated with the Lomb-Scargle period (Ti(LSP)) depends on the ratio of the maximum value of the spectral density to the average value of the spectral density in a given m / z bin.
4. The method of any one of claims 2 or 3, wherein the score associated with the differential period depends on the prominence of the peak height differences.
5. The method of any of the preceding claims, comprising the step of refining the analysis window on the m / z scale (104).
6. The method of the preceding claim, wherein in the refinement step, a limit of the window is evaluated using a period obtained by a Lomb-Scargle calculation.
7. The method according to any of the preceding claims, comprising the step of identifying at least one bacterium in said at least one spectrum obtained by mass spectrometry, if no polymer is detected at the end of the classification step.
8. The method according to any of the preceding claims, comprising the step of sounding an alarm that said at least one polymer has been detected, if said at least one polymer has been detected at the end of the classification step.
9. The method of any of the preceding claims, comprising the step of modeling (107) the signal generated by the polymer.
10. The method of the preceding claim, comprising a step (108) of removing the signal obtained by said modeling step.
11. A computer program comprising instructions for carrying out the method of any preceding claim when said program is executed by a processor. 12 . A non-transitory computer-readable recording medium having a program recorded thereon for implementing the method of claim 1 when the program is executed by a processor.