Method for peak picking in lc / ms
Patent Information
- Authority / Receiving Office
- CA · CA
- Patent Type
- Applications
- Current Assignee / Owner
- BIOCRATES LIFE SCIENCES AG
- Filing Date
- 2025-01-30
- Publication Date
- 2025-08-07
AI Technical Summary
Existing AI algorithms struggle to accurately identify and quantify compounds in complex biological samples due to contaminations, impurities, and matrix effects in LC/MS, leading to reduced performance in peak picking.
A method involving pre-processing steps to generate an artificial chromatogram by subtracting median values from zero samples, selecting retention time windows, applying scaling and smoothing algorithms, and using AI algorithms like U-net to enhance peak detection and quantification accuracy.
The method significantly improves data analysis throughput and accuracy in peak picking and compound quantification, ensuring a safe and automated workflow without increasing computational time.
Abstract
Description
[0001] METHOD FOR PEAK PICKING IN LC / MS
[0002] The present invention relates to a new method for peak picking by using at least one arti ficial intelligence (Al ) algorithm, wherein the presence of at least one compound corresponds to the regions in a Liquid Chromatography / Mass Spectrometry ( LC / MS ) chromatogram by detecting and identi fying peaks with pre-processing steps for cleaning up the chromatogram in order to improve the results of the Al algorithm on complex samples .
[0003] Metabolomics is a comprehensive quantitative measurement of low molecular weight compounds covering systematically the key metabolites , which represent the whole range of pathways of intermediary metabolism . In a systems biology approach, it provides a functional readout of changes determined by genetic blueprint , regulation, protein abundance and modi fication, and environmental influence . The capability to analyse large arrays of metabolites extracts biochemical information reflecting true functional end-points of overt biological events while other functional omics technologies such as transcriptomics and proteomics , though highly valuable , merely indicate the potential cause for phenotypic response . Therefore , they cannot necessarily predict drug ef fects , toxicological response or disease states at the phenotype level unless functional validation is added .
[0004] In general , phenotype is not necessarily predicted by genotype . The gap between genotype and phenotype is spanned by many biochemical reactions , each with individual dependencies to various influences , including drugs , nutrition and environmental factors . In this chain of biomolecules from the genes to phenotype , metabolites are the quanti fiable molecules with the closest link to phenotype . Many phenotypic and genotypic states , such as a toxic response to a drug or disease prevalence are predicted by di f ferences in the concentrations of functionally relevant metabolites within biological fluids and tissues .
[0005] Hereto , metabolomics refers to a method and an apparatus capable of identi fying relevant information in biological samples providing a metabolite profile of a biological sample in a reliable way within a manageable time .
[0006] Metabolomics deals with an apparatus for analyzing a metabolite profile in a biological sample containing at least one metabolite . The apparatus or device includes ( see Figure 9 ) : an Input uni t for inputting the kind of metabolites to be screened, a controlling uni t for determining a parameter set for sample preparation and for mass spectrometry analysis depending on the Input of the kind of metabolites to be screened, a trea tment uni t for preparing the metabolites to be screened depending on the determined parameter set , a mass spectrometer for performing mass spectrometry analysis on prepared metabolites depending on the parameter set , a da tabase for storing results of mass spectrometry analysis and / or parameter sets used for metabolite preparation and for mass spectrometry analysis , an eval ua ti on uni t for evaluating and comparison of the results derived by mass spectrometry using reference results stored in the database ( e . g . reference chromatogram, QC references , etc . ) , and for output of the analysis of the metabolite profile included in the biological sample under consideration .
[0007] This is based on the idea that by providing an apparatus for preparing a biological sample and for analysing metabolites in a biological sample , reproducible conditions for analysing are kept over a wide range of analyses . Moreover, by providing an automated process for preparation of the biological sample and for mass spectrometry analysis , the comprehensive amount of data to be processed could be handled . By adapting the evaluation of the results of mass spectrometry utilising reference results stored in the database , the analyses results derived from drug and / or metabolite profiles could be improved by use of the knowledge collected during previous analyses . Further, quanti fication by utili zing appropriate internal standards ( ISTD) is used for quanti fying the drugs and / or metabolites in a biological sample . By use of automated pre- analytical procedures , the amount of results is reduced . A plurality of standard operational procedures ( SOP ) is used for treating and handling a plurality of biological samples in a standardi zed way, depending on the metabolites to be screened .
[0008] Liquid Chromatography / Mass Spectrometry ( LC / MS ) is a widely used analytical technique in metabolomics due to its sensitivity, speci ficity, and the ability to handle a wide range of compounds . Metabolomics using LC / MS typically involves a systematic workflow . Biological samples , such as blood, urine , tissues , or cells , are first extracted to obtain the metabolites . The extracted sample is then separated using liquid chromatography, which separates the metabolites based on their chemical properties . LC is employed for separating complex mixtures of metabolites . Di f ferent LC techniques , such as reversed-phase or hydrophilic interaction chromatography, can be used to enhance the separation of metabolites based on factors like hydrophobicity, polarity, or charge . LC / MS data is processed to identi fy metabolites by comparing experimental data with databases of known mass spectra . I sotopic patterns , fragmentation patterns , and retention times are often used for metabolite identi fication . Confirmation of identities can be achieved using standards or additional techniques like tandem mass spectrometry (MS / MS ) .
[0009] The abundance of metabolites is quanti fied based on the intensity of peaks in the mass spectra or chromatograms . Data analysis involves statistical and bioinformatics methods to identi fy signi ficant changes in metabolite levels , pathways af fected, and potential biomarkers .
[0010] Peaks appear in the data ( graphs ) obtained from mass spectrometers . Their height and area indicate the amount of a compound present . Moreover, the intensity of a peak is a valuable information .
[0011] Hence , a 2D-chromatogram is represented as follows :
[0012] X-axis (Hori zontal Axis ) : Represents time or the volume of the mobile phase ( elution time ) .
[0013] Y-axis (Vertical Axis ) : Represents the signal intensity recorded by the detector, namely mass spectrometry . The intensity of the signal from the analyte is compared to the noise level in the chromatogram . A higher signal-to-noise ratio indicates a more reliable and accurate measurement . It is crucial for both qualitative and quantitative analyses .
[0014] Peaks : Each peak in the chromatogram represents a separated compound of a sample . The height of the peak corresponds to the abundance or concentration of the compound, while the width of the peak indicates the distribution of the interaction strength of the compound with the chromatographic system / column . The height of a peak in a chromatogram corresponds to the maximum intensity of the signal produced by a particular compound . Peak height is determined by concentration of a compound . Peak height is often used for semi-quantitative or qualitative assessments of the concentration of a compound . The area under a peak is a more accurate measure of the quantity of a compound than peak height alone . It takes into account both the intensity and the duration of the signal . The area is calculated by integrating the signal over the entire width of the peak . Larger peak areas generally indicate higher concentrations of the analyte . The process of measuring the area under a peak is known as integration . This involves summing up the signal intensities at each point across the entire peak . The integration process provides a quantitative measure of the amount of a particular compound in the sample .
[0015] The term "peak centre" essentially means the retention time at which the peak of a particular compound is located in the chromatogram . It represents the time at which the concentration of that compound is at its highest . Therefore , each peak centre has a minimum to the left and to the right .
[0016] Retention Time : The time it takes for a particular compound to travel from the point of inj ection to the detector . Retention time is a characteristic property used for identi fication and quanti fication of compounds .
[0017] Baseline : The baseline is the hori zontal line on the chromatogram indicating the signal level in the absence of any sample components . Peaks rise above the said baseline or so called zero line .
[0018] The term " zero sample" means a blank or a control sample that does not contain the analyte of interest . These samples are used to establish a baseline ( supra ) and determine any background signals or interferences that might be present in the chromatographic system .
[0019] In general , for the purpose of the present invention, said metabolomics analysis method comprises the generation of intensity data or peaks for the measurement of metabolites or compounds by mass spectrometry (MS ) , in particular MS- technologies selected from the group of Matrix Assisted Laser Desorption / Ionisation (MALDI ) , Electro Spray Ioni zation (ES I ) , Atmospheric Pressure Chemical Ioni zation (APCI ) , Selected Reaction Monitoring ( SRM) or Multiple Reaction Monitoring (MRM) , Triple Quadrupole LC / MS (QQQ) , Orbitrap LC / MS, coupled to methods of separation, i.e. Liquid Chromatography (LC-MS) . According to the invention HPLC (High-Performance Liquid Chromatography) is also envisaged.
[0020] Peak picking in the context of Liquid Chromatography / Mass Spectrometry (LC / MS) refers to the process of identifying and extracting relevant information from the raw data generated during an LC / MS experiment providing a sample chromatogram. Overall, peak picking is a crucial step in the data analysis workflow of LC / MS, enabling researchers to identify and quantify the components of a sample, providing valuable information for fields such as chemistry, biochemistry, pharmaceuticals, and environmental science, in particular metabolomics .
[0021] Peak picking involves several key steps:
[0022] Data Acquisition: LC / MS instruments generate raw data, consisting of mass spectra and chromatograms, as the sample is analysed. The mass spectra provide information about the mass- to-charge ratio of ions, while the chromatogram shows the intensity of these ions over time.
[0023] Signal Processing: Raw data often contains noise and other artifacts. Signal processing techniques are applied to enhance the quality of the data, including baseline correction and noise reduction.
[0024] Peak Detection: The process of identifying peaks involves recognizing the elevated regions in the chromatogram that correspond to individual compounds. This is typically carried out using computational algorithms that can distinguish true peaks from background noise.
[0025] Peak Deconvolution: In complex samples, multiple compounds may elute at the same time, resulting in overlapping peaks. Peak deconvolution is the process of separating overlapping peaks to identify individual compounds more accurately.
[0026] Compound Identification: Once peaks are detected, the next step is to identify the compounds responsible for those peaks. This is often achieved by comparing the mass spectra and retention times with reference databases or using additional techniques like tandem mass spectrometry (MS / MS) .
[0027] Quantification: After identification, the abundance or concentration of each compound is determined. This is typically carried out by measuring the area under the peak in the chromatogram, a process known as peak integration.
[0028] However, it is well known that samples, particularly biological or complex samples, i.e. with a multitude of compounds and contaminations or impurities, affect the performance of LC / MS and the related sample chromatogram (hereinafter: "sample chromatogram") .
[0029] It is also known that Al algorithms are used to identify compounds, but the performance is lower, particularly for biological samples, complex samples where contaminations or impurities may occur, low concentrations and matrix effects that make it difficult for Al algorithms or human experts to identify the correct peak.
[0030] It is therefore an object of the present invention to provide a new method for peak picking that allows a safe automated workflow with an ideal application of Al algorithms.
[0031] The problem is solved by means of at least one of the technical teachings as set out in the claims.
[0032] In LC / MS, the use of quality control (QC) samples is essential to ensure that the chromatography has been performed correctly. QC samples have a known chemical composition of known compounds and provide high quality signals that can be used for normalisation across multiple measurements and for identifying the correct peaks in 'unknown' samples, i.e. samples of unknown chemical composition and compounds.
[0033] The inventors have now found that an improvement can be achieved when inventive pre-processing steps lead to better results through the generation of a so-called "artificial chromatogram", enabling a safe automated process with a consistent workflow.
[0034] Hence, the present invention refers to a method for peak picking by using at least one Al algorithm wherein the presence of at least one compound corresponds to the regions in a Liquid Chromatography / Mass Spectrometry (LC / MS) chromatogram by detecting and identifying peaks, characterized in that a first original sample chromatogram is provided and the following steps are carried out: a.) generating a second artificial chromatogram having at least one peak by subtracting the median of the zero samples from the first original chromatogram, bl.) selecting at least one peak from a.) which correlates with the retention time of a known compound and b2. ) selecting a rectangular window (region of interest) to the left and right of the highest maximum of the peak, the rectangular window being horizontally to the left and right independently of each other a maximum of less or more 30 seconds of the retention time of bl.) , the window running perpendicular to the baseline, c.) using a scaling algorithm based on a comparison between the highest intensities of the selected peaks from bl.) and a reference chromatogram to modify the window according to b2. ) and setting all intensities outside the modified window to zero, d.) using a smoothing algorithm and removing small noises still present in the peaks from c.) e.) wherein the peak from d.) is picked up by at least one Al algorithm, the corresponding retention time and the first minima left and right of the peak centre are defining a modified window from c.) , which is transferred to the first sample chromatogram and replaces the correlating peak, f.) wherein the minima left and right of the peak centre of the peak e.) are determined by a minimum finder algorithm defining a peak area, and g.) wherein the integration of the peak area according to f.) is performed.
[0035] The inventive approach is based on the novel aspect of combining the use of QCs, which provide reference values for the retention time of the compounds of interest, followed by a complex data cleaning process. Existing prior art approaches analyse all samples (QC samples and unknown samples) independently, whereas the inventors create a refined artificial chromatogram that is used to detect compounds in the unknown samples. This has the advantage of significantly improving the accuracy of peak detection without drastically increasing the computational time required for the AT algorithm.
[0036] The result of these pre-processing steps is a significant increase in data analysis throughput while maintaining accuracy, making peak picking and compound quanti fication safer .
[0037] In a further preferred embodiment of the invention the retention time is normali zed to compensate fluctuations in the settings (pressure , gradient , mobile phase , etc . ) .
[0038] In a preferred embodiment of the invention, the known compounds of step bl . ) refer to QCs which serve as a reference for the retention time at which a compound is expected to elute . The arti ficial chromatogram is cleaned by using the retention time (window) of the QC samples . All values determined outside the selected retention time window are set to values near zero or zero and thus hidden from the peak picking process . Hereto , " zero" means that the value is set to zero or near to zero .
[0039] This ensures that the selected peak is within the retention time range in which it is expected to elute . A further preference is to select the retention time range (window) in any range to the left and right of the desired retention time .
[0040] Step b2 . ) : I f the sample concentration is in the same range as the QC ( known) samples , the peaks obtained should be similar in height , width, and shape . However, when the concentration is low, the peaks tend to become narrower, lose their Gaussian shape and are surrounded by more and more noise . This can reduce the accuracy of an Al peak picking algorithm, as it becomes more di f ficult to find the correct peak .
[0041] Step c . ) : By comparing the height of the highest point of the peak with a reference chromatogram of a QC sample or known compound, a scaling algorithm is used to reduce the retention time (window) region in which the peak is expected to be located . This further removes noise from the arti ficial chromatogram and increases the chances of selecting the correct peak.
[0042] Step d.) : The smoothing algorithm is based on a moving average approach. It is a standard algorithm used for smoothing of chromatograms and is likely present in software used for data processing of chromatograms (e. g. Analyst)
[0043] In preferred embodiment the Al algorithm according to step e.) refers to a U-net based algorithm, trained on LC / MS data of known samples. Similar Al algorithms are present in literature: PeakOnly (https : / / pubs .acs.org / do i / 10.1021 / acs. analchem.9b04811 ) - CNN (U-net based) , PeakBot
[0044] (https : / / academic . oup . com / bio informatics / article / 38 / 13 / 3422 / 65 90644) - CNN network, (https: / / pubs.acs.org / doi / 10.1021 / acs.analchem.9b02983) - CNN.
[0045] Moreover, Machine Learning Algorithms are useful as Al algorithm, where various machine learning algorithms can be trained to recognize patterns associated with peaks in LC / MS data. Some commonly used machine learning algorithms include:
[0046] Support Vector Machines (SVM) : Effective in binary classification tasks and can be trained to distinguish between peak and non-peak regions .
[0047] Random Forests: Ensemble learning method that can handle complex data sets and is capable of peak detection with high accuracy .
[0048] Neural Networks: Deep learning models, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs) , can learn intricate patterns in LC / MS data for peak identification . Step f . ) : The minimum finder algorithm is known to the skilled person and such approaches are commonly used in peak identi fication approaches provided in software used for data processing of chromatograms .
[0049] Examples and Figures
[0050] The invention is further illustrated by the following Figures 1 to 8 .
[0051] In order to identi fy the correct peak by peak picking, a multi-step protocol is implemented according to the invention . First , an arti ficial chromatogram is generated by subtracting the median of the zero samples from the original chromatogram, as shown in Figure 1 ( step a . ) ) .
[0052] This arti ficial chromatogram is subj ected to several cleaning steps in order to increase the detection accuracy for the AI- based detection .
[0053] First , the region of interest is identi fied ( see Figure 2 , step b . ) . An estimate of the region of interest can be obtained by determining the retention time of known ( reference ) compounds . Once the initial region of interest has been defined, it can be refined by scaling it according to the maximum intensity detected within the region of interest window ( step c . ) . These values can be compared to a reference and scaled appropriately, using a non-linear scaling formula, such as :
[0054] / Detected intensity \ SF = A * log, I — - - I
[0055] \Reference intensity / wherein SF belongs to the scaling factor, A is an empirically determined correction factor, the detected intensity represents the maximum peak height within the region of interest and the reference intensity represents the maximum intensity of the reference chromatographic peak . Using this calculated scaling value, the region of interest is either reduced or extended in size, as shown in Figure 3, remaining centred on the retention time position of the known (reference) compound.
[0056] Once the final region of interest was assigned, the artificial chromatogram undergoes two final cleaning steps. The intensity values outside of the region of interest are set to 0 (Figure 4, step c.) ) , eliminating interfering peaks outside the region of interest. The resulting artificial chromatogram is then smoothed (Figure 5, step d.) ) and prepared for interpretation by the AT algorithm.
[0057] The AT algorithm detects the borders of the peak in the artificial chromatogram (Figure 6, step e.) ) , which are then transferred back to the original chromatogram (Figure 7) . Finally, a minimum finder algorithm can be used to locate the minima closest to the peak (Figure 8, step f.) ) .
[0058] This ensures that the peak borders are correctly adapted to the original chromatogram, as the lowest points next to the peak in the original chromatogram are not necessarily the lowest points in the artificial chromatogram, due to the zerosample correction that took place in the first processing step .
[0059] The steps a.) to f.) may be part of an automated workflow. The workflow can be implemented in an apparatus or device as depicted in Figure 9 and described above.
[0060] Hence, the method according to the invention allows the identification, semi-quantification, or quantification of one or more compounds .
Claims
Claims1. A method for peak picking by using at least one Al algorithm wherein the presence of at least one compound corresponds to the regions in a Liquid Chromatography / Mass Spectrometry (LC / MS) chromatogram by detecting and identifying peaks, characterized in that a first original sample chromatogram is provided and the following steps are carried out: a.) generating a second artificial chromatogram having at least one peak by subtracting the median of the zero samples from the first original chromatogram, bl.) selecting at least one peak from a.) which correlates with the retention time of a known compound and b2. ) selecting a rectangular window (region of interest) to the left and right of the highest maximum of the peak, the rectangular window being horizontally to the left and right independently of each other a maximum of less or more 30 seconds of the retention time of bl.) , the window running perpendicular to the baseline, c.) using a scaling algorithm based on a comparison between the highest intensities of the selected peaks from bl.) and a reference chromatogram to modify the window according to b2. ) and setting all intensities outside the modified window to zero, d.) using a smoothing algorithm and removing small noises still present in the peaks from c.) e.) wherein the peak from d.) is picked up by the at least one Al algorithm, the corresponding retention time and thefirst minima left and right of the peak centre are defining a modified window from c.) , which is transferred to the first sample chromatogram and replaces the correlating peak, f.) wherein the minima left and right of the peak centre of the peak e.) are determined by a minimum finder algorithm defining a peak area, and g.) wherein the integration of the peak area according to f.) is performed.
2. A method for peak picking according to claim 1 wherein the Liquid Chromatography / Mass Spectrometry (LC / MS) contains mass spectrometry selected from the group of Matrix Assisted Laser Desorption / Ionisation (MALDI) , Electro Spray Ionization (ESI) , Atmospheric Pressure Chemical Ionization (APCI) , Selected Reaction Monitoring (SRM) or Multiple Reaction Monitoring (MRM) , Triple Quadrupole LC / MS (QQQ) or Orbitrap LC / MS.
3. A method for peak picking according to claim 1 or claim 2 wherein the Liquid Chromatography / Mass Spectrometry (LC / MS) contains High-Performance Liquid Chromatography.
4. A method for peak picking according to any one of claims 1 to 3, wherein at least one compound corresponds to the regions in the chromatogram of a Liquid Chromatography / Mass Spectrometry (LC / MS) is identified by an Al algorithm.
5. A method for peak picking according to any one of claims 1 to 4, wherein the method presents an automated workflow.
6. A method for peak picking according to any one of claims 1 to 5 for the identification, semi-quantification, or quantification of one or more compounds.