Computer-implemented method for identifying at least one peak in a mass spectrometric response curve

By training a model using a deep learning regression architecture and using a convolutional neural network to identify peaks in the mass spectrometry response curve, the problem of automatic peak area identification and determination in LC-MS data processing is solved, improving accuracy and efficiency.

CN115280143BActive Publication Date: 2025-11-25VENTANA MEDICAL SYSTEMS INC +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202180024754.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-27
Filing Date
2021-03-26
Publication Date
2025-11-25
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing liquid chromatography-mass spectrometry (LC-MS) data processing requires manual data review and has a high error rate, making it difficult to reliably and automatically identify and determine peak areas in mass spectrometry response curves.

Method used

The model is trained using a deep learning regression architecture. The peak start and end points of the mass spectrometry response curve are identified by a convolutional neural network. Combined with one-dimensional signal analysis, the peak area is automatically determined.

Benefits of technology

It achieves reliable automatic identification of peaks and accurate quantification of area in mass spectrometry response curves, reducing manual intervention and error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115280143B_ABST
    Figure CN115280143B_ABST
Patent Text Reader

Abstract

A computer-implemented method (110) for identifying at least one peak in a mass spectrometric response curve is presented. The method comprises the steps of: a) providing (112) at least one mass spectrometric response curve by using at least one mass spectrometric device (114); b) evaluating (116) the mass spectrometric response curve by using at least one trained model to identify a start point and an end point of at least one peak of the mass spectrometric response curve, wherein the model is trained using a deep learning regression architecture (118).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to a computer-implemented method for identifying at least one peak in a mass spectral response curve and to a device for monitoring at least one analyte in a sample. The method and device proposed by the present invention can be used in the field of mass spectrometry, in particular for liquid chromatography-mass spectrometry. BACKGROUND

[0002] Current liquid chromatography-mass spectrometry (LC-MS) data processing usually requires a manual data review of all acquired data and due to a high error rate subsequent manual correction of about 5-20% of the results. The above operations are performed by trained LC-MS operators by tedious visual analysis of hundreds of chromatograms.

[0003] Furthermore, automated methods for identifying and characterizing spectral peaks of a spectrum are known.

[0004] WO 2012 / 047417 A1 and US 8,428,889 B2 describe a method for automatically identifying and characterizing spectral peaks of a spectrum produced by an analytical device, the method comprising the steps of receiving a spectrum produced by an analytical device, automatically subtracting a baseline from the spectrum to produce a baseline-corrected spectrum, automatically detecting and characterizing spectral peaks in the baseline-corrected spectrum, and reporting the detected and characterized spectral peaks to a user. A list of adjustments to be made to the detecting and characterizing steps is received from the user. Based on the list of adjustments, exit values used in the detecting and characterizing steps are adjusted. The automatic detecting and characterizing of spectral peaks is repeated in the same spectrum or in a different spectrum.

[0005] US 7,219,038 B2 describes a protocol for analysis and peak identification in spectral data. A Bayesian approach is used to automatically identify peaks in a data set. After identifying a peak shape, the method tests the hypothesis that a given number of peaks can be found in any given data window. If peaks are identified within a given window, a likelihood function is maximized to estimate the position and amplitude of the peaks.

[0006] US 7,720,612 B2 describes a method for resolving convoluted peaks in a chromatogram into one or more component peaks using a peak resolution value. The peak method of the invention determines empirical-based peak resolution values for "well-defined" or "isolated" peaks in the data, and then extrapolates these empirical-based resolution values to peaks in adjacent regions to predict the number of component peaks at a given peak position. The predicted peak resolution values are compared to the observed peak resolution values of low-resolution or convoluted peaks to determine the number of component peaks in the convoluted peaks.

[0007] WO 2019 / 092836 Al describes that in the learned model storage unit, a model constructed by performing deep learning using accurate peak information and an image obtained by imaging a large number of chromatograms are stored in advance as learning data. When inputting chromatogram data of a target sample obtained by an LC measurement unit, an image generation unit images the chromatogram and generates an input image in which one of two regions on either side of the chromatogram curve in the resulting image is filled, a peak position estimation unit inputs pixel values of the input image to a neural network using the learned model and obtains position information on the start and end points of the peak and a peak detection precision as an output. A peak determination unit determines the start and end points of the peak based on the peak detection precision.

[0008] EP 3467493 Al describes that before measurement, as a measurement end condition that is satisfied throughout the input unit, an analyst selects a chromatogram of a peak detection target (a chromatogram at a specific wavelength or across the entire wavelength) and specifies a determination value for the number of peaks. During measurement by the measurement unit, a chromatogram generation unit generates a chromatogram substantially in real time based on collected data, and a peak detection unit detects peaks on the chromatogram. A measurement end condition determination unit counts the number of detected peaks, and determines that the measurement end condition is satisfied when the counted number reaches the determination value of the number of peaks, and a measurement end timing determination unit instructs an analysis control unit to end the measurement when a predetermined time elapses from the determination.

[0009] Risum Anne Bech et al.: “Using deep learning to evaluate peaks in chromatographic data”, TALANTA, ELSEVIER, Amsterdam, NL, vol. 204, 22 May 2019, pages 255-260, XP085747637, ISSN: 0039-9140, DOI: 10.1016 / J.TALANTA.2019.05.053 describes that the analysis of non-target gas chromatography data is very time consuming and there are still many manual steps in the analysis that require expertise in data analysis. One of them is the need to define whether each resolved component represents a peak suitable for integration. Since both the shape and position of the peaks can vary on the elution time axis, this poses a problem that cannot be easily solved by applying a linear classifier such as PLS-DA (partial least squares regression for discriminant analysis). A convolutional neural network classifier is described for handling these shifts and shape variations.

[0010] Although the known methods and apparatuses have advantages, these known methods and apparatuses only give information on whether a peak is present or not. However, it cannot be guaranteed that the area under the peak is reliably determined.

[0011] Problem to be solved

[0012] Therefore, it is desirable to provide a computer-implemented method for identifying at least one peak in a mass spectrometric response curve and an apparatus for monitoring at least one analyte in a sample to address the above-mentioned technical challenges. In particular, a method for identifying at least one peak in a mass spectrometric response curve and an apparatus for monitoring at least one analyte in a sample shall be provided which allow for a reliable and automated determination of the peak area of an analyte peak in a chromatogram. SUMMARY

[0013] The problem is solved by a computer-implemented method for identifying at least one peak in a mass spectrometric response curve and an apparatus for monitoring at least one analyte in a sample.

[0014] As used in the following, the terms "have", "comprise", "include" or "contain" or any arbitrary grammatical variations thereof are used in a non-exclusive way. Thus, these terms can both refer to a situation in which, besides the feature introduced by these terms, no further features are present in the entity described in this context and to a situation in which one or more further features are present. As an example, the expressions "A has B", "A comprises B" and "A includes B" can both refer to a situation in which, besides B, no further element is present in A (i.e. a situation in which A solely and exclusively consists of B) and to a situation in which, besides B, one or more further elements are present in entity A, such as element C, elements C and D or even further elements.

[0015] Further, it shall be noted that the terms "at least one", "one or more" or similar expressions indicating that a feature or element can be present once or more than once usually are used only once in the following claims. In the following, in most cases, when referring to a respective feature or element, the expression "at least one" or "one or more" will not be repeated although the respective feature or element can be present once or more than once.

[0016] Further, as used in the following, the terminology“preferably”,“more preferred”,“particularly”,“more particularly”,“specifically”,“more specifically”, or similar terms are used in conjunction with optional features, and do not limit the replacement possibilities in any way. Thus, features introduced by these terms are optional features and do not mean that the scope of the claims in any way is limited to the preferred embodiment. The application can be performed by using alternative features. Similarly, features introduced by“in an embodiment of the application” or similar expressions are intended to be optional features and do not in any way limit alternative embodiments of the application, do not limit the scope of the application in any way and do not limit the possibility that features introduced in this way are combined with other optional or non-optional features of the application.

[0017] In a first aspect of the application, a computer-implemented method for identifying at least one peak in a mass spectrometric response curve is disclosed.

[0018] The term“computer-implemented method” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to a method involving at least one computer and / or at least one computer network. The computer and / or computer network can comprise at least one processor configured for performing at least one of the method steps of the method according to the present application. Preferably, each method step is performed by the computer and / or computer network. The method can be performed completely automatically, specifically without user interaction. The term“automatically” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to a process performed completely by means of at least one computer and / or at least one computer network and / or at least one machine, in particular, without the need for manual operations and / or user interaction.

[0019] As used herein, the term "mass spectrometry" is a broad term and will be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to an analytical technique for determining the mass-to-charge ratio of ions. Mass spectrometry can be performed using at least one mass spectrometry device. As used herein, the term "mass spectrometry device", also referred to as "mass analyzer", is a broad term and will be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to an analyzer configured for detecting at least one analyte based on mass-to-charge ratio. The mass analyzer can be or can comprise at least one quadrupole analyzer. As used herein, the term "quadrupole mass analyzer" is a broad term and will be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to a mass analyzer comprising at least one quadrupole as mass filter. The quadrupole mass analyzer can comprise a plurality of quadrupoles. For example, the quadrupole mass analyzer can be a triple quadrupole mass spectrometer. As used herein, the term "mass filter" is a broad term and will be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to a device configured for selecting ions injected into the mass filter according to their mass-to-charge ratio m / z. The mass filter can comprise two pairs of electrodes. The electrodes can be rod-shaped, in particular cylindrical. In an ideal case, the electrodes can be hyperbolic. The electrodes can be designed to be identical. The electrodes can be arranged to extend in parallel along a common axis, e.g. the z-axis. The quadrupole mass analyzer can comprise at least one power supply circuit configured for applying at least one direct current (DC) voltage and at least one alternating current (AC) voltage between the two pairs of electrodes of the mass filter. The power supply circuit can be configured for keeping each pair of opposing electrodes at the same potential. The power supply circuit can be configured for periodically changing the sign of the charge of the electrode pairs such that only ions in a certain mass-to-charge ratio m / z range can have stable trajectories. The trajectories of ions within the mass filter can be described by Mathieu differential equations. For measuring ions having different m / z values, the DC voltage and the AC voltage can be adjusted in time such that ions having different m / z values can be transmitted to a detector mass spectrometry device.

[0020] The mass spectrometry device can further comprise at least one ionization source. As used herein, the term “ionization source”, also referred to as “ion source”, is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a device configured for generating ions, e.g. from neutral gas molecules. The ionization source can be or can comprise at least one source selected from the group consisting of at least one gas phase ionization source, such as at least one electron impact (El) source or at least one chemical ionization (CI) source; at least one desorption ionization source, such as at least one plasma desorption (PDMS) source, at least one fast atom bombardment (FAB) source, at least one secondary ion mass spectrometry (SIMS) source, at least one laser desorption (LDMS) source, and at least one matrix assisted laser desorption (MALDI) source; at least one spray ionization source, such as at least one thermal spray (TSP) source, at least one atmospheric pressure chemical ionization (APCI) source, at least one electrospray (ESI), and at least one atmospheric pressure ionization (API) source.

[0021] The mass spectrometry device can comprise at least one detector. As used herein, the term “detector” is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a device configured for detecting input ions. The detector can be configured for detecting charged particles. The detector can be or can comprise at least one electron multiplier.

[0022] The mass spectrometry device, in particular the detector of the mass spectrometry device and / or the at least one evaluation device, can be configured to determine at least one mass spectrum of the detected ions. As used herein, the term “mass spectrum” is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a two-dimensional representation of signal intensity versus mass-to-charge ratio m / z, wherein the signal intensity corresponds to the abundance of a respective ion. The mass spectrum can be a pixelated image. For determining the resulting intensity of a mass spectrum pixel, the signal detected with the detector within a certain m / z range can be integrated. The analytes in the sample can be identified by the at least one evaluation device. Specifically, the evaluation device can be configured for correlating known masses with identified masses or by characteristic fragmentation patterns.

[0023] The mass spectrometry device can be or can include a liquid chromatography mass spectrometry device. The mass spectrometry device can be connected to and / or can include at least one liquid chromatograph. The liquid chromatograph can be used as sample preparation for the mass spectrometry device. Other embodiments of sample preparation are possible, such as at least one gas chromatograph. As used herein, the term “liquid chromatography mass spectrometry device” is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a combination of liquid chromatography and mass spectrometry. The mass spectrometry device can include at least one liquid chromatograph. The liquid chromatography mass spectrometry device can be or can include at least one high performance liquid chromatography (HPLC) device or at least one microfluidic liquid chromatography (pLC) device. The liquid chromatography mass spectrometry device can include a liquid chromatography (LC) device and a mass spectrometry (MS) device, in the present case a mass filter, wherein the LC device and the mass filter are coupled via at least one interface. The interface coupling the LC device and the MS device can include an ionization source configured for generating molecular ions and transferring the molecular ions into a gas phase. The interface can further include at least one ion mobility module arranged between the ionization source and the mass filter. For example, the ion mobility module can be a high field asymmetric waveform ion mobility spectrometry (FAIMS) module.

[0024] As used herein, the term “liquid chromatography (LC) device” is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to an analytical module configured to separate one or more target analytes of a sample from other components of the sample for detection of the one or more analytes using a mass spectrometry device. The LC device can include at least one LC column. For example, the LC device can be a single column LC device or a multi-column LC device having multiple LC columns. The LC column can have a stationary phase through which a mobile phase is pumped in order to separate and / or elute and / or transfer the target analytes. The liquid chromatography mass spectrometry device can further include a sample preparation station for automated pre-treatment and preparation of samples, each sample including at least one target analyte.

[0025] The term "sample" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to an arbitrary test sample, such as a biological sample and / or an internal standard sample. The sample can comprise one or more target analytes. For example, the test sample can be selected from the group consisting of a physiological fluid, including blood, serum, plasma, saliva, ocular lens fluid, cerebrospinal fluid, sweat, urine, milk, ascites fluid, mucus, synovial fluid, peritoneal fluid, amniotic fluid, tissue, cells, etc. The sample can be used directly as obtained from the respective source or can be subjected to a pre-treatment and / or sample preparation workflow. The sample can be pre-treated by adding an internal standard and / or by dilution with another solution and / or by mixing with reagents, etc. For example, generally, the target analytes can be vitamins D, drugs of abuse, therapeutic drugs, hormones, and metabolites. The internal standard sample can be a sample comprising at least one internal standard substance having a known concentration. For respective further details regarding the sample, reference is made to, for example, EP 3425369 A1, the entire disclosure of which is included herein by reference. Other target analytes are possible as well.

[0026] The method comprises the following steps, which can be performed in the given order as an example. However, it should be noted that different orders are possible as well. Further, one or more method steps can be performed once or repeatedly. Further, two or more method steps can be performed simultaneously or in a timely coinciding manner. The method can comprise further method steps which are not listed.

[0027] The method comprises the following steps:

[0028] a) providing at least one mass spectral response curve by using at least one mass spectrometry device;

[0029] b) evaluating the mass spectral response curve by using at least one trained model, thereby identifying a start point and an end point of at least one peak of the mass spectral response curve, wherein the model is trained using a deep learning regression architecture.

[0030] As used herein, the term "mass spectral response curve" is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a one-dimensional representation of signal intensity. A mass spectral response curve has only one dimension. Specifically, the term "one-dimensional" can refer to a time axis and time is the only one independent variable. Thereby, as used herein, the term "one-dimensional" can refer to the fact that the only independent variable in the data is "time" and the dependent variable is "intensity". Notably, the present application can not require two independent variables (e.g. "time" and "mass-to-charge ratio") as is the case with certain mass spectral data processing techniques. Typically, convolutional neural networks are used for two-dimensional data and applications such as for image recognition, where there are two independent variables (x- and y-direction of the image) in contrast to the use of only one-dimensional representations according to the present application. As used herein, the term "providing" is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a process of determining and / or generating and / or making available a mass spectral response curve, in particular by performing at least one measurement with a mass spectrometric device. Thereby, as used herein, the term "providing at least one mass spectral response curve by using at least one mass spectrometric device" is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to retrieving data of a mass spectral response curve obtained from a mass spectrometric device at a particular reception and / or performing at least one measurement with a mass spectrometric device such that data of a mass spectral response curve is determined, which data can be used for further evaluation in step b).

[0031] As used herein, the term "evaluating a mass spectral response curve" is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to an analysis of a mass spectral response curve. The evaluation can comprise identifying at least one peak and / or determining the start and end of a peak and / or determining the peak area of a peak. The evaluation can comprise applying at least one filter and / or using a baseline subtraction technique and / or using at least one fitting routine etc. The evaluation can be performed using at least one evaluation device as will be described in more detail below.

[0032] As used herein, the term "peak" (of a mass spectral response curve) is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to at least one local maximum of a mass spectral response curve.

[0033] As used herein, the term "onset" of a peak is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to the lower peak boundary. The onset can be the point of the time axis that defines the lower peak boundary. After the onset, the mass spectrometric response curve rises to a local maximum. The onset can be the point at which the integration of the peak begins. As used herein, the term "endpoint" of a (mass spectrometric response) curve is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to the upper peak boundary. The endpoint can be the point of the time axis that defines the upper peak boundary. The mass spectrometric response curve falls before reaching the endpoint to the noise and / or baseline level. The endpoint can be the point at which the integration of the peak ends. The onset and the endpoint can be points on the time axis that are identified as peak limits. The values of the onset and the endpoint for a training data set (to be described in more detail below) can be determined by a trained user by manual evaluation. A trained model can provide the onset and the endpoint for further data. The peak area generally can be defined as the integral of the response curve between the onset and the endpoint.

[0034] As used herein, the term "peak identification" is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a qualitative determination of a peak (such as presence or absence) and / or a quantitative determination of a peak (such as determining the peak area of a peak). The determination of a peak area can include the integration of a peak, in particular by using at least one mathematical operation and / or mathematical algorithm to determine the peak area enclosed by the peak of the mass spectrometric response curve. In particular, the integration of a peak can include the identification and / or measurement of curve features of the mass spectrometric response curve. The peak identification includes determining the onset and / or the endpoint. The peak identification can further include one or more of peak detection, peak finding, peak fitting, peak evaluation, baseline determination, and baseline determination. The integration of a peak can allow determining one or more of the following: peak area, retention time, peak height, and peak width. The peak identification can be an automated peak identification, i.e. a peak identification performed by at least one computer and / or computer network and / or machine. In particular, the automated peak identification can be performed without the need for manual operations or interaction with a user.

[0035] The term "trained model" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a model for identifying peaks in a mass spectrometric response curve, which model is trained on at least one training data set, also referred to as training data. In particular, the trained model is trained on existing data that is expert a priori classified. This enables to provide an automated peak identification with enhanced reliability and less susceptibility to variations and errors. The trained model can comprise an architecture and a set of weights for various filters or nodes defined by the architecture. The architecture of the CNN can reflect a complex relationship between the shape of the response curve and the start and end position of the peaks.

[0036] The method can comprise at least one training step, wherein in the training step the trained model is trained on at least one training data set. The training step can be an offline training, while the peak identification in step b) of the proposed method can be an online peak identification. In particular, the training step can be performed prior to performing steps a) and b). The term "online" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to during a measurement process using a mass spectrometric device.

[0037] The model is trained using a deep learning regression architecture. The term "deep learning" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to at least one method related to machine learning. Deep learning can be based on at least one artificial neural network. The term "deep learning regression architecture" as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to a deep learning architecture configured for solving a regression problem.

[0038] The deep learning regression architecture can comprise a convolutional neural network. The convolutional neural network can be a multi-layer convolutional neural network. The convolutional neural network can comprise a plurality of convolutional layers. A convolutional layer is a one-dimensional layer, i.e. the convolution is applied in one dimension, the time domain. Typically, a convolutional layer is a standard building block in a convolutional neural network and is therefore well known to the person skilled in the art. In mathematics, a convolutional layer corresponds to an operation of convolving input data with a convolution kernel (see e.g. en.wikipedia.org / wiki / Convolutional_neural_network). As used herein, the term “convolutional layer” is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to an operation of convolving input data with a one-dimensional convolution kernel. The convolutional layer can be followed by a plurality of fully connected layers. The convolutional neural network can comprise a plurality of pooling layers. The structure of a convolutional neural network is generally known to the person skilled in the art, such as from en.wikipedia.org / wiki / Convolutional_neural_network#Convolutional. It is generally known to the person skilled in the art to use a convolutional neural network for object recognition in images. However, the present invention proposes a new approach of using a convolutional neural network for one-dimensional signal analysis.

[0039] A convolutional neural network (CNN) can be configured for solving a regression problem. For solving a regression problem, the convolutional neural network can comprise a regression layer as final layer, in particular in contrast to the usual classification softmax layer. The regression layer can be a fully connected layer. The regression layer can have a linear or sigmoid activation. Thus, the present invention proposes to use a one-dimensional convolutional neural network in a regression framework. Specifically, the present invention proposes to use a convolutional neural network for fitting a complex function that maps an input to a peak position. However, a convolutional neural network can not be used for classifying a one-dimensional signal.

[0040] The training step can comprise the following substeps:

[0041] i) providing at least one training data set comprising a plurality of input mass spectrometric response curves and corresponding ground truth values;

[0042] ii) determining at least one model by using a deep learning regression architecture on the training data set, wherein the determination of the model comprises determining a model architecture of the model and at least one parameter.

[0043] Step i) can comprise providing more than 100, preferably more than 1000 input mass spectral response curves. For example, 1270 vitamin D3 curves can be used to train the model. The plurality of input mass spectral response curves provided in step i) can be determined by performing multiple measurements using the mass spectrometry device. For example, the plurality of input mass spectral response curves provided in step i) can be or can comprise LC-MS data from a particular analyte, such as from vitamin D2 or from vitamin D3. The term“true label value” as used herein is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not limited to a special or customized meaning. The term specifically can refer, without limitation, to the actual or true value of the start and end point of a peak of a corresponding input mass spectral response curve. The true label value can be indicative of the position of the peak. For example, the true label value can be the position of the peak or the position of the start and end point of the peak. The true label value can be provided by a trained LC-MS operator.

[0044] Injection into the mass spectrometry device can result in four mass spectral response curves, i.e. two analyte mass spectral response curves and two internal standard mass spectral response curves. The training data set can be provided in the form of a five-channel vector comprising the aggregated time vector, the two analyte mass spectral response curves and the two internal standard mass spectral response curves. The raw time vector of the four response curves comprises time steps that at least slightly deviate from curve to curve. The term“aggregated time vector” can refer to a time vector on which all four response curves are interpolated on the same time grid. The training data set can be provided as input (also referred to as input data) to the convolutional neural network. For example, given an input mass spectral response curve of length N, the input can be a 2xN matrix, where the first row represents N intensity values, such as the y-values of a chromatogram, and the second row can represent N time values, such as the x-values of a chromatogram. Aggregation can enable the convolutional neural network to propagate information between different mass spectral response curves, such that, for example, in case a peak on a given analyte curve is particularly weak, the peak positions of the other curves can be used to inform the position of the weak curve.

[0045] The training of the model can comprise at least one normalization step. The normalization step can comprise normalizing the input data with respect to time. The normalization step can comprise shifting the time values such that the expected retention time can be at t=0. The normalization step can comprise cropping the input data to a fixed time window around the expected retention time. The normalization step can comprise normalizing the intensity values Y themselves by:

[0046]

[0047] The training of the model can comprise at least one augmentation step. To achieve generalization across analytes, the convolutional neural network can be trained using augmented data that takes into account the shift and scaling differences in the mass spectrometric response curve data. The augmentation step can avoid overtraining. The augmentation step can comprise a position augmentation and / or a scaling augmentation. The position augmentation can comprise shifting the peak position by a predetermined constant value. The position augmentation can comprise using a sliding window. This can take into account the possibility of peaks broadening and narrowing. The scaling augmentation can comprise scaling the peak value by a predetermined value, such as 1.2. For each input mass spectrometric response curve, the training data set can be supplemented by randomly generating a new data set using the position augmentation and randomly generating three new data sets using the scaling augmentation.

[0048] The method can comprise using a deep learning regression architecture on the normalized and / or augmented training data set together with the true label values.

[0049] The deep learning regression architecture can be a convolutional neural network (CNN) built in Python by the Keras library with TensorFlow as backend. For the Keras library in Python with TensorFlow, see https: / / www.tensorflow.org / , https: / / de.wikipedia.org / wiki / TensorFlow and https: / / keras.io / or https: / / de.wikipedia.org / wiki / Keras. For example, the following settings can be used: Adaptive Moment Estimation (Adam) can be used as optimizer. The loss function can be mean squared error. The number of epochs can be 500, the batch size 16 and the patience early stopping 100.

[0050] The training step can further comprise at least one testing step using at least one test data set. The testing step can comprise a validation of the trained model. As used herein, the term "test data set" is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to at least one mass spectrometric response curve and the corresponding true label value not included in the training data set. The testing step can comprise determining the onset and the end of at least one peak of the mass spectrometric response curve of the test data set using the determined model. The performance of the determined model can be determined based on the determined onset and end of the mass spectrometric response curve of the test data set and the true label value.

[0051] The method can comprise determining, in particular automatically, the peak area of the peak of the mass spectrometric response curve provided in step i) by using the identified onset and end.

[0052] In a further aspect of the application, a computer program or identifying at least one peak in a mass spectrometric response curve is disclosed, the computer program comprising instructions which, when the program is executed by a computer or computer network, cause the computer or computer network to carry out steps a) and b) of the method according to the application, fully or partially, such as according to any one of the embodiments disclosed above and / or according to any one of the embodiments disclosed in further detail below. Step a) can comprise steps which can at least partially be carried out by a user, such as sample preparation steps. However, embodiments are possible in which all steps of the method according to the application are carried out fully automatically. Thus, in particular, one, more than one or even all method steps a) and b) as indicated above can be carried out by using a computer or computer network, preferably by using a computer program. In particular, the computer program can be stored on a computer-readable data carrier and / or on a computer-readable storage medium.

[0053] Similarly, a computer-readable storage medium is disclosed, comprising instructions which, when executed, cause a computer or computer network to carry out steps a) and b) of the method according to the application, fully or partially, such as according to any one of the embodiments disclosed above and / or according to any one of the embodiments disclosed in further detail below.

[0054] As used herein, the term "computer-readable storage medium" can in particular refer to a non-transitory data storage device, such as a hardware storage medium having stored thereon computer-executable instructions. The computer-readable data carrier or storage medium can in particular be or can comprise a storage medium such as a random access memory (RAM) and / or a read-only memory (ROM).

[0055] The computer program can also be embodied as a computer program product. As used herein, a computer program product can refer to a program as a tradable product. The product can generally exist in any format, such as a paper format, or on a computer-readable data carrier and / or computer-readable storage medium. In particular, the computer program product can be distributed over a data network.

[0056] Further disclosed and proposed herein is a data carrier having a data structure stored thereon, which, after loading into a computer or computer network, such as into a working memory or main memory of the computer or computer network, can carry out the method according to one or more of the embodiments disclosed herein.

[0057] It is further disclosed and proposed herein a computer program product having stored on a machine readable carrier program code means to perform the method according to one or more embodiments disclosed herein when the program is executed on a computer or computer network. As used herein, a computer program product refers to a program as a tradable product. The product can generally be present in any format, such as a paper format, or on a computer readable data carrier and / or computer readable storage medium. In particular, the computer program product can be distributed over a data network.

[0058] It is further disclosed and proposed herein a modulated data signal comprising instructions readable by a computer system or computer network for performing the method according to one or more embodiments disclosed herein.

[0059] With reference to the computer implemented aspects of the present application, one or more of the method steps, or even all of the method steps, of the method according to one or more embodiments disclosed herein can be performed by using a computer or computer network. Thus, generally speaking, any of the method steps including providing and / or processing data can be performed by using a computer or computer network. Generally speaking, these method steps can include any of the method steps except for method steps that typically require manual operations, such as providing a sample and / or performing certain aspects of an actual measurement.

[0060] In particular, it is further disclosed herein:

[0061] - a computer or computer network comprising at least one processor, wherein the processor is adapted to perform the method according to one of the embodiments described in the present specification,

[0062] - a computer loadable data structure adapted to perform the method according to one of the embodiments described in the present specification when the data structure is executed on a computer,

[0063] - a computer program adapted to perform the method according to one of the embodiments described in the present specification when the program is executed on a computer,

[0064] - a computer program comprising program means for performing the method according to one of the embodiments described in the present specification when the computer program is executed on a computer or on a computer network,

[0065] - a computer program comprising program means according to the preceding embodiment, wherein the program means are stored on a computer readable storage medium,

[0066] - a storage medium, wherein a data structure is stored on the storage medium and wherein the data structure is adapted to perform the method according to one of the embodiments described in this description after being loaded into a main memory and / or a working memory of a computer or computer network, and

[0067] - a computer program product having program code means, wherein the program code means can be stored on or in a storage medium or on or in a storage medium for performing the method according to one of the embodiments described in this description when the program code means are executed on a computer or computer network.

[0068] In a further aspect of the application, a device for monitoring at least one analyte in a sample is disclosed. The device comprises:

[0069] - at least one mass spectrometry device configured for providing at least one mass spectrometric response curve;

[0070] - at least one evaluation device configured for evaluating the mass spectrometric response curve by using at least one trained model, thereby identifying a start point and an end point of at least one peak of the mass spectrometric response curve, wherein the model is trained using a deep learning regression architecture.

[0071] The device can be configured for performing the method for identifying at least one peak in a mass spectrometric response curve according to any one of the preceding embodiments relating to the method. For the definition of features of the device and optional features of the device, reference can be made to one or more of the embodiments of the method as disclosed above or as disclosed in further detail below.

[0072] As generally used herein, the term "evaluation device" is a broad term and is to be given its ordinary and customary meaning to a person of ordinary skill in the art and is not to be limited to a special or customized meaning. The term specifically can refer, without limitation, to any device configured for performing a specified operation. The evaluation device can comprise at least one processing unit. The processing unit can be any logic circuitry configured for performing basic operations of a computer or system; and / or generally a device configured for performing a calculation or a logic operation. In particular, the processing unit can be configured for processing basic instructions that drive the computer or system. As an example, the processing unit can comprise at least one arithmetic logic unit (ALU), at least one floating point unit (FPU) such as a math coprocessor or a numeric coprocessor, a plurality of registers, in particular registers configured for providing operands to the ALU and storing the results of the operations, and a memory such as LI and L2 cache memory. In particular, the processing unit can be a multi-core processor. Specifically, the processing unit can be or can comprise a central processing unit (CPU). Additionally or alternatively, the processing unit can be or can comprise a microprocessor, thus, in particular, the elements of the processing unit can be contained in one single integrated circuit (IC) chip. Additionally or alternatively, the processing unit can be or can comprise one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs) or the like. The processing unit specifically can be configured, such as by software programming, for performing one or more evaluation operations.

[0073] The evaluation device can be configured for performing step b) of the method according to the present application as described in detail above or in more detail below. The evaluation device can further be configured for performing the training steps described in detail above or in more detail below.

[0074] Summarizing and without excluding further possible embodiments, the following embodiments can be envisaged:

[0075] Embodiment 1 : A computer-implemented method for identifying at least one peak in a mass spectrometric response curve, the method comprising the steps of:

[0076] a) providing at least one mass spectrometric response curve by using at least one mass spectrometric device;

[0077] b) evaluating the mass spectrometric response curve by using at least one trained model, thereby identifying a start point and an end point of at least one peak of the mass spectrometric response curve, wherein the model is trained using a deep learning regression architecture.

[0078] Embodiment 2: The method according to the preceding embodiment, wherein the deep learning regression architecture is a convolutional neural network model.

[0079] Embodiment 3: The method according to the preceding embodiments, wherein the convolutional neural network is a multi-layer convolutional neural network.

[0080] Embodiment 4: The method according to the preceding embodiments, wherein the convolutional neural network comprises a plurality of convolutional layers, wherein the convolutional layers are one-dimensional layers.

[0081] Embodiment 5: The method according to any one of the two preceding embodiments, wherein the convolutional neural network comprises a regression layer as a final layer, wherein the regression layer has a linear or sigmoid activation.

[0082] Embodiment 6: The method according to any one of the three preceding embodiments, wherein the convolutional neural network comprises a plurality of pooling layers.

[0083] Embodiment 7: The method according to any one of the preceding embodiments, wherein the method comprises at least one training step, wherein the training step comprises the following sub-steps:

[0084] i) providing at least one training data set comprising a plurality of input mass spectrometric response curves and corresponding true label values;

[0085] ii) determining at least one model by using a deep learning regression architecture on the training data set, wherein the determination of the model comprises determining a model architecture and at least one parameter of the model.

[0086] Embodiment 8: The method according to the preceding embodiments, wherein the training data set is provided in form of a five-channel vector comprising an aggregation time vector, two analyte mass spectrometric response curves and two internal standard mass spectrometric response curves.

[0087] Embodiment 9: The method according to any one of the two preceding embodiments, wherein the training of the model comprises at least one normalization step and / or at least one augmentation step.

[0088] Embodiment 10: The method according to any one of the two preceding embodiments, wherein the training step further comprises at least one testing step using at least one test data set, wherein the testing step comprises determining the onset and the termination of at least one peak of the mass spectrometric response curves of the test data set using the determined model, wherein the performance of the determined model is determined based on the determined onset and termination of the mass spectrometric response curves and the true label values of the test data set.

[0089] Embodiment 11 : The method according to any one of the preceding embodiments, wherein the method comprises determining the peak area of a peak of a mass spectrometric response curve by using the identified onset and termination.

[0090] Embodiment 12: Computer program for identifying at least one peak in a mass spectrometric response curve, which, when executed by a computer or a computer network, causes the computer or the computer network to carry out, entirely or partially, the method for identifying at least one peak in a mass spectrometric response curve according to any one of the preceding embodiments relating to a method, wherein the computer program is configured to carry out at least step b) of the method for identifying at least one peak in a mass spectrometric response curve according to any one of the preceding embodiments relating to a method.

[0091] Embodiment 13: Computer program product having program code means, wherein the program code means are stored on or in a storage medium for, when said program code means are executed on a computer or on a computer network, carrying out at least step b) of the method for identifying at least one peak in a mass spectrometric response curve according to any one of the preceding embodiments relating to a method.

[0092] Embodiment 14: An apparatus for monitoring at least one analyte in a sample, comprising:

[0093] at least one mass spectrometric device configured for providing at least one mass spectrometric response curve;

[0094] at least one evaluation device configured for evaluating the mass spectrometric response curve by using at least one trained model, thereby identifying a start point and an end point of at least one peak of the mass spectrometric response curve, wherein the model is trained using a deep learning regression architecture.

[0095] Embodiment 15: The apparatus according to the preceding embodiments, wherein the apparatus is configured for carrying out the method for identifying at least one peak in a mass spectrometric response curve according to any one of the preceding embodiments relating to a method. BRIEF DESCRIPTION OF DRAWINGS

[0096] Further optional features and embodiments will be disclosed in more detail in the subsequent embodiment description. Therein, individual optional features can be realized in a separate manner as well as in any arbitrary feasible combination, as will be appreciated by the skilled person. The scope of the present application is not limited to the preferred embodiments. Embodiments are schematically depicted in the accompanying drawings. Therein, same reference signs in these drawings refer to same or functionally equivalent elements.

[0097] In the drawings:

[0098] Figure 1 An embodiment of a computer-implemented method for identifying at least one peak in a mass spectrometric response curve according to the present application is shown;

[0099] Figure 2An embodiment of a device for monitoring at least one analyte in a sample according to the present invention is shown;

[0100] Figure 3 An embodiment of a deep learning regression architecture is shown;

[0101] Figure 4 An embodiment of a training step according to the present invention is shown;

[0102] Figures 5A to 5C An embodiment of a normalization and cropping is shown; and

[0103] Figures 6A to 6C An embodiment of an enhancement and scaling is shown. DETAILED DESCRIPTION

[0104] Figure 1 A flowchart of an embodiment of a computer-implemented method 110 for identifying at least one peak in a mass spectrometric response curve according to the present invention is shown highly schematically. The method can comprise the following steps:

[0105] a) (denoted with reference sign 112) providing at least one mass spectrometric response curve by using at least one mass spectrometric device 114;

[0106] b) (denoted with reference sign 116) evaluating the mass spectrometric response curve by using at least one trained model, thereby identifying a start point and an end point of at least one peak of the mass spectrometric response curve, wherein the model is trained 118 using a deep learning regression architecture.

[0107] Figure 2 An embodiment of a device 120 for monitoring at least one analyte in a sample according to the present invention (comprising the mass spectrometric device 114) is shown. The mass spectrometric device 114 can be configured for detecting at least one analyte based on mass-to-charge ratio. The mass spectrometric device 114 can be or can comprise at least one quadrupole analyzer comprising at least one quadrupole 122 as mass filter. The quadrupole mass analyzer can comprise a plurality of quadrupoles. For example, the quadrupole mass analyzer can be a triple quadrupole mass spectrometer. The quadrupole mass analyzer can comprise at least one power supply circuit (not shown) configured for applying at least one direct current (DC) voltage and at least one alternating current (AC) voltage between two pairs of electrodes of the quadrupole. The power supply circuit can be configured for keeping each pair of opposite electrodes at the same potential. The power supply circuit can be configured for periodically changing the charge sign of the electrode pairs such that only ions within a certain mass-to-charge ratio m / z range can have stable trajectories. The trajectories of ions within the mass filter can be described by Mathieu differential equations. For measuring ions having different m / z values, the DC voltage and the AC voltage can be adjusted in time such that ions having different m / z values are transmitted to a detector 124 of the mass spectrometric device 114.

[0108] The mass spectrometry device 114 can further comprise at least one ionization source 126. The ionization source 126 can be configured for generating ions, e.g. from neutral gas molecules. The ionization source 126 can be or can comprise at least one source selected from the group consisting of at least one gas phase ionization source, such as at least one electron impact (El) source or at least one chemical ionization (CI) source; at least one desorption ionization source, such as at least one plasma desorption (PDMS) source, at least one fast atom bombardment (FAB) source, at least one secondary ion mass spectrometry (SIMS) source, at least one laser desorption (LDMS) source, and at least one matrix assisted laser desorption (MALDI) source; at least one spray ionization source, such as at least one thermal spray (TSP) source, at least one atmospheric pressure chemical ionization (APCI) source, at least one electrospray (ESI), and at least one atmospheric pressure ionization (API) source.

[0109] The mass spectrometry device 114 can comprise at least one detector 124. The detector can be configured for detecting input ions. The detector 124 can be configured for detecting charged particles. The detector 124 can be or can comprise at least one electron multiplier.

[0110] The mass spectrometry device 114, in particular the detector 124 and / or the at least one evaluation device 128 of the mass spectrometry device 114, can be configured for determining at least one mass spectrum of the detected ions. The mass spectrum can be a pixelated image. For determining the resulting intensity of a mass spectrum pixel, the signal detected with the detector within a certain m / z range can be integrated. Analytes in the sample can be identified by the at least one evaluation device 128. In particular, the evaluation device 128 can be configured for correlating known masses with the identified masses, or by characteristic fragmentation patterns.

[0111] The mass spectrometry device 114 can specifically be or can comprise a liquid chromatography mass spectrometry device. The mass spectrometry device 114 can be connected to and / or can comprise at least one liquid chromatograph. The liquid chromatograph can be used as sample preparation for the mass spectrometry device 114. Other embodiments of sample preparation are possible, such as at least one gas chromatograph. The mass spectrometry device 114 can comprise at least one liquid chromatograph. The liquid chromatography mass spectrometry device can be or can comprise at least one high performance liquid chromatography (HPLC) device or at least one microfluidic liquid chromatography (pLC) device. The liquid chromatography mass spectrometry device can comprise a liquid chromatography (LC) device and a mass spectrometry (MS) device, in the present case a mass filter, wherein the LC device and the mass filter are coupled via at least one interface. The interface coupling the LC device and the MS device can comprise an ionization source configured for generating molecular ions and for transferring the molecular ions into a gas phase. The interface can further comprise at least one ion mobility module arranged between the ionization source and the mass filter. For example, the ion mobility module can be a high-field asymmetric waveform ion mobility spectrometry (FAIMS) module.

[0112] The LC device can be configured for separating one or more target analytes of a sample from other components of the sample for detecting the one or more analytes using the mass spectrometry device 114. The LC device can comprise at least one LC column. For example, the LC device can be a single column LC device or a multi-column LC device having a plurality of LC columns. The LC column can have a stationary phase through which a mobile phase is pumped in order to separate and / or elute and / or transfer the target analytes. The liquid chromatography mass spectrometry device can further comprise a sample preparation station for automated pre-treatment and preparation of samples, each sample comprising at least one target analyte.

[0113] The sample can be any test sample, such as a biological sample and / or an internal standard sample. The sample can comprise one or more target analytes. For example, the test sample can be selected from the group consisting of physiological fluids, including blood, serum, plasma, saliva, ocular lens fluid, cerebrospinal fluid, sweat, urine, milk, ascites, mucus, synovial fluid, peritoneal fluid, amniotic fluid, tissue, cells, etc. The sample can be used directly as obtained from the respective source or can be pre-treated and / or subjected to a sample preparation workflow. The sample can be pre-treated by adding an internal standard and / or by dilution with another solution and / or by mixing with a reagent or the like. For example, generally, the target analytes can be vitamins D, drugs of abuse, therapeutic drugs, hormones and metabolites. The internal standard sample can be a sample comprising at least one internal standard substance having a known concentration. For respective further details regarding the sample, reference is made to, for example, EP 3425369 A1, the entire disclosure of which is hereby incorporated by reference. Other target analytes are possible.

[0114] The mass spectrometric response curve provided in step a) 112 can be a one-dimensional representation of signal intensity. The mass spectrometric response curve has only one dimension. The mass spectrometric response curve can be provided by determining and / or generating and / or making available the mass spectrometric response curve, in particular by performing at least one measurement with the mass spectrometric device 114.

[0115] The evaluation of the mass spectrometric response curve in step b) 116 can comprise performing at least one analysis of the mass spectrometric response curve. The evaluation 116 can comprise identifying at least one peak and / or determining a start point and an end point of the peak and / or determining a peak area of the peak. The evaluation 116 can comprise applying at least one filter and / or using a baseline subtraction technique and / or using at least one fitting routine, etc.

[0116] The peak of the mass spectrometric response curve can be at least one local maximum of the mass spectrometric response curve. The start point of the peak can be a lower peak boundary. The start point can be a point of the time axis defining the lower peak boundary. After the start point, the mass spectrometric response curve rises to the local maximum. The start point can be a point at which the integration of the peak starts. The end point of the mass spectrometric response curve can be an upper peak boundary. The end point can be a point of the time axis defining the upper peak boundary. The mass spectrometric response curve falls before reaching the noise and / or baseline level at the end point. The end point can be a point at which the integration of the peak ends. The start point and the end point can be points on the time axis identified as peak limits. The values of the start point and the end point for a training data set (to be described in more detail below) can be determined by a trained user by manual evaluation. The trained model can provide the start point and the end point for further data. The peak area can generally be defined as the integral of the response curve between the start point and the end point.

[0117] Step b) 116 can comprise the identification of the peak. The identification of the peak can comprise a qualitative determination of the peak (such as presence or absence) and / or a quantitative determination of the peak (such as determining a peak area of the peak). The determination of the peak area can comprise a peak integration, in particular by using at least one mathematical operation and / or mathematical algorithm to determine the peak area enclosed by the peak of the mass spectrometric response curve. In particular, the integration of the peak can comprise an identification and / or measurement of the curve characteristics of the mass spectrometric response curve. The peak identification comprises determining the start point and / or the end point. The peak identification can further comprise one or more of a peak detection, a peak finding, a peak fitting, a peak evaluation, a baseline determination, and a baseline determination. The peak integration can allow determining one or more of a peak area, a retention time, a peak height, and a peak width. The peak identification can be an automatic peak identification, i.e. a peak identification performed by at least one computer and / or computer network and / or machine. In particular, the automatic peak identification can be performed without manual operation or interaction with a user.

[0118] The trained model used in step b) 116 can be or can comprise a model for identifying peaks in a mass spectrometric response curve, which model is trained on at least one training data set, also referred to as training data. In particular, the trained model is trained on existing data that is expert-prior classified. This enables to provide an automated peak identification with enhanced reliability and less susceptibility to variations and errors. The trained model can comprise an architecture and a set of weights for various filters or nodes defined by the architecture. The architecture of the CNN can reflect the complex relationship between the shape of the response curve and the start and end position of the peaks.

[0119] The method 110 can comprise at least one training step 130, wherein in the training step 130 the trained model is trained on at least one training data set. The training step 130 can be an offline training, while the identification of peaks in step b) 116 of the proposed method can be an online identification of peaks. In particular, the training step 130 can be performed before performing steps a) 112 and b) 116.

[0120] The model is trained using a deep learning regression architecture 118. The deep learning regression architecture 118 can be a deep learning architecture configured for solving a regression problem.

[0121] The deep learning regression architecture 118 can comprise a convolutional neural network. The convolutional neural network can be a multi-layer convolutional neural network. Figure 3 An embodiment of the deep learning regression architecture 118 is shown. The convolutional neural network can comprise at least one feature learning portion 132. The feature learning portion can comprise a plurality of convolutional layers 134. A convolutional layer 134 is a one-dimensional layer, i.e. the convolution is applied to a one-dimensional time domain. Each convolutional layer 134 can comprise at least one rectified linear unit, also referred to as ReLU. The convolutional neural network 118 can comprise a plurality of pooling layers 136. The structure of a convolutional neural network is generally known to the person skilled in the art, such as from https: / / en.wikipedia.org / wiki / Convolutional_neural_network#Convolutional. The use of a convolutional neural network for object recognition in images is generally known to the person skilled in the art. However, the present invention proposes a new approach for one-dimensional signal analysis using a convolutional neural network.

[0122] The convolutional neural network 118 can be configured for solving a regression problem. For solving a regression problem, the convolutional neural network 118 can comprise a regression layer 138 as a final layer, in particular in contrast to the usual classification softmax layer. The convolutional neural network 118 can comprise at least one flattening layer 140.

[0123] The regression layer 138 can be a fully connected layer 139. The regression layer 138 can have a linear or sigmoid activation. Thus, the present invention proposes to use a one-dimensional convolutional neural network in a regression framework. Specifically, the present invention proposes to use a convolutional neural network 118 to fit a complex function that maps the input to the peak position. However, convolutional neural networks 118 can not be used to classify one-dimensional signals.

[0124] Figure 4 An embodiment of the training step 130 is shown. The training step can comprise the following sub-steps:

[0125] i) (denoted with reference sign 142) providing at least one training data set comprising a plurality of input mass spectrometric response curves and corresponding ground truth values;

[0126] ii) (denoted with reference sign 144) determining at least one model by using a deep learning regression architecture 118 on the training data set, wherein the determination of the model comprises determining a model architecture of the model and at least one parameter of the model.

[0127] Step i) 142 can comprise providing more than 100, preferably more than 1000 input mass spectrometric response curves. For example, the model can be trained using 1270 vitamin D3 curves. The plurality of input mass spectrometric response curves provided in step i) 142 can be determined by performing a plurality of measurements using a mass spectrometry device. For example, the plurality of input mass spectrometric response curves provided in step i) 142 can be or can comprise LC-MS data from a specific analyte, such as LC-MS data from vitamin D2 or from vitamin D3. The ground truth values can be actual or true values of the start and end points of the peaks of the corresponding input mass spectrometric response curves. The ground truth values can be indicative of the position of the peaks. The ground truth values can be provided by a trained LC-MS operator.

[0128] The injection into the mass spectrometry device 114 can result in four mass spectrometry response curves, namely two analyte mass spectrometry response curves and two internal standard mass spectrometry response curves. The training data set can be provided in the form of a five-channel vector comprising the aggregated time vector, the two analyte mass spectrometry response curves, and the two internal standard mass spectrometry response curves. The raw time vector of the four response curves comprises time steps that at least slightly deviate from curve to curve. The aggregated time vector can be a time vector in which all four response curves are interpolated on the same time grid. The training data set can be provided as input (also referred to as input data) to the convolutional neural network. For example, given an input mass spectrometry response curve of length N, the input can be a 2xN matrix, where the first row represents N intensity values, such as the y-values of a chromatogram, and the second row can represent N time values, such as the x-values of a chromatogram. The aggregation can enable the convolutional neural network to propagate information between different mass spectrometry response curves, such that, for example, in case a peak on an analyte curve is particularly weak, the peak position of other curves can be used to inform the position of the weak curve.

[0129] The training 130 of the model can comprise at least one normalization step 146. The normalization step 146 can comprise normalizing the input data with respect to time. The normalization step 146 can comprise shifting the time values such that the expected retention time can be at t=0, Figure 5B The normalization step 146 can comprise cropping the input data to a fixed time window around the expected retention time. The normalization step 146 can comprise normalizing the intensity values Y themselves by:

[0130]

[0131] Figures 5A to 5C An embodiment of the normalization and cropping is shown. Figure 5A The intensity I is shown as a function of the time (in minutes) of the raw input data. Figure 5B The input data after time normalization is shown, and Figure 5C The input data after further normalization and cropping of the intensity is shown. The circles 148 represent the true annotation values for the peak start, and the corresponding circles 150 represent the true annotation values for the peak end position.

[0132] The training 130 of the model can comprise at least one augmentation step 152. To achieve generalization across analytes, the convolutional neural network can be trained using augmented data that takes into account shift and scaling differences in the mass spectrometric response curve data. The augmentation step 152 can avoid overfitting. The augmentation step 152 can comprise a position augmentation and / or a scaling augmentation. The position augmentation can comprise shifting the peak position by a predetermined constant value. The position augmentation can comprise using a sliding window. This can take into account the possibility of peaks broadening and narrowing. The scaling augmentation can comprise scaling the peak value by a predetermined value, such as 1.2. For each input mass spectrometric response curve, the training data set can be supplemented by randomly generating a new data set using the position augmentation and randomly generating three new data sets using the scaling augmentation. Figure 6B The original input data is shown with true label values 148 and 150. Figure 6A The position augmentation is shown, where the original curve 154 and the augmented curve 156 are depicted. Figure 6C The scaling augmentation by a factor of 1.2 is shown.

[0133] The deep learning regression architecture 118 can be a convolutional neural network (CNN) built in Python with TensorFlow as backend by the Keras library. For the Keras library in Python with TensorFlow, see https: / / www.tensorflow.org / , https: / / de.wikipedia.org / wiki / TensorFlow and https: / / keras.io / or https: / / de.wikipedia.org / wiki / Keras. For example, the following settings can be used: The adaptive moment estimation (Adam) can be used as optimizer. The loss function can be the mean squared error. The number of epochs can be 500, the batch size 16 and the patience early stopping 100.

[0134] The training step 130 can comprise using the deep learning regression architecture 118 on the normalized and / or augmented training data set together with the true label values. As output 158, the model architecture of the model and at least one parameter of the model can be provided.

[0135] The training step 130 can further comprise at least one testing step 160 using at least one test data set. The testing step 160 can comprise a validation of the trained model. The test data set can comprise at least one, preferably a plurality of mass spectrometric response curves and corresponding true label values not included in the training data set. The testing step 160 can comprise determining the onset and the end of at least one peak of the mass spectrometric response curves of the test data set using the determined model. The performance of the determined model can be determined based on the determined onset and end of the mass spectrometric response curves of the test data set and the true label values.

[0136] In the experimental setup, the deep learning regression architecture 118 was trained on 1462 curves, specifically fragment curves, from chromatograms of 488 different samples. The chromatograms contained true labeled values. For the final validation, 10% or 49 samples were reserved for the testing step 160. For the cross-validation, 439 samples were selected from the 488 samples to train the algorithm for the testing step 160. For the cross-validation, the 439 samples were divided into 5 groups. Four of the 5 groups were used for training 130 and one group was used for testing 160. All possible permutations occurred. The method 110 according to the present application was found to generalize well to analytes and improve the accuracy of the LC-MS, and thus is expected to enhance the reliability of automatic peak identification.

[0137] The following table shows the performance (peak position R 2 ) when the model was trained on a specific analyte and measurement system (denoted as enhanced data system 2) and applied to another analyte and measurement system (denoted as enhanced data system 1). It contains similar analytes, vitamin D2 and vitamin D3, and one completely different substance, testosterone.

[0138]

[0139] List of reference signs

[0140] 110 method

[0141] 112 step a)

[0142] 114 mass spectrometry device

[0143] 116 step b)

[0144] 118 deep learning regression architecture

[0145] 120 device

[0146] 122 quadrupole

[0147] 124 detector

[0148] 126 ionization source

[0149] 128 evaluation device

[0150] 130 training step

[0151] 132 feature learning part

[0152] 134 convolutional layer

[0153] 136 pooling layer

[0154] 138 regression layer

[0155] 139 fully connected layer

[0156] 140 flatten layer

[0157] 142 step i)

[0158] 144 step ii)

[0159] 146 normalization step

[0160] 148 circle

[0161] 150 circle

[0162] 152 enhancement step

[0163] 154 original curve

[0164] 156 enhanced curve

[0165] 158 output

[0166] 160 test step.

Claims

1. A computer-implemented method (110) for identifying at least one peak in a mass spectrometry response curve, the method comprising the following steps: a) Provide at least one mass spectrometry response curve by using at least one mass spectrometry device (114); b) Identifying the start and end points of at least one peak of the mass spectrometry response curve by evaluating (116) the mass spectrometry response curve using at least one trained model, wherein the model is trained using a deep learning regression architecture (118). The deep learning regression architecture (118) includes a convolutional neural network and the convolutional neural network includes a regression layer (138) as the final layer, the convolutional neural network includes multiple convolutional layers (134), the convolutional layer (134) is a one-dimensional layer, and the mass spectrum response curve is a one-dimensional representation of signal intensity.

2. The method (110) according to claim 1, wherein the regression layer (138) has linear or sigmoid activation.

3. The method (110) according to claim 1, wherein the method includes at least one training step (130), wherein the training step includes the following sub-steps: i) Provide at least one training dataset, which includes multiple input mass spectrometry response curves and corresponding ground-labeled values; ii) Determine at least one model by using the deep learning regression architecture (118) on the training dataset, wherein determining the model includes determining the model architecture and at least one parameter of the model.

4. The method (110) according to claim 3, wherein the training dataset is provided in the form of a five-channel vector including an aggregated time vector, two analytical mass spectrometry response curves and two internal standard mass spectrometry response curves.

5. The method (110) according to claim 3, wherein the training (130) of the model includes at least one normalization step (146) and / or at least one enhancement step (152).

6. The method (110) according to claim 4 or 5, wherein the training step (130) further comprises at least one testing step (160) using at least one test dataset, wherein the testing step (160) comprises using a determined model to determine the start and end points of at least one peak of the mass spectrometry response curve of the test dataset, wherein the performance of the determined model is determined based on the determined start and end points of the mass spectrometry response curve of the test dataset and the true labeled values.

7. The method (110) of claim 1, wherein the method includes determining the peak area of ​​the peak of the mass spectrometry response curve by using an identified start point and end point.

8. A computer-readable storage medium comprising instructions that, when executed, cause a computer or computer network to perform the method according to any one of claims 1-7.

9. An apparatus (120) for monitoring at least one analyte in a sample, the apparatus comprising: - At least one mass spectrometry device (114), said at least one mass spectrometry device being configured to provide at least one mass spectrometry response curve; - At least one evaluation device (128) configured to evaluate the mass spectrometry response curve by using at least one trained model, thereby identifying the start and end points of at least one peak of the mass spectrometry response curve, wherein the model is trained using a deep learning regression architecture (118). The deep learning regression architecture (118) includes a convolutional neural network and the convolutional neural network includes a regression layer (138) as the final layer, the convolutional neural network includes multiple convolutional layers (134), the convolutional layer (134) is a one-dimensional layer, and the mass spectrum response curve is a one-dimensional representation of signal intensity.

10. The apparatus (120) according to claim 9, wherein the apparatus (120) is configured to perform the method according to any one of claims 1-5.

11. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Automated clinical diagnostic system and method

    EP3425369A1

  • Chromatograph device

    EP3467493A1

  • Automatic peak identification method

    US7219038B2

  • Methods for resolving convoluted peaks in a chromatogram

    US7720612B2

  • Methods of automated spectral peak detection and quantification having learning mode

    US8428889B2