Method for automatic quality check of chromatographic and / or mass spectral data

The method uses analyte-specific machine learning models to automate mass spectrometry data quality checking, addressing manual review inefficiencies by accurately classifying data reliability, thus enhancing data analysis efficiency and accuracy.

JP2026010113APending Publication Date: 2026-01-21F HOFFMANN LA ROCHE & CO AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025174545
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-06
Filing Date
2025-10-16
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Current mass spectrometry data processing requires significant manual review due to high error rates, with existing automated methods relying on limited sample sizes and specific laboratory settings, failing to fully replace manual peak review.

Method used

A computer-implemented method using analyte-specific trained machine learning models to classify chromatographic and mass spectral data quality, eliminating the need for manual inspection by predicting peak integration failures through regression models trained on historical and semi-synthetic datasets.

Benefits of technology

Automated quality checking reduces manual review by accurately classifying data as reliable or unreliable, improving efficiency and accuracy in mass spectrometry data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026010113000001_ABST
    Figure 2026010113000001_ABST
Patent Text Reader

Abstract

To provide a method for automatic quality check of chromatographic data and / or mass spectral data.SOLUTION: Providing processed chromatographic and / or mass spectral data obtained by a mass spectrometer, and classifying the quality of the chromatographic and / or mass spectral data by applying a trained machine learning model to the chromatographic and / or mass spectral data, wherein the trained machine learning model uses one regression model. The trained machine learning model is trained on a training dataset comprising historical chromatographic and / or mass spectral data and / or semi-synthetic chromatographic and / or mass spectral data. The trained machine learning model is an analyte-specific machine learning model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method, a test system, a computer program and a computer program product for automated quality check of chromatographic and / or mass spectral data. The proposed method and device can be used in the technical field of mass spectrometry, in particular in liquid chromatography-mass spectrometry. [Background technology]

[0002] Current mass spectrometry (MS) data processing typically requires manual data review of all acquired data and subsequent manual correction of approximately 5–20% of results due to a high error rate. This is performed by trained operators through tedious visual analysis of hundreds of plots. Manually flagging unreliable data obtained using MS instruments such as liquid chromatography coupled to mass spectrometry (LC-MS) or tandem mass spectrometry (LC-MS / MS) is time-consuming. However, only a few solutions exist for flagging unreliable results generated by automated peak integration. The proposed method aims to reduce the amount of manual review by focusing on problematic results. However, a significant portion of the data still must be corrected and, in some cases, manually reintegrated.

[0003] Furthermore, some of these methods rely on machine learning techniques, which rely on real-world training datasets that are tailored to specific laboratory settings, subjectively labeled as "good" or "bad," and have limited sample sizes.

[0004] For example, www.indigobio.com / ascent / describes the ASCENT peak processor, which may be implemented when manual inspection is still necessary. ASCENT notifies peaks to be reviewed and presents a set of flags to focus on the peaks. This approach may reduce manual peak review, but does not replace it. Similarly, Yu M, Bazydlo LAL, Bruns DE, Harrison JH Jr., “Streamlining Quality Review of Mass Spectrometry Data in the Clinical Laboratory by Use of Machine Learning,” in Learning,” Arch Pathol Lab Med. 2019 Aug;143(8):990-998. doi:10.5858 / arpa.2018-0238-OA describes determining whether classification models created using standard machine learning algorithms can validate analytically acceptable MS results, thereby reducing manual review requirements. The proposed technique may reduce, but does not replace, manual peak review. Toghi Eshghi S, Auger P, Mathews WR, “Quality assessment and interference detection in targeted mass spectrometry data using machine learning,” Clin Proteomics. 2018 Oct 6;15:33. doi:10.1186 / s12014-018-9209-x describes an algorithm that utilizes supervised machine learning to identify peaks with interferences or chromatographic defects based on a set of peaks annotated by expert analysts. Analyzing targeted proteomics data using TargetedMSQC reduces the time spent manually inspecting peaks and improves both the speed and accuracy of interference detection. ,The proposed technique may reduce manual peak review, but does not replace it. [Prior art documents] [Non-patent literature]

[0005] Non-patent document 1: www.indigobio.com / ascent / Non-patent document 2: Yu M, Bazydlo LAL, Bruns DE, Harrison JH Jr., "Streamlining Quality Review of Mass Spectrometry Data in the Clinical “Laboratory by Use of Machine Learning”, Arch Pathol Lab Med.2019 August;143(8):990-998.doi:10.5858 / arpa.2018-0238-OA Non-Patent Document 3: Toghi Eshghi S, Auger P, Mathews WR, “Quality assessment and interference detection in targeted mass spectrometry data using machine learning,” Clin Proteomics. 2018 October 6;15:33. doi:10.1186 / s12014-018-9209-x Summary of the Invention [Problem to be solved by the invention]

[0006] It is therefore an object of the present invention to provide a method, test system, computer program and computer program product for automated quality checking of chromatographic and / or mass spectral data that avoids the above-mentioned drawbacks of known methods, devices, computer programs and computer program products. In particular, a method and device are provided that allow replacing manual peak review.

[0007] overview This problem is addressed by a method, a test system, a computer program and a computer program product for automated quality check of chromatographic and / or mass spectral data having the features of the independent claims. Advantageous embodiments, which may be implemented alone or in any combination, are set out in the dependent claims as well as in the specification as a whole.

[0008] When used below, the terms "have," "comprise," or "include," or any grammatical variations thereof, are used in a non-exclusive manner. Thus, these terms may refer both to a situation in which, besides the features introduced by these terms, no further features are present in the entity described in this context, and to a situation in which one or more further features are present. As an example, the expressions "A has B," "A comprises B," and "A includes B" may refer both to a situation in which, besides B, no other elements are present in A (i.e., a situation in which A consists solely and exclusively of B), and to a situation in which, besides B, one or more further elements are present in entity A, such as element C, elements C and D, and even further elements.

[0009] Furthermore, it should be noted that the terms "at least one," "one or more," or similar expressions indicating that a feature or element may be present one or more times are typically used only once when introducing each feature or element. In the following, in most cases, when referring to each feature or element, the expressions "at least one" or "one or more" will not be repeated, despite the fact that each feature or element may be present one or more times.

[0010] Furthermore, when used hereinafter, the terms "preferably," "more preferably," "particularly," "more particularly," "particularly," "more particularly," or similar terms are used in conjunction with optional features without limiting the possibility of substitution. Thus, features introduced by these terms are optional features and are not intended to constrain the scope of the claims in any way. The present invention may also be implemented by using alternative features, as recognized by those skilled in the art. Similarly, features introduced by "in an embodiment of the invention" or similar expressions are intended to be optional features without any limitation regarding alternative embodiments of the invention, without any limitation regarding the scope of the invention, and without any limitation regarding the possibility of combining the feature introduced in this way with other optional or non-optional features of the invention. [Means for solving the problem]

[0011] In a first aspect, a computer-implemented method for automated quality checking of chromatographic and / or mass spectral data is proposed.

[0012] The term "computer-implemented method" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically, but not exclusively, refer to a method involving at least one computer and / or at least one computer network. The computer and / or computer network may comprise at least one processor configured to perform at least one of the method steps of the method according to the present invention. Preferably, each of the method steps is performed by the computer and / or computer network. The method may be performed completely automatically, specifically without user interaction. The terms "automatically" and "automated" as used herein are broad terms and should be given their ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically, but not exclusively, refer to a process performed completely by at least one computer and / or computer network and / or machine, specifically without manual action and / or user interaction.

[0013] The term "mass spectrometry data," as used herein, is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically, but is not limited to, refer to data obtained by using at least one mass spectrometer, particularly at least one mass spectrum.

[0014] The term "chromatographic data," as used herein, is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically refer, but is not limited to, data obtained by using at least one chromatographic device, such as at least one liquid chromatograph. The chromatographic data may include at least one chromatogram.

[0015] The term "mass spectrometry" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically, but is not limited to, refer to an analytical technique for determining the mass-to-charge ratio of ions. Mass spectrometry may be performed using at least one mass analyzer. As used herein, the term "mass analyzer" refers to a method for analyzing ions. The term "mass analyzer," also referred to as a "mass analyzer," is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically, but is not limited to, refer to an analyzer configured to detect at least one analyte based on mass-to-charge ratio. A mass analyzer may be or may comprise at least one quadrupole analyzer. As used herein, the term "quadrupole mass analyzer" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically, but is not limited to, refer to a mass analyzer comprising at least one quadrupole as a mass filter. A quadrupole mass analyzer may comprise multiple quadrupoles. For example, a quadrupole mass analyzer may be a triple quadrupole mass analyzer. As used herein, the term "mass filter" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically, but not exclusively, refer to a device configured to select ions for injection into a mass filter according to their mass-to-charge ratio (m / z). The mass filter may include two pairs of electrodes. The electrodes may be rod-shaped, particularly cylindrical. Ideally, the electrodes may be hyperbolic. The electrodes may be identically designed. The electrodes may be arranged to extend parallel along a common axis, e.g., the z-axis. The quadrupole mass analyzer may include at least one power supply circuit configured to apply at least one direct current (DC) voltage and at least one alternating current (AC) voltage between the two pairs of electrodes of the mass filter. The power supply circuit may be configured to maintain each opposing electrode pair at the same potential. The power supply circuit may be configured to periodically change the sign of the charge of the electrode pair to enable stable trajectories only for ions within a specific mass-to-charge ratio (m / z). The trajectories of ions in the mass filter may be described by the Mathieu differential equation.To measure ions of different m / z values, the DC and AC voltages may be varied in time to send ions with different m / z values ​​to the detector of the mass spectrometer.

[0016] The mass analyzer may further include at least one ionization source. As used herein, the term "ionization source," also referred to as "ion source," is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically refer to, but is not limited to, a device configured to generate ions from, for example, neutral gas molecules. The ionization source may be or comprise at least one source selected from the group consisting of at least one gas phase ionization source, such as at least one electron impact (EI) source or at least one chemical ionization (CI) source, at least one desorption ionization source, such as at least one plasma desorption (PDMS) source, at least one fast atom bombardment (FAB) source, at least one secondary ion mass spectrometry (SIMS) source, at least one laser desorption (LDMS) source, and at least one matrix-assisted laser desorption (MALDI) source, at least one spray ionization source, such as at least one thermospray (TSP) source, at least one atmospheric pressure chemical ionization (APCI) source, at least one electrospray (ESI) source, and at least one atmospheric pressure ionization (API) source.

[0017] A mass analyzer may include at least one detector. As used herein, the term "detector" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically, but is not limited to, refer to a device configured to detect incoming ions. A detector may be configured to detect charged particles. A detector may be or include at least one electron multiplier.

[0018] The mass analyzer, particularly the detector and / or at least one processing unit of the mass analyzer, may be configured to determine at least one mass spectrum of the detected ions. As used herein, the term "mass spectrum" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. Specifically, the term may refer, but is not limited to, a two-dimensional representation of signal intensity versus mass-to-charge ratio (m / z), where the signal intensity corresponds to the abundance of each ion. The mass spectrum may be a pixelated image. Signals detected by the detector within a specific m / z range may be integrated to determine the resulting intensity of the pixels of the mass spectrum. Analytes in the sample may be identified by the processing unit. Specifically, the processing unit may be configured to correlate known masses to identified masses or via characteristic fragmentation patterns.

[0019] The mass spectrometer may be or comprise a liquid chromatography-mass spectrometer. The mass spectrometer may be connected to and / or comprise at least one liquid chromatograph. The liquid chromatograph may be used for sample preparation for the mass spectrometer. Other embodiments of sample preparation, such as at least one gas chromatograph, may be possible. As used herein, the term "liquid chromatography-mass spectrometer" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to a special or customized meaning. The term may specifically refer to, but is not limited to, a combination of liquid chromatography and mass spectrometry. The mass spectrometer may comprise at least one liquid chromatograph. The liquid chromatography-mass spectrometer may be or comprise at least one high-performance liquid chromatography (HPLC) device or at least one micro-liquid chromatography (μLC) device. The liquid chromatography-mass spectrometer may comprise a liquid chromatography (LC) device and a mass spectrometry (MS) device, in this case a mass filter, where the LC device and the mass filter are coupled via at least one interface. The interface connecting the LC device and the MS device can include an ionization source configured to generate molecular ions and transfer the molecular ions to the gas phase. The interface can further include at least one ion mobility module disposed between the ionization source and the mass filter. For example, the ion mobility module can be a high-field asymmetric waveform ion mobility spectrometry (FAIMS) module.

[0020] As used herein, the term "liquid chromatography (LC) system" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically refer to, but is not limited to, an analytical module configured to separate one or more analytes of a sample from other components of the sample for detection of the one or more analytes using a mass spectrometer. An LC system may include at least one LC column. For example, an LC system may be a single-column LC system or a multi-column LC system having multiple LC columns. An LC column may have a stationary phase through which a mobile phase is pumped to separate and / or elute and / or transfer analytes. A liquid chromatography-mass spectrometer may further include a sample preparation station for automated sample preparation and preparation, each sample containing at least one analyte.

[0021] As used herein, the term "quality check" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may particularly refer to, but is not limited to, the process of distinguishing between reliable and unreliable automated peak integration. The quality check may include determining information on whether the peak integration process is complete, i.e., whether a calculated nominal signal is available, whether the data quality was suitable for automatic peak integration, and whether the calculated nominal signal and readouts are reliable.

[0022] The term "quality" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. Specifically, the term may refer to, but is not limited to, a measure of the reliability of automated peak integration performed on data provided by an MS instrument and / or an LC instrument. The classified quality may be used to distinguish between acceptable and unacceptable chromatographic and / or mass spectral data. Specifically, quality may be classified as good (acceptable) for reliable automated peak integration and as poor (unacceptable) for unreliable automated peak integration. The quality classification may include distinguishing between reliable and unreliable automated peak integration. Quality may depend on several factors, such as noise level, background, interferences that could not be resolved from the target peak, retention time shifts, peak width, and the presence or absence of an internal standard signal.

[0023] The method includes, by way of example, the following steps, which may be performed in the given order. However, it should be noted that different orders are possible. Furthermore, one or more of the method steps may be performed once or repeatedly. Furthermore, two or more of the method steps may be performed simultaneously or with overlapping timing. The method may include additional method steps not listed.

[0024] The method comprises: a) providing processed chromatographic and / or mass spectral data obtained by at least one mass spectrometer; b) classifying the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data, wherein the trained machine learning model uses at least one regression model, and the trained machine learning model is trained on at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model. Includes:

[0025] As used herein, the term "processed chromatographic data and / or mass spectral data" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically refer to, but is not limited to, chromatographic data and / or mass spectral data that has been subjected to automated peak integration. Regarding automated peak integration, see International Publication No. WO 2021 / 023865 A1, the entire contents of which are incorporated by reference.

[0026] As used herein, the term "providing" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term specifically refers to, but is not limited to, processing, particularly by performing at least one measurement using a mass spectrometer and subsequently processing the data. The term "processing" may refer to the process of determining and / or generating and / or making available processed chromatographic data and / or mass spectral data. Accordingly, the term "providing processed chromatographic data and / or mass spectral data," as used herein, is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. Specifically, the term may refer, without limitation, to retrieving processed chromatographic data and / or mass spectral data obtained from a mass spectrometer upon a particular receipt, and / or performing at least one measurement and processing using a mass spectrometer, thereby determining processed chromatographic data and / or mass spectral data.

[0027] The term "classifying" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art, without being limited to any special or customized meaning. This term may specifically refer to, but is not limited to, the process of classifying chromatographic and / or mass spectral data into at least two categories, for example, good or reliable for reliable automated peak integration, and bad or unreliable for unreliable automated peak integration. The classification is performed by applying at least one trained machine learning model. Thus, according to the present invention, at least one machine learning model can be used to predict peak integration failures, providing a fully automated decision regarding the release of results. Thus, the proposed method eliminates the need for manual inspection of data.

[0028] The term "machine learning model" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically, but not exclusively, refer to a mathematical model that can be trained on at least one training dataset using machine learning, particularly deep learning or other forms of artificial intelligence. The term "machine learning" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically, but not exclusively, refer to a method of using artificial intelligence (AI) to automatically build a model. The training may be performed using at least one machine learning system. The term "machine learning system" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically, but not exclusively, refer to a system or unit comprising at least one processing unit, such as a processor, microprocessor, or computer system, configured for machine learning, particularly for executing logic in a given algorithm. The machine learning system may be configured to implement and / or execute at least one machine learning algorithm, which is configured to build a trained machine learning model. The machine learning system may be part of the mass spectrometer and / or may be executed by an external processor, such as a cloud.

[0029] The trained machine learning model uses at least one regression model. As used herein, the term "regression model" is a broad term and should be given its ordinary and customary meaning to those skilled in the art, and should not be limited to any special or customized meaning. The term may specifically, but is not limited to, refer to a predictive model configured to analyze the relationship between a target variable and independent variables in a dataset. The target variable for chromatographic data may be a continuous deviation from an expected result value. For mass spectral data, the target variable may be dichotomous information regarding whether the result is valid or not. The regression model may be a random forest, e.g., as described in Breiman L., Random forests, Machine Learning, 2001, 45(1):5-32; gradient boosting forest, e.g., as described in Friedman, JH (2001); greedy function approximation, e.g., as described in "A Gradient Boosting Machine", The Annals of Statistics, 29(5):1189-1232; partial least squares, e.g., as described in Wold, H. (1985); partial least squares, e.g., as described in Kotz, Samuel, Johnson, Norman The regression model may be at least one selected from partial least squares, such as that described in Tibshirani, R. (ed.), Encyclopedia of statistical sciences, 6, New York: Wiley, pp. 581-591; Lasso regression, such as that described in Tibshirani, R. (1996), Regression Shrinkage and Selection via the lasso, Journal of the Royal Statistical Society, Series B (methodological), Wiley, 58(1):267-88; logistic regression, such as that described in Hosmer, D., Lemeshow, S.: Applied logistic regression, Wiley, New York 2000; or Bayesian regression, such as that described in Box, G.E.P., Tiao, G.C. (1973), Bayesian Inference in Statistical Analysis, Wiley. For example, the regression model may be selected from gradient boosting forests or random forests. The regression model may be, for example, gradient boosting forests. The regression model may be, for example, random forests.

[0030] The trained machine learning model is an analyte-specific trained machine learning model. For example, the analyte is at least one target substance selected from the group consisting of vitamin D, drugs of abuse, therapeutic drugs, hormones, and metabolites to be quantified from a sample. The term "sample" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically refer to any test sample, such as, but not limited to, a biological sample and / or an internal standard sample. The sample may contain one or more analytes. For example, the test sample may be selected from the group consisting of physiological fluids, including blood, serum, plasma, saliva, ocular lens fluid, cerebrospinal fluid, sweat, urine, milk, ascites, mucus, synovial fluid, peritoneal fluid, amniotic fluid, tissue, cells, etc. The sample may be used directly as obtained from its respective source or may be subjected to pretreatment and / or sample preparation workflow. For example, the sample may be pretreated by adding an internal standard and / or diluting with another solution and / or mixing with a reagent, etc. For example, the analyte may generally be vitamin D, drugs of abuse, therapeutic agents, hormones, and metabolites. The internal standard sample may be a sample containing at least one internal standard substance with a known concentration. For further details regarding samples, see, for example, EP 3425369, the entire disclosure of which is incorporated herein by reference. Other analytes are also possible.

[0031] The machine learning model may use a feature set. The feature set considered to be beneficial for data and peak integration quality may include standard MS quality parameters such as peak asymmetry or ion ratio, ratios of parameters between different transitions, such as the retention time ratio between the analyte quantifier and the internal standard quantifier, features for assessing the quality of the peak fit, such as residual ratios or peak fit uncertainty, and further design features describing noise, background, and peak shape. The feature set may include peak area, peak background, relative background, ion ratio, Q4 ratio, retention time ratio, peak asymmetry, asymmetry ratio, peak width, peak width ratio, area of ​​integration residual, confidence interval of peak area, mass shift, full width at half maximum. The features may include at least one feature selected from the group consisting of: signal-to-noise ratio, single-cycle median ratio, single-cycle median ion ratio, peak height, peak fit mean square error, fit intensity correlation, Earth Mover's Distance, and the deviation of any of the mentioned features when derived from the processed data, i.e., integrated peaks and raw data, e.g., the difference in retention time between the fitted peak and the raw signal. Peak background may refer to the intensity of the estimated background in the peak interval. Relative background may refer to the ratio of the peak background to the peak height. Ion ratio may refer to the area of ​​the analyte or internal standard (ISTD) quantifier relative to the area of ​​the analyte or the area of ​​the ISTD quantifier. The Q4 ratio may be given by Q4 = (area of ​​analyte quantifier / area of ​​analyte quantifier) / (area of ​​ISTD quantifier / area of ​​ISTD quantifier). The retention time ratio may refer to one or more of RT_analyte_qualifier / RT_analyte_quantifier, RT_IStd_qualifier / RT_ISTD_quantifier, or RT_analyte_quantifier / RT_ISTD_quantifier, where RT_analyte_qualifier is the retention time of the analyte qualifier, RT_analyte_quantifier is the retention time of the analyte quantifier, RT_ISTD_qualifier is the retention time of the ISTD qualifier, and RT_ISTD_quantifier is the retention time of the ISTD quantifier. Peak asymmetry may be defined according to USP40 guidelines (also referred to herein as USP40), see http: / / pharmacopeia.cn / v29240 / usp29nf24s0_c621_viewall.html, in particular Figure 2.The asymmetry ratio may refer to one or more of asymmetry_analyte_qualifier / asymmetry_analyte_quantifier, asymmetry_ISTD_qualifier / asymmetry_ISTD_quantifier, or asymmetry_analyte_quantifier / asymmetry_ISTD_quantifier, where asymmetry_analyte_qualifier is the asymmetry of the peak of the analyte qualifier, asymmetry_ISTD_qualifier is the asymmetry of the peak of the ISTD qualifier, and asymmetry_ISTD_qualifier is the asymmetry of the peak of the ISTD quantifier. The peak width ratio may refer to one or more of width_analyte_qualifier / width_analyte_quantifier, width_ISTD_qualifier / width_ISTD_quantifier, or width_analyte_quantifier / width_ISTD_quantifier, where width_analyte_qualifier is the peak width of the analyte qualifier, width_analyte_quantifier is the peak width of the analyte quantifier, width_ISTD_qualifier is the peak width of the ISTD qualifier, and width_ISTD_quantifier is the peak width of the ISTD quantifier. The signal-to-noise ratio may be defined according to USP40. The single-cycle median ratio may refer to the median ratio of the intensity of the analyte quantifier to the intensity of the ISTD quantifier. The single-cycle median ion ratio may refer to one or more medians of the ratio of the intensity of the analyte quantifier to the intensity of the analyte qualifier or the ratio of the intensity of the ISTD quantifier to the intensity of the ISTD qualifier. The peak fit mean squared error may be given by the mean [(smoothed intensity / area of ​​fitted intensity / area)2]. Fit intensity correlation may refer to one or more of cor(smoothed intensity, fitted intensity) or cor(preprocessed intensity, fitted intensity).For information on Earthmover distance, see, for example, https: / / en.wikipedia.org / wiki / Earth_mover%27s_distance. A rich set of features can be derived from the chromatography data and / or mass spectrometry data and used to build a regression model. Training the model can include determining feature rankings. Training the model can include selecting features.

[0032] The features of the feature set may be combined in a machine learning model to predict area ratio deviations as equivalent to peak integration failures. Regression models, such as random forests and gradient boosting, have been found to perform well with reasonable model complexity in terms of evaluation time and required disk space. Model parameters, such as the type of algorithm, the number of features, and the number and size of trees, may be adjusted by resampling techniques.

[0033] In the case of random forests, it has been found that increasing the number of features improves the performance of random forests. In the case of gradient boosting forests, it has been found that decreasing the number of features improves the performance of gradient boosting forests. Feature selection may be performed to select top features that are "stable" across many data splits and / or models. The method may also include feature engineering, including evaluation of newly created features. For example, in the case of gradient boosting forests, 50 features may be used with a minimum leaf size of 50 and 400 trees.

[0034] The regression model result in step b) may be a percentage deviation of the area ratio from a known true value. For classification, at least one threshold may be used to generate a binary result for classification. If the regression model result is greater than the threshold, the data may be classified as bad; otherwise, if the regression model result is below the threshold, the data may be classified as good. For example, the threshold may be 10%.

[0035] The method may comprise at least one training step, step c), which may comprise training a machine learning model based on the training dataset.

[0036] The term "training" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art, and should not be limited to any special or customized meaning. This term may specifically, but is not limited to, refer to the process of constructing a trained machine learning model, particularly the process of determining the model's parameters, especially the weights. Training may include determining and / or updating the model's parameters. A trained machine learning model may be at least partially data-driven. As used herein, the term "at least partially data-driven model" is a broad term and should be given its ordinary and customary meaning to those skilled in the art, and should not be limited to any special or customized meaning. This term may specifically, but is not limited to, refer to the fact that a model includes a data-driven model portion and another model portion based on physical chemistry laws, etc. Training may be performed on historical chromatographic and / or mass spectral data and / or semi-synthetic chromatographic and / or mass spectral data. Training may also include retraining a trained model after acquiring additional chromatographic and / or mass spectral data, for example, during operation of an MS and / or LC-MS instrument.

[0037] The trained machine learning model is trained on at least one training dataset including historical chromatographic and / or mass spectrometric data and / or semi-synthetic chromatographic and / or mass spectrometric data. The training dataset may be generated by manually classifying the historical chromatographic and / or mass spectrometric data and / or semi-synthetic chromatographic and / or mass spectrometric data into two categories.

[0038] The training step may include training the machine learning model for different analytes. Step b) may be performed during assay development for a plurality of different assays, and the trained machine learning models for the different assays are stored in at least one databank. The databank may include a data processing configuration file, enabling automatic flagging of peak integration results on the instrument. The method may also include at least one selection step performed before step b), in which one trained machine learning model is selected from trained machine learning models trained on the analyte used to obtain the provided chromatographic data and / or mass spectral data.

[0039] The trained machine learning model may be suitable for different analytes with similar chromatography. The training step may include training the machine learning model for different chromatography types. For different chromatography types, separate models may be used, for example, for standard chromatography, in which peak fitting may be applied, non-standard chromatography, in which boundary detection must be applied, and when an internal standard having exactly the same retention time as the analyte and a retention time offset exists between the analyte and the ISTD is not available.

[0040] The term "historical chromatographic data and / or mass spectral data" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically refer to, but is not limited to, measurements obtained by using at least one mass spectrometer. The historical data may be actual data. The historical chromatographic data and / or mass spectral data may measure several analytes and include data from different instruments with different scenarios. An example of a historical training dataset may include approximately 500 chromatographic measurements involving five different analytes measured on two instruments from one system and three instruments from another system during an 11-week period.

[0041] The training dataset includes semi-synthetic chromatographic data and / or mass spectral data, also referred to as a semi-synthetic dataset. As used herein, the term "semi-synthetic chromatographic data and / or mass spectral data" is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. This term may specifically refer to, but is not limited to, simulated chromatographic data and / or mass spectral data based on historical chromatographic data and / or mass spectral data. Semi-synthetic chromatographic data and / or mass spectral data may be generated by applying and / or simulating defined disturbances to actual measured chromatographic data and / or mass spectral data. The semi-synthetic chromatographic data and / or mass spectral data may include modified historical chromatographic data and / or mass spectral data. The historical chromatographic data and / or mass spectral data may be modified by one or more of introducing at least one interference, introducing background, introducing at least one shift in retention time, altering peak width, or replacing an internal standard signal with a chromatogram from a duplicate blank sample. Semi-synthetic simulation approaches combine the benefits of knowing the truth in simulation studies with the provision of datasets with real-world characteristics. Using simulated datasets for model training has several advantages over real data, including the ability to objectively define the true state of the measurements, the ability to investigate rare cases and "gray zones," and scalability in terms of sample size. To resemble the real data as closely as possible, semi-synthetic approaches are employed, where real measurements are modified in a controlled manner.

[0042] A semi-synthetic dataset may be generated as follows: A (manually curated) actual chromatogram with clear peaks and reliable integration results is selected and then modified to resemble a challenging situation for peak integration. The generation of the semi-synthetic dataset may include considering one or more of the following conditions: interference, background, retention time shift, peak width, and missing internal standard signal. For example, to account for interference, the fitted intensity of the actual internal standard peak is added to the raw intensity of the neighboring analyte peak. The distance between peaks allows for various resolutions to be explored. The height of the artificial interference peak may be expanded or contracted to simulate various relative peak heights between the peak of interest and the interference. For example, to account for background, a step function is first generated to simulate a varying background signal, and the step height is drawn from a uniform distribution. The maximum step height may control the magnitude of the simulated background. A background fit is then applied to the step function, and the resulting curve is added to the actual chromatogram intensities. A curvature parameter in the background fit allows for manipulation of the curvature of the artificial background. For example, retention time variations can be easily simulated by shifting the actual signal along the time scale to account for retention time shifts. For example, to account for peak widths, the peak fit is rescaled by changing each parameter of the fitting function. To maintain the area under the peak, the intensity is rescaled. Rescaled noise from the original data is then added to the new peak fit. For example, to account for missing internal standard signals, the chromatogram of the internal standard is replaced with the chromatogram of a double blank sample.

[0043] Simulated data, i.e., semi-synthetic chromatographic data and / or mass spectral data, may have a much higher percentage of bad cases and a much higher percentage of borderline cases than real data, i.e., historical synthetic chromatographic data and / or mass spectral data. Including a portion of the real data for training can improve model performance. Other portions of the real data may be used to test the trained model. The real data may be a manually labeled, authentic dataset.

[0044] The method may include at least one testing step, which includes testing the trained model. The testing step may include testing the trained model against at least one test dataset. The testing step may include obtaining performance characteristics of the trained model, such as accuracy, false positive rate, and false negative rate. To evaluate predictive performance, the model may be tested using simulated data and / or against real data, particularly a manually labeled real dataset. The test dataset may include simulated data and / or real data.

[0045] For example, the training data set may include a first semi-synthetic data set, such as including 7062 measurements, and the test data set may include a second semi-synthetic data set, such as including 3638 measurements.

[0046] For example, the training data set may include both a semi-synthetic data set and a portion of real data labeled "good." The training data set may also include another portion of real data labeled "good" and real data labeled "bad."

[0047] An exemplary peak shape of an analyte (e.g., Testosterone) A machine learning model was trained on a semi-synthetic dataset. The machine learning model was trained on 241 manually labeled real measurements obtained from 10 sample runs on different instruments. 121 were manually labeled as bad and 120 were manually labeled as good. A quality check of the peak integration using the trained machine learning model correctly classified all 120 "good" measurements. Five of the 121 "bad" measurements were classified as "good" by the trained machine learning model. The accuracy was determined to be 0.9793, with a false positive rate of 0.0000 and a false negative rate of 0.0413.

[0048] The trained machine learning model may then be deployed to predict the quality state of new measurements, as performed in step b). The trained machine learning models for different analytes and / or different chromatography types may be transferred to a data processing configuration file. The data processing configuration file may be stored in at least one data storage device of the mass spectrometer, which may enable automatic flagging of peak integration results on the mass spectrometer.

[0049] The method may include assigning a flag to the chromatographic data and / or mass spectral data as acceptable or unacceptable based on the classified quality. A measure of how affected the data is by the introduced "disturbing factors" may be the percentage deviation of the area ratio result calculated for the created semi-synthetic data from the area ratio of the original actual data set. The area ratio deviation represents a continuous result of the regression model. A gold standard for error handling can then be defined, for example, by flagging measurements with an area ratio deviation greater than 10%. The binary flag serves as a truth state when evaluating predictive performance in terms of accuracy and false positive / false negative rates. The method may also include providing at least one piece of information to a user via at least one user interface in response to the flag of the chromatographic data and / or mass spectral data. The term "user interface" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to a special or customized meaning. The term may refer to an element or unit configured to interact with its environment, such as, but not limited to, exchanging information unidirectionally or bidirectionally, such as to exchange one or more data or commands. For example, a user interface may be configured to share information with a user and receive information by a user. A user interface may be a feature that interacts visually with a user, such as a display, or a feature that interacts acoustically with a user. A user interface may include, by way of example, one or more of a graphical user interface, a data interface, such as a wireless and / or wired data interface.

[0050] In a further aspect, a test system is proposed, configured to carry out the method according to the invention. For definitions of the features of the test system and for any features of the test system, reference may be made to one or more of the embodiments of the method disclosed above or disclosed in more detail below. The test system may be part of a mass spectrometer.

[0051] The test system is - at least one communication interface configured to receive processed chromatographic and / or mass spectral data obtained by at least one mass spectrometer; - at least one processing device configured to classify the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data, the trained machine learning model using at least one regression model, at least one processing device, wherein the trained machine learning model is trained with at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model; and - at least one user interface configured to provide a user with information about the classified quality; Equipped with.

[0052] The test system may be configured to perform steps a)-b) and optionally step c) of the method according to the invention.

[0053] The term "communications interface" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to any special or customized meaning. The term may specifically, but not exclusively, refer to an item or element forming a boundary configured to transfer information. In particular, a communications interface may be configured to transfer information from a computing device, e.g., a computer, e.g., for the purpose of transmitting or outputting information to another device. Additionally or alternatively, a communications interface may be configured to transfer information to a computing device, e.g., a computer, e.g., for the purpose of receiving information. A communications interface may specifically provide a means for transferring or exchanging information. In particular, a communications interface may provide a data transfer connection, e.g., Bluetooth, NFC, inductive coupling, etc. By way of example, a communications interface may be or comprise at least one port comprising one or more of a network or internet port, a USB port, and a disk drive. A communications interface may be at least one web interface.

[0054] Further disclosed and proposed herein is a computer program comprising computer-executable instructions for carrying out the method according to the present invention in one or more of the embodiments encompassed herein when the program is run on a computer or computer network, in particular a test system. In particular, the computer program may be stored on a computer-readable data carrier and / or a computer-readable storage medium.

[0055] As used herein, the terms "computer-readable data carrier" and "computer-readable storage medium" may specifically refer to non-transitory data storage means such as a hardware storage medium having computer-executable instructions stored thereon. A computer-readable data carrier or storage medium may specifically be or include a storage medium such as a random access memory (RAM) and / or a read-only memory (ROM).

[0056] Thus, in particular, one, two or more or even all of the above method steps a)-b) and optionally step c) may be carried out by using a computer or a computer network, preferably by using a computer program.

[0057] It is further disclosed and proposed herein that a program, when executed on a computer or computer network, in particular a test system, generates program code steps for carrying out the method according to the invention in one or more of the embodiments encompassed herein. In particular, the program code means may be stored on a computer readable data carrier and / or a computer readable storage medium.

[0058] Further disclosed and proposed herein is a data carrier having stored thereon a data structure which, after being loaded into a computer or computer network, such as into the working memory or main memory of the computer or computer network, is capable of carrying out the methods according to one or more of the embodiments disclosed herein.

[0059] Further disclosed and proposed herein is a computer program product having program code means stored on a machine-readable carrier for performing a method according to one or more of the embodiments disclosed herein when the program is executed on a computer or computer network, particularly a test system. As used herein, a computer program product refers to a program as a tradeable product. The product may generally be present in any format, such as a paper format, or on a computer-readable data carrier and / or computer-readable storage medium. In particular, the computer program product may be distributed via a data network.

[0060] Finally, disclosed and suggested herein is a modulated data signal containing instructions readable by a computer system or computer network for carrying out a method according to one or more of the embodiments disclosed herein.

[0061] With reference to computer implementations of the present invention, one or more or all of the method steps of the methods according to one or more of the embodiments disclosed herein may be performed using a computer or a computer network. Thus, in general, any of the method steps involving providing and / or manipulating data may be performed using a computer or a computer network. In general, these method steps may include any method steps, typically excluding method steps that require manual operations, such as providing a sample and / or certain aspects of performing the actual measurement.

[0062] Specifically, in this specification, a computer or computer network comprising at least one processor, the processor being configured to execute a method according to one of the embodiments described herein; - a computer-loadable data structure adapted to perform a method according to one of the embodiments described herein while the data structure is being executed on a computer; a computer program adapted to carry out a method according to one of the embodiments described herein while the program is running on a computer; a computer program comprising program means for carrying out a method according to one of the embodiments described herein while said computer program is running on a computer or a computer network; a computer program comprising program means according to the preceding embodiments, the program means being stored on a computer-readable storage medium; and a storage medium, the storage medium having a data structure stored thereon, the data structure being adapted to perform a method according to one of the embodiments described herein after being loaded into a main memory and / or a working memory of a computer or a computer network; Storage media and a computer program product having program code means, which may be stored on a storage medium or is stored on a storage medium, for performing a method according to one of the embodiments described herein when the program code means is executed on a computer or a computer network; is further disclosed.

[0063] In summary, without excluding further embodiments, the following embodiments may be envisaged:

[0064] Embodiment 1. A computer-implemented method for automated quality checking of chromatographic and / or mass spectral data, comprising: a) providing processed chromatographic and / or mass spectral data obtained by at least one mass spectrometer; b) classifying the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data, wherein the trained machine learning model uses at least one regression model, and the trained machine learning model is trained on at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model. A method comprising:

[0065] Embodiment 2. The method of embodiment 1, wherein the analyte is at least one target substance selected from the group consisting of vitamin D, drugs of abuse, therapeutic agents, hormones, and metabolites to be quantified from the sample.

[0066] Embodiment 3. The method of any one of embodiments 1 or 2, wherein the regression model is at least one regression model selected from the group consisting of Random Forest, Gradient Boosting Forest, Partial Least Squares, Lasso Regression, Logistic Regression, and Bayesian Regression.

[0067] Embodiment 4. The method of any one of embodiments 1 to 3, wherein the regression model is at least one regression model selected from the group of gradient boosting forest or random forest.

[0068] Embodiment 5. The method of any one of embodiments 1 to 4, wherein the regression model is a Gradient Boosting Forest.

[0069] Embodiment 6. The method of any one of embodiments 1 to 5, wherein the regression model is a random forest.

[0070] Embodiment 7. The method of any one of embodiments 1 to 6, which is performed fully automatically.

[0071] Embodiment 8. The classified quality is classified into acceptable chromatographic data and / or mass spectral data and unacceptable chromatographic data and / or mass spectral data. 8. The method of any one of embodiments 1 to 7, comprising a step of assigning a flag to the chromatographic data and / or mass spectral data as acceptable or unacceptable based on the classified quality, which is used to distinguish between chromatographic data and mass spectral data.

[0072] Embodiment 9. The method of any one of embodiments 1 to 8, comprising providing at least one piece of information to a user via at least one user interface in response to flags in the chromatographic data and / or mass spectral data.

[0073] Embodiment 10. The method of any one of embodiments 1 to 9, wherein the machine learning model uses a feature set, the feature set including at least one feature selected from the group consisting of peak area, peak background, relative background, ion ratio, Q4 ratio, retention time ratio, peak asymmetry, asymmetry ratio, peak width, peak width ratio, area of ​​integrated residual, confidence interval of peak area, mass shift, full width at half maximum, signal to noise ratio, median single cycle ratio, median single cycle ion ratio, peak height, peak fit mean squared error, fit intensity correlation, earthmover distance, and the deviation of any of the above features when derived from processed data and raw data.

[0074]

[0033] Embodiment 11 c) At least one training step, wherein the training step comprises training a machine learning model based on a training dataset. 11. The method of any one of embodiments 1 to 10, comprising:

[0075]

[0023] Embodiment 12. The method of embodiment 11, wherein the training step comprises training machine learning models for different analytes.

[0076] Embodiment 13. The method of embodiment 12, wherein the training step is performed during assay development for a plurality of different assays, and the trained machine learning models for the different assays are stored in at least one databank.

[0077] Embodiment 14: The method of any one of embodiments 12 or 13, comprising at least one selection step performed before step b), wherein in the selection step, one trained machine learning model is selected from trained machine learning models trained on the analyte used to obtain the provided chromatographic data and / or mass spectral data.

[0078] Embodiment 15. The method of any one of embodiments 1 to 14, wherein the training data set is generated by manually classifying historical chromatographic and / or mass spectral data and / or semi-synthetic chromatographic and / or mass spectral data into two categories.

[0079] Embodiment 16. The method of any one of embodiments 1 to 15, wherein the semi-synthetic chromatographic and / or mass spectral data comprises modified historical chromatographic and / or mass spectral data, and the historical chromatographic and / or mass spectral data is modified by one or more of introducing at least one interference, introducing background, introducing at least one shift in retention time, modifying peak widths, replacing internal standard signals with chromatograms from double blank samples.

[0080] Embodiment 17: A test system configured to perform the method according to any one of embodiments 1 to 16, comprising: - at least one communication interface configured to receive processed chromatographic and / or mass spectral data obtained by at least one mass spectrometer; - at least one processing device configured to classify the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data, wherein the trained machine learning model uses at least one regression model, and the trained machine learning model is trained on at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model; and - at least one user interface configured to provide a user with information about the classified quality; A test system comprising:

[0081] Embodiment 18: A test system as described in embodiment 17, configured to perform steps a) to b) and optionally step c) of the method as described in any one of embodiments 1 to 16.

[0082] Embodiment 19: A computer program comprising instructions which, when executed by a test system as described in embodiment 17 or 18, cause the test system to perform steps a) to b) and optionally step c) of the method as described in any one of embodiments 1 to 16.

[0083] Embodiment 20: A computer-readable storage medium containing instructions that, when executed by a test system described in embodiment 17 or 18, cause the test system to perform steps a) to b) and optionally step c) of the method described in any one of embodiments 1 to 16. [Brief explanation of the drawings]

[0084] Further optional features and embodiments are disclosed in more detail in the following description of the embodiments, preferably in conjunction with the dependent claims. In this description, each optional feature may be realized alone as well as in any feasible combination, as understood by a person skilled in the art. The scope of the present invention is not limited by the preferred embodiments. The embodiments are schematically illustrated in the figures. In the embodiments, the same reference numerals in these figures refer to the same or functionally equivalent elements. [Figure 1] FIG. 1 illustrates one embodiment of a method for automated quality checking of chromatographic and / or mass spectral data according to the present invention. [Figure 2] 1 is a schematic diagram of the development and deployment of a trained machine learning model. [Figure 3a] FIG. 1 illustrates a simulation scenario. [Figure 3b] FIG. 1 illustrates a simulation scenario. [Figure 3c] FIG. 1 illustrates a simulation scenario. [Figure 3d] FIG. 1 illustrates a simulation scenario. [Figure 3e] FIG. 1 illustrates a simulation scenario. [Figure 4] FIG. 1 illustrates the definition of regression model results in terms of percent deviation from the original area ratio. [Figure 5] 1 illustrates an embodiment of a mass spectrometer including a test system according to the present invention. [Figure 6] FIG. 1 illustrates an example of model optimization. DETAILED DESCRIPTION OF THE INVENTION

[0085] FIG. 1 is a flow diagram of a computer-implemented method for automated quality checking of chromatographic and / or mass spectral data. The method comprises: a) providing processed chromatographic and / or mass spectral data obtained by at least one mass analyzer 112 (denoted by reference numeral 110); b) classifying the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data (denoted by reference numeral 114), wherein the trained machine learning model uses at least one regression model, the trained machine learning model is trained on at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model. Includes:

[0086] The mass spectral data may be data, in particular at least one mass spectrum, obtained by using at least one mass analyzer 112. The chromatographic data may be at least one chromatogram.

[0087] The quality check may be a process that distinguishes between reliable and unreliable automated peak integration. The quality check may include determining whether the raw data reduction process is complete, whether the data quality was suitable for automated peak integration, and whether the calculated nominal signals and readouts are reliable. The quality may be a measure of the reliability of the automated peak integration performed on the data provided by the MS instrument and / or LC instrument 112. The classified quality may be used to distinguish between acceptable and unacceptable chromatographic and / or mass spectral data. Specifically, the quality may be classified as good (acceptable) for reliable automated peak integration and as poor (unacceptable) for unreliable automated peak integration. The quality classification may include distinguishing between reliable and unreliable automated peak integration. The quality may depend on several factors, such as noise level, background, interference, retention time shift, peak width, and the presence or absence of an internal standard signal.

[0088] The processed chromatographic data and / or mass spectral data may be chromatographic data and / or mass spectral data that has been subjected to automated peak integration. Regarding automated peak integration, see WO 2021 / 023865 A1, the entire contents of which are incorporated by reference.

[0089] Providing in step a) 110 may include determining and / or generating and / or making available processed chromatographic data and / or mass spectral data, in particular by performing at least one measurement using a mass analyzer and then processing the data. Providing the processed chromatographic data and / or mass spectral data may include retrieving the processed chromatographic data and / or mass spectral data obtained from the mass analyzer 112 upon specific receipt and / or performing at least one measurement and processing using the mass analyzer 112, thereby determining the processed chromatographic data and / or mass spectral data.

[0090] The classification in step b) 114) may include classifying the chromatographic and / or mass spectral data into at least two categories, such as good or reliable for reliable automated peak integration and bad or unreliable for unreliable automated peak integration. The classification is performed by applying at least one trained machine learning model. Thus, according to the present invention, at least one machine learning model is used to predict peak integration failures, allowing for a fully automated decision regarding the release of results. The proposed method thus allows for eliminating the need for manual inspection of the data.

[0091] The trained machine learning model uses at least one regression model 116. The regression model 116 may be a predictive model configured to analyze the relationship between a target variable and independent variables in a dataset. The target variable for chromatographic data may be a continuous deviation from an expected result value. For mass spectrometric data, the target variable may be dichotomous information regarding whether the result is valid or not. The regression model 116 may be any of the following: Random forests, as described, for example, in Breiman L., Random forests, Machine Learning, 2001, 45(1):5-32; Gradient Boosting Forests, as described, for example, in Friedman, JH (2001); Greedy function approximation, as described, for example, in "A Gradient Boosting Machine", The Annals of Statistics, 29(5):1189-1232; Partial least squares, as described, for example, in Wold, H. (1985), Partial least squares, in Kotz, Samuel, Johnson, Norman L. (eds.), Encyclopedia of statistical sciences, 6. New York: Wiley, pp. 581-591; and Partial least squares, as described, for example, in Tibshirani, R. (1996), Regression Shrinkage and Selection via the lasso, Journal of the Royal Statistical Society. Series. B(methodological). Wiley. 58(1):267-88, Lasso regression as described, for example, in Hosmer, D., Lemeshow, S.: Applied logistic regression, Wiley, New York 2000, or Bayesian regression as described, for example, in Box, G.E.P., Tiao, G.C. (1973), Bayesian Inference in Statistical Analysis. Wiley.For example, the regression model 116 is selected from a gradient boosting forest or a random forest. For example, the regression model 116 is a gradient boosting forest. For example, the regression model 116 is a random forest.

[0092] The trained machine learning model is an analyte-specific trained machine learning model. For example, the analyte is at least one target substance selected from the group consisting of vitamin D, drugs of abuse, therapeutic drugs, hormones, and metabolites to be quantified from a sample. The sample may be any test sample, such as a biological sample and / or an internal standard sample. The sample may contain one or more analytes. For example, the test sample may be selected from the group consisting of physiological fluids including blood, serum, plasma, saliva, ocular lens fluid, cerebrospinal fluid, sweat, urine, milk, ascites, mucus, synovial fluid, peritoneal fluid, amniotic fluid, tissue, cells, etc. The sample may be used directly as obtained from its respective source or may be subjected to pretreatment and / or sample preparation workflows. For example, the sample may be pretreated by adding an internal standard and / or diluting it with another solution and / or mixing it with a reagent, etc. For example, the analyte may generally be vitamin D, drugs of abuse, therapeutic drugs, hormones, and metabolites. The internal standard The quasi-sample may be a sample containing at least one internal standard with a known concentration. For further details regarding samples, see, for example, EP 3425369 A1, the entire disclosure of which is incorporated herein by reference. Other analytes are also possible.

[0093] The machine learning model may use a feature set 118. The feature set 118, which may be beneficial to data and peak integration quality, may include standard MS quality parameters such as peak asymmetry or ion ratio, ratios of parameters between different transitions, e.g., the retention time ratio between the analyte quantifier and the internal standard quantifier, features for assessing the quality of the peak fit, e.g., residual ratio or peak fit uncertainty, and additional design features describing noise, background, and peak shape. The feature set 118 may include at least one feature selected from the group consisting of peak area, peak background, relative background, ion ratio, Q4 ratio, retention time ratio, peak asymmetry, asymmetry ratio, peak width, peak width ratio, area of ​​integrated residual, confidence interval of peak area, mass shift, full width at half maximum, signal-to-noise ratio, median single-cycle ratio, median single-cycle ion ratio, peak height, peak fit mean square error, fit intensity correlation, earthmover distance, and deviation of any of the mentioned features when derived from processed data, i.e., integrated peaks and raw data, e.g., the difference in retention time between the fitted peak and the raw signal. Peak background may refer to the intensity of the estimated background in the peak interval. Relative background may refer to the ratio of peak background to peak height. Ion ratio may refer to the area of ​​the analyte or internal standard (ISTD) quantifier relative to the area of ​​the analyte or the area of ​​the ISTD quantifier. The Q4 ratio may be given by Q4 = (area of ​​analyte quantifier / area of ​​analyte quantifier) / (area of ​​ISTD quantifier / area of ​​ISTD quantifier).The retention time ratio may refer to one or more of RT_analyte_qualifier / RT_analyte_quantifier, RT_ISTD_qualifier / RT_ISTD_quantifier or RT_analyte_quantifier / RT_ISTD_quantifier, where RT_analyte_qualifier is the retention time of the analyte qualifier, RT_analyte_quantifier is the retention time of the analyte quantifier, RT_ISTD_qualifier is the retention time of the ISTD qualifier, and RT_ISTD_quantifier is the retention time of the ISTD quantifier. Peak asymmetry may be defined according to USP40. The asymmetry ratio may refer to one or more of asymmetry_analyte_qualifier / asymmetry_analyte_quantifier, asymmetry_ISTD_qualifier / asymmetry_ISTD_quantifier, or asymmetry_analyte_quantifier / asymmetry_ISTD_quantifier, where asymmetry_analyte_qualifier is the asymmetry of the peak of the analyte qualifier, asymmetry_ISTD_qualifier is the asymmetry of the peak of the ISTD qualifier, and asymmetry_ISTD_qualifier is the asymmetry of the peak of the ISTD quantifier. The peak width ratio may refer to one or more of width_analyte_qualifier / width_analyte_quantifier, width_ISTD_qualifier / width_ISTD_quantifier, or width_analyte_quantifier / width_ISTD_quantifier, where width_analyte_qualifier is the peak width of the analyte qualifier, width_analyte_quantifier is the peak width of the analyte quantifier, width_ISTD_qualifier is the peak width of the ISTD qualifier, and width_ISTD_quantifier is the peak width of the ISTD quantifier. The signal-to-noise ratio may be defined according to USP40.The median single cycle ratio may refer to the median ratio of the intensity of the analyte quantifier to the intensity of the ISTD quantifier. The median single cycle ion ratio may refer to the ratio of the intensity of the analyte quantifier to the intensity of the analyte qualifier. , or may refer to one or more median values ​​of the ratio of the intensity of the ISTD quantifier to the intensity of the ISTD qualifier. The peak fit mean square error may be given by the mean [(smoothed intensity / area of ​​fitted intensity / area)²]. The fit intensity correlation may refer to one or more of cor(smoothed intensity, fitted intensity) or cor(preprocessed intensity, fitted intensity). Regarding Earthmover distance, see, for example, https: / / en.wikipedia.org / wiki / Earth_mover%27s_distance. A rich set of features can be derived from the chromatographic data and / or mass spectral data and used to build a regression model. Training the model may include determining feature rankings. Training the model may include selecting features.

[0094] FIG. 2 shows a schematic diagram of the development and deployment of a trained machine learning model, in this case, regression model 116. Features from feature set 118 may be combined in regression model 116 to predict area ratio deviations as the equivalent of peak integration failures. The trained regression model may then be deployed to predict the quality state of new measurements, as performed in step b) 114. Shown from left to right in FIG. 2 are feature set 118, an exemplary regression model 116, and the application of the trained regression model 116 to exemplary processed chromatographic data and / or mass spectral data. In the top right plot, the processed chromatographic data and / or mass spectral data is classified as good in step b), and in the bottom right plot, it is classified as bad.

[0095] Regression models 116, such as random forests and gradient boosting, have been found to perform well at reasonable model complexity in terms of evaluation time and required disk space. Model parameters such as the type of algorithm, number of features, number and size of trees may be adjusted by resampling techniques.

[0096] In the case of random forests, it has been found that increasing the number of features improves the performance of random forests. In the case of gradient boosting forests, it has been found that decreasing the number of features improves the performance of gradient boosting forests. Feature selection may be performed to select top features that are "stable" across many data splits and / or models. The method may also include feature engineering, including evaluation of newly created features. For example, in the case of gradient boosting forests, 50 features may be used with a minimum leaf size of 50 and 400 trees.

[0097] The method may comprise at least one training step, step c) 120. The training step may comprise training a machine learning model based on a training dataset.

[0098] Training may include the process of building a trained machine learning model, particularly determining the model's parameters, particularly weights. Training may include determining and / or updating the model's parameters. Training may be performed on historical chromatographic and / or mass spectral data and / or semi-synthetic chromatographic and / or mass spectral data. Training may also include retraining the trained model after acquiring additional chromatographic and / or mass spectral data, for example, during operation of an MS and / or LC-MS instrument.

[0099] The training step 120 may include training machine learning models for different analytes. The training step 120 may be performed during assay development for a number of different assays, where the trained machine learning models for the different assays are trained for at least one data bank. The databank may include a data processing configuration file, enabling automatic flagging of peak integration results on the instrument. The method may include at least one selection step, performed before step b), e.g., as part of step c), in which one trained machine learning model is selected from trained machine learning models trained on the analyte used to obtain the provided chromatographic data and / or mass spectral data.

[0100] The trained machine learning model may be suitable for different analytes with similar chromatography. The training step may include training the machine learning model for different chromatography types. For different chromatography types, separate models may be used, for example, for standard chromatography, in which peak fitting may be applied, non-standard chromatography, in which boundary detection must be applied, and when an internal standard having exactly the same retention time as the analyte and a retention time offset exists between the analyte and the ISTD is not available.

[0101] The historical chromatographic and / or mass spectral data may include measurements obtained using at least one mass spectrometer. The historical data may be actual data. The historical chromatographic and / or mass spectral data may include data from different instruments measuring several analytes and having different scenarios. An example of a historical training dataset may include approximately 500 chromatographic measurements involving five different analytes measured on two instruments from one system and three instruments from another system during an 11-week period.

[0102] The training dataset includes semi-synthetic chromatographic and / or mass spectral data, also referred to as a semi-synthetic dataset. The semi-synthetic chromatographic and / or mass spectral data may be simulated based on historical chromatographic and / or mass spectral data. The semi-synthetic chromatographic and / or mass spectral data may be generated by applying and / or simulating defined disturbances to actual measured chromatographic and / or mass spectral data. The semi-synthetic chromatographic and / or mass spectral data may include modified historical chromatographic and / or mass spectral data. The historical chromatographic and / or mass spectral data may be modified by one or more of the following: introducing at least one interference; introducing background; introducing at least one shift in retention time; modifying peak width; or replacing the internal standard signal with a chromatogram from a double blank sample. The semi-synthetic simulation approach combines the benefits of knowing the truth in simulation studies with providing a dataset with real-world characteristics. Using simulated datasets for model training has several advantages over real data, including the ability to objectively define the true state of the measurement, the ability to investigate rare cases and "gray zones," and scalability in terms of sample size. A semi-synthetic approach is employed, whereby the actual measurements are modified in a controlled way in order to resemble the real data as closely as possible.

[0103] Figure 3, a–e, shows different simulation scenarios. The top row shows the actual data, and the bottom row shows the actual data plus introduced disturbances. In Figure 3a, at least one interference was introduced by varying the transition, position, resolution, and relative height. In Figure 3b, a retention time shift was introduced by varying the shift. In Figure 3c, a background was introduced by changing the height and curvature. In Figure 3d, the peak width was varied by changing the scale factor. In Figure 3e, a missing ISTD signal was simulated.

[0104] A semi-synthetic dataset may be generated as follows: A (manually curated) actual chromatogram with clear peaks and reliable integration results is selected and then modified to resemble a challenging situation for peak integration. The generation of the semi-synthetic dataset may include considering one or more of the following conditions: interference, background, retention time shift, peak width, and missing internal standard signal. For example, to account for interference, the fitted intensity of the actual internal standard peak is added to the raw intensity of the neighboring analyte peak. The distance between peaks allows for various resolutions to be explored. The height of the artificial interference peak may be expanded or contracted to simulate different relative peak heights between the peak of interest and the interference. For example, to account for background, a step function is first generated to simulate a varying background signal, and the step height is drawn from a uniform distribution. The maximum step height can control the magnitude of the simulated background. A background fit is then applied to the step function, and the resulting curve is added to the actual chromatogram intensities. A curvature parameter in the background fit allows for manipulation of the curvature of the artificial background. For example, retention time variations can be easily simulated by shifting the actual signal along the time scale to account for retention time shifts. For example, to account for peak widths, the peak fit is rescaled by changing each parameter of the fitting function. To maintain the area under the peak, the intensity is rescaled. Rescaled noise from the original data is then added to the new peak fit. For example, to account for missing internal standard signals, the chromatogram of the internal standard is replaced with the chromatogram of a double blank sample.

[0105] The simulated data, i.e., chromatographic data and / or mass spectral data, may have a much higher proportion of bad cases and a much higher proportion of borderline cases than the real data, i.e., historical synthetic chromatographic data and / or mass spectral data. Including a portion of the real data for training can improve model performance. Another portion of the real data may be used to test the trained model. The real data may be a manually labeled, true dataset.

[0106] The method may include at least one testing step, which includes testing the trained model. The testing step may include testing the trained model against at least one test dataset. The testing step may include obtaining performance characteristics of the trained model, such as accuracy, false positive rate, and false negative rate. To evaluate predictive performance, the model may be tested using simulated data and / or against real data, particularly a manually labeled real dataset. The test dataset may include simulated data and / or real data.

[0107] An exemplary machine learning model for an analyte with a typical peak shape (e.g., testosterone) was trained on a semi-synthetic dataset. The machine learning model was trained on 241 manually labeled actual measurements obtained from 10 sample runs on different instruments. 121 were manually labeled as bad and 120 were manually labeled as good. A quality check of the peak integration using the trained machine learning model correctly classified all 120 "good" measurements. Five of the 121 "bad" measurements were classified as "good" by the trained machine learning model. The accuracy was determined to be 0.9793, with a false positive rate of 0.0000 and a false negative rate of 0.0413.

[0108] A measure of the extent to which the data are affected by the introduced "disturbing factors" is created. The area ratio result may be the percent deviation of the area ratio result calculated for the semi-synthetic data from the area ratio of the original real data set. The area ratio deviation represents a continuous outcome of the regression model. A gold standard for error handling can then be defined by flagging measurements, for example, considering a threshold of area ratio deviation greater than 10%. The binary flag serves as the true state when evaluating prediction performance in terms of accuracy and false positive / false negative rate. Figure 4 illustrates the definition of the regression model result in terms of percent deviation from the original area ratio. The top panel of Figure 4 shows five integrated peaks, labeled A-E. The bottom plot of Figure 4 shows the percent area ratio deviation as a continuous outcome for prediction, for A-E. Additionally, a threshold of >10% is indicated.

[0109] The trained machine learning model may then be deployed to predict the quality state of new measurements, as performed in step b). The trained machine learning models for different analytes and / or different chromatography types may be transferred to a data processing configuration file. The data processing configuration file may be stored in at least one data store of the mass analyzer 112, which may enable automatic flagging of peak integration results of the mass analyzer 112.

[0110] 5 illustrates one embodiment of a mass spectrometer 112 including a test system 122 in accordance with the present invention. at least one communication interface 124 configured to receive processed chromatographic and / or mass spectral data obtained by at least one mass spectrometer 112; - at least one processing device 126 configured to classify the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data, wherein the trained machine learning model uses at least one regression model 116, the trained machine learning model is trained on at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model; and - at least one user interface 128 configured to provide a user with information about the classified quality; Equipped with.

[0111] Figure 6 shows an example of model optimization. The table contains area under the curve (AUC) values ​​derived by data resampling for different model settings: Gradient Boosting Forest (GBR) in the left block, Random Forest Regression (RFR) in the right block, the number of estimators ("num_est" = number of trees) in the columns, the number of dimensions ("d" = number of features) and the minimum leaf size ("msl" = size of the trees) in the rows. Darker colors and larger values ​​indicate better model performance. [Explanation of symbols]

[0112] 110 Step a) 112 Mass spectrometer 114 Step b) 116 Regression Models 118 feature sets 120 Step c) 122 Test System 124 Communication Interface 126 Processing equipment 128 User Interface

Claims

1. 1. A computer-implemented method for automated quality checking of chromatographic and / or mass spectral data, comprising: a) providing (110) processed chromatographic and / or mass spectral data obtained by at least one mass spectrometer (112); b) classifying the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data (114), wherein the trained machine learning model uses at least one regression model (116), the trained machine learning model is trained on at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model; A method comprising:

2. 10. The method of claim 1, wherein the analyte is at least one target substance selected from the group consisting of vitamin D, drugs of abuse, therapeutic drugs, hormones, and metabolites to be quantified from a sample.

3. 3. The method of claim 1, wherein the regression model (116) is at least one regression model selected from the group consisting of: random forest, gradient boosting forest, partial least squares, Lasso regression, logistic regression, and Bayesian regression.

4. 4. The method according to any one of claims 1 to 3, which is carried out fully automatically.

5. 5. The method of claim 1, wherein the classified quality is used to distinguish between acceptable and unacceptable chromatographic and / or mass spectral data, the method comprising flagging the chromatographic and / or mass spectral data as acceptable or unacceptable based on the classified quality.

6. 6. The method of claim 5, comprising providing at least one piece of information to a user via at least one user interface (128) in response to the flags in the chromatographic data and / or mass spectral data.

7. 7. The method of any one of claims 1 to 6, wherein the machine learning model uses a feature set (118), the feature set (118) comprising at least one feature selected from the group consisting of peak area, peak background, relative background, ion ratio, Q4 ratio, retention time ratio, peak asymmetry, asymmetry ratio, peak width, peak width ratio, area of ​​integrated residual, confidence interval of peak area, mass shift, full width at half maximum, signal to noise ratio, median single cycle ratio, median single cycle ion ratio, peak height, peak fit mean square error, fit intensity correlation, earthmover distance, and the deviation of any of the foregoing features when derived from processed data and raw data.

8. c) at least one training step (120), said training step (120) comprising training said machine learning model based on said training dataset; One training step (120) 8. The method of claim 1, comprising:

9. The method of claim 8 , wherein the training step (120) comprises training machine learning models for different analytes.

10. 10. The method of claim 1, wherein the training data set is generated by manually classifying the historical chromatographic and / or mass spectral data and / or the semi-synthetic chromatographic and / or mass spectral data into two categories.

11. 11. The method of any one of claims 1 to 10, wherein the semi-synthetic chromatographic and / or mass spectral data comprises modified historical chromatographic and / or mass spectral data, wherein the historical chromatographic and / or mass spectral data is modified by one or more of introducing at least one interference, introducing background, introducing at least one shift in retention time, modifying peak width, replacing an internal standard signal with a chromatogram from a double blank sample.

12. A test system (122) configured to perform the method of any one of claims 1 to 11, comprising: at least one communication interface (124) configured to receive processed chromatographic and / or mass spectral data obtained by at least one mass spectrometer (112); at least one processing device (126) configured to classify the quality of the chromatographic data and / or mass spectral data by applying at least one trained machine learning model to the chromatographic data and / or mass spectral data, wherein the trained machine learning model uses at least one regression model (116), the trained machine learning model is trained on at least one training dataset comprising historical chromatographic data and / or mass spectral data and / or semi-synthetic chromatographic data and / or mass spectral data, and the trained machine learning model is an analyte-specific trained machine learning model; and at least one user interface (128) adapted to provide a user with information about said classified quality; A test system (122) comprising:

13. 13. A test system (122) according to claim 12, configured to carry out steps a) to b) and optionally step c) of the method according to any one of claims 1 to 11.

14. 14. A computer program comprising instructions which, when executed by a test system (122) according to claim 12 or 13, cause the test system to perform steps a) to b) and optionally step c) of the method according to claim 1.

15. 14. A computer-readable storage medium containing instructions that, when executed by a test system (122) according to claim 12 or 13, cause the test system to perform steps a) to b) and optionally step c) of the method according to claim 1.