A method for predicting one or more sample metric values ​​using machine learning.

By extracting representative features from LC-MS data and applying machine learning, the method addresses the inefficiencies of current proteomic data processing, enabling rapid and accurate prediction of sample metrics for QC runs.

DE112024001902T5Pending Publication Date: 2026-02-19THERMO FISHER SCI BREMEN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112024001902
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-28
Filing Date
2024-04-16
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Current methods for processing proteomic data in liquid chromatography-mass spectrometry (LC-MS) are time-consuming and computationally intensive, particularly during quality control (QC) runs, requiring significant resources and time to determine sample metrics like the number of protein groups.

Method used

A method involving data reduction of raw LC-MS data to extract representative features, followed by using machine learning to predict sample metrics such as the number of protein groups, peptide groups, or peptide spectrum matches, reducing the need for extensive database searches.

Benefits of technology

This approach significantly reduces the time required to estimate sample metrics, allowing for faster validation of instrument performance and quality control, while preserving the accuracy and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for predicting one or more sample metric values ​​using machine learning is provided. The method involves determining a variety of properties from raw data obtained by analyzing a proteomics sample in a liquid chromatography-mass spectrometer (LC-MS). The method further includes predicting one or more sample metric values ​​by processing these properties using one or more trained machine learning models.
Need to check novelty before this filing date? Find Prior Art

Description

AREA OF INVENTION

[0001] The invention relates to a novel method for predicting sample metrics through a data reduction step combined with machine learning. In particular, the invention relates to reducing processing time by using only a representative subset of the information contained in the raw mass spectrometry data files. The invention can be used during quality control (QC) to assess the performance of liquid chromatography-mass spectrometry (LC-MS) instruments. The invention can also be used to estimate the quality of the raw data generated by liquid chromatography-mass spectrometry (LC-MS) instruments. STATE OF THE ART

[0002] In mass spectrometry, a sample is ionized and the resulting ions are categorized according to their mass-to-charge ratio (m / z). The output of a typical mass spectrometer is a spectrum showing the m / z distribution detected by the system, which must be carefully interpreted to provide information about the sample.

[0003] Fig.Figure 1 shows a schematic diagram of a mass spectrometer that can be used to carry out the invention. Such a mass spectrometer is described in detail in WO2012 / 160001. The diagram shows the mass spectrometer 2, in which ions are generated from a sample in an ion source (not shown), which can be a conventional ion source such as an electrospray. Ions can be generated in the ion source as a continuous current, as in the electrospray, or in pulsed form, as in a matrix-assisted laser desorption / ionization source, or "MALDI" source. The sample that is ionized in the ion source can originate from an attached instrument such as a liquid chromatograph (not shown). The ions pass through a heated capillary 4, are transferred through a pure RF S-lens 6, and pass through the S-lens exit lens 8.The ions in the ion beam are next transferred through an injection flat pole 10 and a curved flat pole 12; these are purely RF devices for transferring the ions, with the RF amplitude being adjusted depending on the mass. The ions then pass through a pair of lenses and enter a mass-resolving quadrupole 18.

[0004] The differential RF and DC voltages of the quadrupole 18 are controlled such that either all ions are transferred (RF-only mode) or ions of a specific m / z range are selected for transfer by applying RF and DC voltages according to the Mathieu stability diagram. It is understood that in other embodiments, a pure RF quadrupole or multipole can be used as ion guide instead of the mass-resolving quadrupole 18; however, the spectrometer would then lack the mass selection capability prior to analysis. In still other embodiments, an alternative mass resolution device, such as a linear ion trap, a magnetic sector, or a time-of-flight analyzer, can be used instead of the quadrupole 18. Such a mass resolution device could be used for mass selection and / or ion fragmentation.Returning to the arrangement shown, the ion beam, transmitted through the quadrupole 18, exits the quadrupole through a quadrupole exit lens 20 and is switched on and off by a split lens 22. The ions are then transferred through a transfer multipole 24 (HF only; the HF amplitude can be adjusted depending on the mass) and collected in a curved linear ion trap (C-trap) 26. The ions are captured radially in the C-trap by applying an HF voltage to the curved rods of the trap in a known manner. The C-trap is extended in an axial direction (thus defining a trap axis) in which the ions enter the trap. The voltage at the exit lens 28 of the C-trap can be adjusted so that ions cannot pass through it and are thus stored in the C-trap 26. Similarly, after the desired ion filling time (or number of ion pulses, e.g.,Once the ion beam reaches the C-trap (at MALDI), the voltage at the entrance lens 30 of the C-trap is adjusted so that no ions can escape from the trap and no more ions are injected into the C-trap. More precise gating of the incident ion beam is enabled by the split lens 22.

[0005] Ions stored in the C-trap 26 can be ejected orthogonally to the axis of the trap (orthogonal ejection) by pulsed direct current into the C-trap, so that the ions are injected in this case via the Z-lens 32 and the deflector 33 into a mass analyzer 34, which in this case is an electrostatic orbital trap, and more precisely an orbitrap. (TM)This is a Thermo Fisher Scientific FT mass analyzer. The orbital trap 34 comprises an inner electrode 40 extending along the orbital trap axis and a split pair of outer electrodes 42, 44 surrounding the inner electrode 40 and defining a capture volume between them. Ions are trapped in this volume and oscillate by orbiting the inner electrode 40, to which a capture voltage is applied, while oscillating back and forth along the trap axis. The pair of outer electrodes 42, 44 acts as detection electrodes to detect an image current induced by the oscillation of the ions in the capture volume, thus providing a detected signal. The outer electrodes 42, 44 therefore constitute a first detector of the system.The outer electrodes 42, 44 typically function as a differential pair of detection electrodes and are coupled to respective inputs of a differential amplifier (not shown), which in turn is part of a digital data acquisition system (not shown) for receiving the detected signal. The detected signal can be processed using a Fourier transform to obtain a mass spectrum.

[0006] The mass spectrometer 2 further comprises a collision or reaction cell 50 downstream of the C-trap 26. Ions collected in the C-trap 26 can either be ejected orthogonally as a pulse to the mass analyzer 34 without entering the collision or reaction cell 52, or the ions can be passed axially to the collision or reaction cell for processing before the processed ions are returned to the C-trap for subsequent orthogonal ejection to the mass analyzer. In this case, the exit lens 28 of the C-trap is adjusted such that ions can enter the collision or reaction cell 50, and ions can be injected into the collision or reaction cell by means of a corresponding voltage gradient between the C-trap and the collision or reaction cell (e.g., the collision or reaction cell can be set to a negative potential for positive ions).The collision energy can be controlled by this voltage gradient. The collision or reaction cell 50 includes a multipole 52 for receiving the ions. The collision or reaction cell 50 can, for example, be pressurized with a collision gas to enable fragmentation (collision-induced dissociation) of the ions contained therein, or it can contain a source of reactive ions for electron transfer dissociation (ETD) of the ions contained therein. Applying a suitable voltage to an exit lens 54 of the collision cell prevents the ions from axially exiting the collision or reaction cell 50.The C-trap exit lens 28 at the other end of the collision or reaction cell 50 also functions as an entry lens for the collision or reaction cell 50 and can be adjusted to prevent ions from exiting while being processed in the collision or reaction cell. In other embodiments, the collision or reaction cell 50 may have its own separate entry lens. After processing in the collision or reaction cell 50, the potential of the cell 50 can be shifted to expel ions back into the C-trap for storage (with the C-trap exit lens 28 adjusted to allow the ions to return to the C-trap); for example, the voltage shift of the cell 50 can be increased to expel positive ions back to the C-trap. The ions thus stored in the C-trap can then be injected into the mass analyzer 34 as previously described.

[0007] The mass spectrometer 2 optionally further comprises an electrometer 60, which is located downstream of the collision or reaction cell 50 and can be reached by the ion beam through an aperture 62 in the exit lens 54 of the collision cell. The electrometer 60 can be either a collector plate or a Faraday cup and is connected to a charge-sensitive amplifier with high gain. However, it is understood that in other arrangements the electrometer 60 can be a different type of charge-measuring device. Preferably, the electrometer is a differential type, which reduces noise pickup from other nearby electrical sources. A first input of the electrometer is arranged to receive current or charge from the ion source, while another input is arranged to have a similar capacitance, dimensions, and orientation to the first input but does not receive any ion current or charge.The electrometer 60 thus represents an optional second detector of the system, which is independent of the first detector, namely the image current detection electrodes 42, 44 of the mass analyzer 34. In some arrangements, the collision or reaction cell 50 may not be present, in which case the electrometer 60 is preferably located downstream of the C-trap behind the C-trap exit lens 28.

[0008] It is understood that the path of the ion beam through the spectrometer and in the mass analyzer takes place under suitable evacuation conditions, as are known in the art, with different vacuum levels being suitable for different parts of the spectrometer.

[0009] It is understood that any other mass spectrometer may be equally suitable for use in connection with this invention. For example, the Orbitrap mass analyzer could be replaced by a time-of-flight analyzer, or the mass spectrometer 2 could be a triple quadrupole mass spectrometer, an ion trap mass spectrometer, or a quadrupole or ion trap quadrupole time-of-flight mass spectrometer. The mass spectrometer could comprise more than one mass analyzer; for example, the mass spectrometer could be the hybrid instrument consisting of an Orbitrap described in EP3410463. TM -Mass analyzer and time-of-flight mass analyzer.

[0010] The mass spectrometer 2 is controlled by a control unit, for example, a suitably programmed computer (not shown), which controls the operation of various components and, for example, sets the voltages to be applied to the various components and receives and processes data from various components, including the detector(s). The computer is configured to use a known algorithm to determine the settings (e.g., injection time or number of ion pulses) for injecting ions into the C-trap for analytical scans in order to achieve the desired ion concentration (i.e., the number of ions) while avoiding space charge effects and simultaneously optimizing the statistics of the data collected from the analytical scan. The algorithm may be based on previous measurements of the mass analyzer 34 or electrometer 60.

[0011] Mass spectrometry is particularly useful for the identification of proteins (or other compounds in general, hereinafter referred to as "proteins") in an organic sample, a field known as proteomics, which is an important aspect of biological and medical research.

[0012] In a method for analyzing organic samples, sometimes called "bottom-up" mass spectrometry, proteins are pre-digested into their peptide components (or, more generally, any components of a compound, referred to hereafter as "peptides"), which are then analyzed and classified in a mass spectrometer. Several databases exist that can provide the spectra expected for a given peptide. Therefore, if only a small number of proteins are present in the original sample, it is relatively easy to compare the spectrum of the peptides identified by the mass spectrometer with known spectra of various predefined proteins and identify the closest match, thus identifying the protein most likely present in the sample.

[0013] The development of large protein databases has made it possible to identify many otherwise unidentified proteins by comparing information from their analysis, such as their sequences or mass spectra, with information in or from the database. Advances in high-throughput peptide analysis techniques, such as robotic gel band excision and digestion, and matrix-assisted laser desorption / ionization mass spectrometry (MALDI), have enabled the collection of large datasets characterizing a vast number of experimental proteins. This information can be compared with information in databases of known proteins to identify such experimental proteins.

[0014] Mass spectrometry (MS) is particularly well-suited for the analysis of these peptides, especially in combination with liquid chromatography (LC). Using LC / MS, the peptides of proteolytically digested proteins are separated using LC techniques. Subsequently, a mass spectrometer analyzes the peptides according to their relative mass-to-charge ratio (m / z), generating a characteristic spectrum of peaks for the peptide, which may belong to one or more proteins. Tandem mass spectrometry (MS / MS) allows a single peptide of a protein to be selected and subjected to collision-induced dissociation (CID) or another fragmentation technique. CID generates fragment ions that can then be sorted according to their mass-to-charge ratio, resulting in a characteristic spectrum for the selected peptide.By repeatedly applying liquid chromatography-tandem mass spectrometry (LC-MS / MS), a large number of spectra can be generated, each characterizing a variety of different peptides.

[0015] A protein characterized using methods such as LC-MS / MS can be identified by comparing its experimental data, for example, the mass spectra of its peptides, with characteristic data, such as the theoretical mass spectra for peptides of previously identified ("known") proteins. By comparing the experimental data of an unknown peptide with theoretically derived properties of known peptide sequences, both the unknown peptide and the unknown protein to which the unknown peptide belongs can be identified.

[0016] Searchable protein databases are available, for example, on the website of the National Center for Biotechnology Information (NCBI) (http: / / www.ncbi.nlm.nih.gov). These include databases with information on nucleotide sequences and amino acid sequences of proteins.

[0017] To evaluate MS / MS data for peptides using a nucleotide or protein sequence database, sequences in the database representing proteins can be subdivided into sequences representing the peptides that would result from actual proteolytic digestion of the proteins. Subsequently, a theoretical spectrum can be generated for each peptide of a protein represented in the database, based on the peptide sequence. This theoretical spectrum includes mass-to-charge peaks that would be expected if the protein in the database were subjected to MS / MS and the peptide of interest were selected for characterization. Each theoretical peptide spectrum for proteins represented in the database can be compared with observed peptide spectra for an unknown protein.The similarity of the theoretical peptide spectra to the unknown peptide spectra can then be used to determine the identity of the unknown protein. The search engines SEQUEST and MASCOT implement such a routine for protein identification. See, for example, Eng JK, McCormack AL, and Yates JR 3rd, “An Approach to Correlate Tandem Mass Spectral Data of Peptides with Amino Acid Sequences in a Protein Database,” J. Am. Soc. Mass. Spectrom., 1994, 5: 976–989.

[0018] Modern software used in proteomics performs many different steps to analyze raw files. In general, these steps may include (but are not limited to): (i) reading the raw file to obtain MS and MS / MS information, (ii) extracting relevant information from the MS / MS spectrum, (iii) performing a database search as described above (SEQUST), (iv) extracting information for quantification [if applied], (v) performing false detection to distinguish between true and false positives, and (vi) reporting sample metrics such as "number of proteins," "number of protein groups," "number of peptide groups," and "number of peptide spectrum matches."

[0019] These sample metrics can be described as follows: • Number of peptide spectrum matches (PSMs): ◯ Is the number of MS / MS spectra that have been matched for a specific protein with peptide sequences. • Success rate ◯ Calculated by dividing the number of peptide spectrum matches (PSMs) by the number of MS / MS scans acquired. • Number of peptide groups ◯ Is the number of peptide sequences obtained after grouping peptide spectrum matches based on their sequence and modification. • Number of proteins: ◯ Represents the number of identified proteins. • Number of protein groups: • A reduced number of proteins. A protein group is a set of proteins that cannot be uniquely identified by individual peptides. These are grouped together as a protein group.

[0020] Processing proteomic data as described above can be a time-consuming and computationally intensive step. This is especially true for modern search engines like CHIMERYS. Resources (time, money, and / or electrical energy) should ideally be invested only in previously acquired raw files that are expected to provide useful sample metrics (such as the number of protein groups). Furthermore, new raw files should only be acquired if the LC-MS setup is performing as expected. This latter approach is often ensured by running quality control (QC) runs. QC runs are used to validate the performance of the LC-MS application system. Here, QC runs are recorded using standardized procedures that closely approximate the LC-MS settings used to collect data from real / biological samples.Therefore, a sample similar to the actual biological samples is used in these runs. Using the data obtained during the QC run, it can be determined whether the spectrometer is functioning correctly. If not, maintenance can be performed on the spectrometer to restore its performance. QC runs should be performed as close as possible to the runs analyzing the actual samples. These QC runs can therefore be performed immediately before, immediately after, or between runs of biological / actual samples. To ensure the validity of the QC runs and reduce the overall time for the LC-MS sequence, it is important to process and evaluate the data files from the QC runs as quickly as possible to minimize the waiting time of an LC-MS setup while awaiting the QC run results.

[0021] Currently, three approaches are used to reduce the waiting time of an LC-MS setup while raw files are analyzed for QC purposes or to estimate whether the quality of a biological sample is high enough to justify investing additional time in processing this data: a) Only a portion of the database search workflow is executed.

[0022] In this case, the quantification portion of the database search is usually omitted to reduce the overall search time. A disadvantage of this approach is that the database search is limited to identification (e.g., protein groups). If quantification results are also required (e.g., to further assess the quality of a QC run), the database search must be repeated to obtain results for both identification and quantification. This can tie up additional hardware resources that could be used to process other raw data. b) Only the OC files are run with a different search engine.

[0023] Using a speed-optimized search engine / algorithm can significantly reduce the required processing time. Disadvantages of this approach include (i) an increase in the number of processed result files, (ii) the need to trigger this tool via an external process, and (iii) the need to match the speed search results with the results of the original workflow. c) In another approach, the file size of the raw files generated by the mass spectrometer (MS) is used as a very approximate QC metric.

[0024] Current methods for processing proteomic data have the disadvantages described above. Therefore, a new method that is fast, lightweight, and reliable is desired. SUMMARY

[0025] The object of the invention is to provide an improved method for reducing the time required to process data from proteomics runs (such as QC runs) in order to predict commonly used sample metrics (e.g., "number of protein groups") using machine learning in order to assess their data quality. The proposed methods can be used independently of a database search approach for the rapid analysis of raw proteomics files acquired using a standardized LC-MS procedure. Based on the results, the data can be further processed using modern post-processing software employed in proteomics.

[0026] One or more sample-related metrics can be obtained using state-of-the-art methods by processing the raw files generated by the LC-MS using bioinformatics software such as Thermo Fisher Scientific's "Proteome Discoverer" software.

[0027] The one or more sample metrics for proteomics (and proteomics quality control) may include: • Number of proteins; • Number of protein groups; • Number of peptide groups; • Number of peptide spectrum matches; and / or • Success rate (=number of peptide spectrum matches / number of MS / MS scans performed).

[0028] However, running bioinformatics software packages requires significant computing resources and time. Instead of processing data files acquired with standardized LC-MS methods using computationally intensive database searches (or a speed-optimized search engine as in other state-of-the-art methods), the proposed methods extract representative data from the raw data files. Subsequently, machine learning approaches are used to predict one or more sample metrics.

[0029] In this way, the time required to estimate one or more sample metrics (see list above) can be reduced compared to state-of-the-art methods.

[0030] Therefore, a method is desired that includes data extraction from the raw file(s), data complexity reduction, and training a machine learning model to quickly predict sample metrics in order to avoid time-consuming and complex calculations in proteomics applications.

[0031] According to the present invention, a method for predicting one or more sample metric values ​​by machine learning is provided. The method comprises: Determining a variety of properties from raw data obtained by analyzing a proteomics sample in a liquid chromatography-mass spectrometer (LC-MS); and

[0032] Prediction of one or more sample metric values ​​by processing the multitude of properties determined from the raw file using one or more trained machine learning models.

[0033] The multitude of properties can be representative of the performance of the LC-MS. In other words, the properties can provide information about the quality of the raw data obtained.

[0034] Since the raw data are obtained by analyzing a proteomics sample in a liquid chromatography-mass spectrometer (LC-MS), the raw data can be referred to as LC-MS raw data.

[0035] The proteomics sample can be a standardized proteomics sample. In other words, the composition of the sample can be known. Analyzing standardized proteomics samples is useful for validating instrument performance, calibrating instrument parameters, and training data analysis methods such as machine learning models.

[0036] In one example, the proteomics sample could be a quality control (QC) sample. In this case, the one or more sample metric values ​​could include one or more QC metrics. These one or more QC metrics can be used to determine a QC result during a QC run, thereby validating the instrument's performance.

[0037] A key improvement over the prior art lies in the manipulation of raw data to reduce complexity (and data size) while preserving useful characteristic information that enables the predictive methods to deliver reliable results. This manipulation of raw data (extracting representative features from the raw data) combined with a machine learning algorithm to process these features and predict one or more sample metric values, such as the number of protein groups, is not taught anywhere in the prior art.

[0038] The sample metric to be predicted by a data reduction step and subsequent machine learning can be the number of protein groups in the proteomics sample identifiable by LC-MS. Alternatively, the sample metric can be the number of peptide groups in the proteomics sample identifiable by LC-MS. Alternatively, the sample metric can be the number of peptide spectrum matches (PSMs) in the proteomics sample identifiable by LC-MS. Alternatively, the sample metric can be the number of peptide spectrum matches divided by the number of MS / MS spectra acquired in the proteomics sample identifiable by LC-MS (defined as success rate = PSMs / number of MS / MS spectra).

[0039] The number of protein groups is a sample metric used in state-of-the-art proteomics procedures to assess the quality of raw files, particularly in QC runs. The value of this metric can be related to a combination of other parameters (including the number of peptide groups and peptide spectrum matches) and filtering options. Therefore, the value of this metric is relatively robust in a controlled environment (i.e., using the same standardized sample and the same standardized LC-MS procedure). Although the value of this sample metric is a combination of several functions, a decrease in the number of protein groups indicates that the LC-MS system (including sample and solvent) is not performing as expected. In contrast to the number of protein groups, the number of peptide spectrum matches (i.e., the number of peptide groups) is based on a combination of factors.how many MS / MS spectra could be assigned to a peptide sequence) is not based on any of the other sample metrics described above (i.e. peptide groups, proteins or protein groups), but is directly related to the quality of the MS / MS spectra.

[0040] The quality of the collected MS / MS spectra depends strongly and directly on the performance of the LC-MS system (including the solvent, sample, column, LC and MS; and the LC-MS procedure used). Since a specific peptide spectrum match (PSM) is associated with a particular MS / MS event during sample elution, this PSM (and its subsequent success rate) provides an identification-based sample metric that can be effectively visualized in an LC-MS runtime plot (i.e., number of peptide spectrum matches (y-axis) versus retention time (x-axis)). This plot allows for a visual assessment of whether two QC runs exhibit different characteristics based on the number of PSMs obtained at a given retention time. Such a characteristic difference could indicate a shift toward longer retention times.

[0041] In some examples, sample metric values ​​(such as protein count, protein group count, peptide group count, peptide spectrum match count, or success rate) can also be calculated by Proteome Discoverer. In some cases, the sample metric values ​​are interrelated. For example, peptide spectrum matches (“PSMs”) can be combined when specific functions match to form a peptide group. Multiple peptide groups can form a protein group when specific functions match. In other words, a number of protein groups can be the result of PSMs and peptide groups. In some cases, the number of peptide groups and PSMs can be affected by smaller changes (LC, MS, sample, etc.) than the other metrics. Therefore, the number of protein groups may be a preferred sample metric.

[0042] Other characteristics for assessing the quality of a raw data file include the number of unfragmented (MS1) and fragmented (MS2) spectra. An MS spectrum can be composed of many peaks of varying intensity (in some cases, hundreds of peaks per scan). The complexity of the raw data can be reduced before processing (in some cases, to an absolute minimum). In some examples, this data reduction can be achieved by reporting a median, absolute, or relative value. All of these data reduction functions are extracted from the raw data file (e.g., by exporting values ​​from the scan header or directly from the MS spectrum). An example of data extracted from the scan header is the total ion current (TIC).Information extracted from MS or MS / MS spectra can be the total number of peaks per spectrum or the mean intensity of the peaks within a given spectrum.

[0043] Information extracted from the raw file can be further grouped at the MS-n level.

[0044] This level typically lies between 1 and 2 (i.e., MS1 and MS2) for routine identification-based analysis using LC-MS in proteomics experiments. Further grouping can be based on other specific functions. For example, data could be further subdivided if the ion injection time during acquisition of an MS or MS / MS spectrum was below or at the maximum permissible ion injection time.

[0045] This differs from commonly used methods for obtaining sample metric values ​​(such as the number of protein groups), where information from (ideally) every MS or MS / MS spectrum (i.e., m / z and its intensity pair) is used as input functions. In the approach described above, only one value for a function is used as a representative value to describe a raw file.

[0046] Further information and examples of the various data reduction steps are listed below: 1. Median

[0047] For example, the total ion current (TIC) MS1 changes throughout the proteomics run depending on the retention time when peptides (and other substances) are detected by the mass spectrometer. The median value of all MS1 TIC values ​​should differ only very slightly between different raw files acquired with the same sample and the same LC-MS method. 2. Absolute value

[0048] For example, the number of MS1 scans, which is an aggregated value, is reported as the "Total number of MS1 scans". This value remains very constant even with a constant LC-MS setup and is not negatively affected. 3. Relative value

[0049] For example, the ratio of the number of MS1 scans to the number of MS2 scans (MS1 / MS2) can be another useful property of the raw data.

[0050] The complexity of the data contained in the raw file can be reduced by extracting these characteristic features directly from the raw file (e.g., from the scan header of the raw file) while eliminating the retention time domain from the processed data. This reduced dataset can be paired with results from a database search engine to train a machine learning model. This trained machine learning model can then be used to predict the sample metric (e.g., number of protein groups) from data that the machine learning model has not previously seen ("new data"). The new data can be extracted from a new raw file that was not used to train the ML model. To determine the sample metric value(s) (e.g.,To quickly predict the number of protein groups, the new data provided to the trained model can include functions extracted from the new raw data file, rather than providing the entire new raw data file to the model. The reduced data can be processed much faster in commonly used proteomics tools (state-of-the-art methods) than the raw data.

[0051] The multitude of properties can include functions extracted directly from the scan header (e.g., total ion current), functions extracted from the MS or MS / MS scan (e.g., number of peaks in a scan), grouped functions (e.g., whether an MS / MS scan is less than or equal to the maximum allowable injection time), or calculated values ​​(e.g., ratio of MS1 to MS2 scans). Information classically used in prior art methods (i.e., m / z and intensity pairs from MS or MS / MS spectra) may not be required in the proposed methods, and therefore it is not strictly necessary to extract these properties and use them as input for the machine learning model.

[0052] The multitude of properties may include one or more of the following: a number of unfragmented scans; a number of fragmented scans; a ratio of unfragmented scans to fragmented scans; a number of peaks in fragmented and / or unfragmented scans; a mean total ion current for fragmented and / or unfragmented scans; a mean ion injection time for fragmented and / or unfragmented scans; a mean charge state value for fragmented and / or unfragmented scans; a mean intensity state value for fragmented and / or unfragmented scans; a medium resolution value for fragmented and / or unfragmented scans; a mean noise level for fragmented and / or unfragmented scans; an average current, voltage and / or pressure value; a minimum and / or maximum pressure value; a fill rate of the automatic gain control and / or an average fill value of the automatic gain control; an average base value; a mass-to-charge ratio of the precursor; a high signal-to-noise ratio; a medium scan resolution; an isolation window for the total ion current; and a total ion current for fragmented scans.

[0053] Preferably, only representative features are extracted from the raw data file and used as input (both for training the machine learning model and for predicting the sample metric value using a pre-trained machine learning model). In other words, the features used provide information about the LC-MS performance. Sources for extracting this representational information can include sections within the raw data file, such as the scan header, or from the MS and MS / MS scans.

[0054] Properties from the raw file can be further grouped into subcategories such as the MSn level (usually n = 1, 2) or the ion injection time (below or at the maximum permissible ion injection time).

[0055] If each raw file is considered as a whole, there can be one input value per property and raw file. If each raw file is divided into a multitude of retention-time slices, there can be one input value per property and retention-time slice per raw file.

[0056] The machine learning model can be trained using training data obtained from a variety of different LC-MS combinations. Furthermore, the training data can be obtained from one or more LC-MS instruments exhibiting a variety of different LC-MS performance states (training data can be collected from a single LC-MS instrument over a long period, thus reflecting a degradation in the instrument's performance). Obtaining data from a variety of LC-MS combinations and performance states can improve the model's performance.

[0057] The data size of the multitude of properties can be at least 10 times smaller than the data size of the raw data, preferably 100 times smaller, more preferably 1000 times smaller.

[0058] The machine learning algorithm can be used to predict one or more sample metrics (e.g., number of protein groups) using only properties related to the raw data as input. The machine learning algorithm may only process the multitude of properties (i.e., after reducing the data to a representative value) and may not directly process the raw data. In other words, the raw data may not be provided as input to the machine learning algorithm.

[0059] The machine learning algorithm can predict the sample metric in a shorter time than state-of-the-art methods that rely on database search software (such as Proteome Discoverer).

[0060] The time required to determine the multitude of properties from the raw data can be shorter than the time required to obtain the raw data by analyzing the proteomics sample, preferably at least ten times shorter.

[0061] To reduce the time required to determine the sample metric (which, as described above, may be the number of protein groups, the number of peptide groups, or peptide spectrum matches), the invention predicts the sample metric by inputting (data-reduced) functions (which can be quickly extracted from the raw files) into a trained machine learning model.

[0062] The multitude of features (functions extracted from the raw data) can be extracted from the raw data in a short time (e.g., 5 minutes or less, depending on the raw data size and the methods used). This time is short compared to the overall runtime of the LC-MS run (which can be around 80 minutes in some cases). In this way, the processing of the raw data from the proteomics run can be completed quickly, allowing hardware resources to be focused solely on processing high-quality raw data. If the raw data is of poor quality, it may not be worthwhile to invest time and resources in further processing. It may be preferable to skip data analysis altogether if the raw data is determined to be of low quality.Alternatively, the sample metric value predicted by machine learning can be used in combination for quality control.

[0063] The properties can be extracted from the raw files without tying up a significant portion of the available computing resources, as the process is not resource-intensive.

[0064] One or more of the numerous properties can be determined from the raw data, whereas LC-MS obtains the raw data by analyzing the proteomic sample (on-the-fly processing). More preferably, all of the numerous properties can be determined from the raw data, whereas LC-MS obtains the raw data by analyzing the proteomic sample.

[0065] Thus, the invention offers a fast and efficient method for providing a predicted, high-quality, sample-related metric value (such as the number of protein groups) for LC-MS systems that perform proteomics experiments in a controlled environment.

[0066] The machine learning model can be a supervised machine learning algorithm. In other words, during the model's training, in addition to the training data (the properties extracted from the raw training data), a sample-based metric determined using alternative (more accurate and time-consuming) methods can be provided.

[0067] Training data can be provided to implement the aforementioned procedures. The training data may include: a variety of properties extracted from raw training data; and one or more appropriate sample metrics (e.g., number of protein groups) obtained from a database search software (such as Proteome Discoverer).

[0068] Validation data can be provided to optimize the trained machine learning model. This validation data can include: a variety of properties extracted from raw validation data; and one or more appropriate sample metrics obtained from database search software (such as Proteome Discoverer).

[0069] Test data can be provided to evaluate the trained machine learning model. The test data can include: a variety of properties extracted from raw test data; and one or more appropriate sample metrics obtained from database search software (such as Proteome Discoverer).

[0070] Predictive data can be provided to the trained machine learning model. This predictive data can include: a multitude of properties extracted from raw data.

[0071] The process may further involve determining one or more functions from the raw data using a database search (such as an ultrafast database search). Predicting one or more sample metric values ​​may involve providing the multitude of properties and the one or more functions to the machine learning model.

[0072] A procedure for training a machine learning model to predict a sample metric value is provided. The procedure includes: Determining a variety of properties from raw training data obtained by analyzing a training sample in an LC-MS; Processing the raw training data via a database search to determine the sample metric value for the training sample; and Training the machine learning model using the multitude of properties and the sample metric value for the training sample, which was determined via database search, so that the machine learning model is trained to predict the sample metric value for a subsequent proteomics sample, based on a multitude of properties determined from raw data obtained by analyzing the subsequent proteomics sample in an LC-MS.

[0073] The numerous properties determined from the raw data obtained by analyzing the subsequent proteomics sample may be the same as those determined from the raw training data obtained by analyzing the proteomics training sample (e.g., number of MS1 scans, number of MS2 scans, MS1 / MS2 scan ratio, and the like). However, the values ​​of these properties differ between the samples.

[0074] In some examples, the model can be trained with a subset of the properties.

[0075] In some examples, a model trained on a specific set of properties (the "training set") requires an input value for each property in the training set to predict a sample metric value. In other examples, the model can be used to predict a sample metric value based on values ​​for a subset of the specific set of properties (the "prediction subset"). In this case, values ​​for one or more properties from the training set that are missing from the prediction subset can be replaced with a dummy value. The dummy value can be zero, a median value from all the training data, or another placeholder value. By adding these dummy values ​​to the values ​​in the prediction subset, the model can be provided with a set of values ​​that includes one value for each property in the training set to predict the sample metric value.Because some functions are only "placeholders," the prediction may be less accurate than if values ​​for each property in the training set were provided for the prediction. The procedure may also involve training the machine learning model using one or more of the following functions: Functions from the LC (e.g., pressure values); Functions from one or more calibration files; and / or Features from another database search engine.

[0076] The procedure may also include validating the machine learning model by: Determining a variety of properties from raw validation data obtained by analyzing a validation sample in an LC-MS; Predicting a sample metric value for the validation sample by processing the multitude of properties using the machine learning model; Processing the raw validation data via a database search to determine the sample metric value for the validation sample; and Comparing the sample metric value determined by processing the raw validation data via database search with the sample metric value predicted by processing the multitude of properties using the machine learning model.

[0077] The procedure can include determining whether the predicted value (which can be obtained relatively quickly) is sufficiently close to the determined value (which, while accurate, can be obtained slowly). In other words, the machine learning model can be validated by feeding the trained machine learning model properties obtained from raw validation data to provide a predicted sample metric value, and also by accurately determining the sample metric value by feeding the raw validation data to a database search engine. The sample metric value could be the number of protein groups, peptide groups, PSMs, the median success rate, and the like.

[0078] The LC-MS method used to collect the training data can be the same as the LC-MS method used to collect the validation, testing, and prediction data.

[0079] The LC-MS can be the same LC-MS used to collect the training dataset, or a different LC-MS. It is not strictly necessary for the validation, test, or prediction dataset to be collected using the same LC-MS.

[0080] The accuracy of the model prediction can be improved if the datasets (training data, validation data, test data and / or prediction data) originate from the same LC-MS instrument and / or use the same standardized LC-MS procedures (this is described in detail below with regard to applications B, C and D in the Fig. 12, Fig. 13 and Fig. Figure 14 shows where different LC-MS systems were used, but with the same standardized LC-MS procedures).

[0081] The process can further involve determining one or more functions from the raw training data using a database search (such as an ultrafast database search). The machine learning model can be trained using the multitude of properties, the sample metric value for the training sample determined via the database search, and the one or more functions determined from the database search. The machine learning model can therefore be configured to predict the sample metric value for a subsequent proteomics sample based on the multitude of properties determined from the raw data and one or more functions determined from the raw data using a database search.

[0082] The step of predicting one or more sample metric values ​​using machine learning can be accomplished using one or more different machine learning models. Examples of usable machine learning models include: linear regression, XGBoost, Random Forest, deep learning, and / or a combination thereof (this will be described in detail below in relation to…). Fig. 4, Fig. 5 and Fig. 7 shown).

[0083] Predicting one or more sample metric values ​​can include predicting a single sample metric (e.g., number of protein groups) or predicting multiple sample metric values ​​simultaneously (e.g., number of protein groups, number of proteins, number of peptide groups, number of peptide spectrum matches, and / or success rate) by providing the same sample functions / properties extracted from the raw file as input data for one or more trained machine learning models. This will be described in detail below. Fig. 4, Fig. 5 and Fig. 7 shown.

[0084] The analysis of the proteomic sample by LC-MS can include performing an LC-MS run. The raw data can include data relating to a retention time during the LC-MS run. The procedure can further include dividing the retention time into a multitude of retention time slices. The procedure can also include manipulating the raw data to provide data for each of the multitude of slices. The multitude of properties can include a multitude of properties for each retention time slice.

[0085] As described above, the procedure can further include slicing the raw file into sections based on retention time (e.g., 1-minute blocks). This approach helps to (i) increase the volume of training data and (ii) provide more insight into the LC-MS run by providing a predicted value for the number of peptide spectrum matches (PSMs) at discrete and specific time intervals. This approach includes an additional processing step to obtain the multitude of properties. Here, a property is calculated based on the specific slice of the raw file (e.g., between retention times of 2 and 3 minutes; referred to as "local"), but also on the retention time range from start to the specific time interval (e.g., from retention time 0 to 3 minutes; referred to as "global"). This process can be performed for the sample metric values ​​and for the multitude of properties.

[0086] The sample metric value can be used to determine the data quality of raw data obtained by LC-MS. In other words, the sample metric value can be used to determine the result of a data quality test for the LC-MS run, which can be either "pass" or "fail." A data quality result should be "fail" if the data quality is so low that further analysis of the raw data is not worthwhile.

[0087] Furthermore, a method for determining the data quality of raw data obtained by analyzing a proteomics sample in a liquid chromatography-mass spectrometer (LC-MS) is provided. The method includes predicting one or more sample metric values ​​using the procedures described above. Based on the one or more sample metric values, the method further includes determining whether the raw data quality exceeds a threshold. If the raw data quality exceeds a threshold, the method further includes performing additional data analysis on the raw data. Otherwise, if the raw data quality does not exceed the threshold, the method further includes skipping additional data analysis on the raw data.

[0088] The sample metric value can be used to validate a QC run. In other words, the sample metric value can be used to determine a result for the QC run, which can be either "pass" or "fail." A QC run should result in a "fail" if maintenance (or other action) is required.

[0089] A method for determining a quality control (QC) result for a liquid chromatography-mass spectrometry (LC-MS) process is provided. The method comprises predicting one or more sample metric values ​​in a QC run by analyzing a proteomic sample in a liquid chromatography-mass spectrometer (LC-MS) using one of the methods described above. The proteomic sample is the QC sample. The one or more sample metric values ​​comprise a QC sample metric. The method further comprises determining, based on the QC sample metric, whether the QC result is pass or fail.

[0090] Several approaches are available to determine when maintenance is required (and thus whether the result should be "pass" or "fail") based on the QC sample metric. One approach defines that cleaning / maintenance is required when performance drops by a threshold (e.g., 10%) of the original performance. For example, the number of protein groups can be used as the QC sample metric to provide a performance marker / indicator. If the QC sample metric falls below a threshold (e.g., 90%) of the original value, it can be assumed that cleaning / maintenance of the MS is required.The result of a QC run can therefore be “fail” if the QC sample metric value falls below the threshold percentage of the original value (which can be determined in a first QC run after cleaning / maintenance or can be an average of several QC runs, e.g. the first ten QC runs).

[0091] There may also be situations where the QC sample metric suddenly drops significantly more than the threshold. This could be due to a problem or error in the LC, the sample, or another part of the MS. In this case, the QC result should be marked as "fail," allowing the sequence to be stopped, the problem to be resolved, and measurements to be resumed in a new sequence. Alternatively, a sudden drop in the QC metric could be an anomaly.

[0092] In another approach, LC-MS measurements can therefore be continued after the QC sample metric value has dropped in order to determine the following: (i) whether the decline in performance is temporary or permanent (e.g. following a gradual decline in performance) and (ii) whether an unexpected recovery is observed (e.g. after a sudden drop in performance).

[0093] As described above, if the result of a QC run falls below a predefined threshold, the prediction of the QC sample metric value can be repeated in further cycles to verify whether the measurement of "lower performance" is correct (or whether it is anomalous or recoverable). Further cycles can also be performed after the threshold has been violated for other reasons, as explained below.

[0094] If the performance falls below the threshold, whether the sequence is continued (with a "fail" result) or continued (with a "pass" result) may depend on the degree of deviation of the QC sample metric value below the threshold and the specific situation. For example, if a long-term measurement of biological samples is underway and a QC run at the very end of the long-term measurement gradually falls below the threshold, but the deviation below the threshold is still small, it may be determined that the measurements of the biological samples should continue and that maintenance of the LC-MS system should be performed after all biological samples in the same sequence have been measured by LC-MS.

[0095] However, if the deviation of the QC sample metric value between the QC run result and the original level (or the acceptance threshold level) is too high, it may be determined that the sequence of long-term measurements should be interrupted to perform maintenance. The QC run is then repeated and evaluated. If the QC sample results in a "pass" rating, the measurement of the biological samples can continue.

[0096] A procedure for performing a liquid chromatography-mass spectrometry sequence is provided. The sequence comprises one or more cycles, each cycle comprising the following steps: Performing a quality control (QC) run and determining a QC result using the procedure described above, If the QC result is "fail", end the sequence without further cycles; and If the QC result is "pass", perform one or more sample analysis runs (in which biological samples are analyzed) and proceed to the next cycle.

[0097] In this procedure, a cycle is defined as one QC run and one or more sample analysis runs. The sequence continues until the result of a QC run is "fail." The sequence is then terminated so that maintenance can be performed.

[0098] A liquid chromatography-mass spectrometer (LC-MS) will be provided, which is also configured to perform the procedures described above.

[0099] Computer software comprising instructions which, when executed by a computer's processor, cause the computer to perform the procedures described above, is also provided. PARAGRAPHS

[0100] The following numbered paragraphs provide further illustrative examples: Paragraph 1. Procedure for estimating a quality control metric, QC metric, wherein the procedure comprises: Determining a variety of properties from raw data obtained by analyzing a QC sample in a liquid chromatography-mass spectrometer (LC-MS); and Estimating the QC metric by processing the multitude of properties using a machine learning algorithm. Paragraph 2. Procedure according to paragraph 1, wherein the QC metric comprises one of the following: a number of protein groups in the QC sample that are identifiable by LC-MS; a number of peptide groups in the QC sample that are identifiable by LC-MS; and a number of peptide spectrum matches in the QC sample that are identifiable by LC-MS. Paragraph 3. Procedure according to paragraph 1 or paragraph 2, wherein the plurality of characteristics includes one or more of the following: a number of unfragmented, MS1 scans; a number of fragmented MS2 scans; a ratio of unfragmented scans to fragmented scans, MS1 / MS2; a number of peaks in unfragmented scans; a number of peaks from fragmented scans; a mean total ion current for unfragmented scans; a mean total ion current for fragmented scans; a mean ion injection time for unfragmented scans; a mean ion injection time for fragmented scans; a mean charge state value for unfragmented scans; a mean charge state value for fragmented MSn scans; a mean intensity state value for unfragmented scans; a mean intensity value for fragmented MSn scans; a medium resolution value for unfragmented scans; a mean resolution value for fragmented, MSn scans; a medium noise level for unfragmented scans; a mean noise value for fragmented MSn scans; an average current value; a medium voltage value; a medium pressure value; a minimum pressure value; a maximum pressure value; a filling rate with automatic gain control; a medium fill level with automatic gain control; a medium base value; a mass-to-charge ratio, m / z ratio, of precursors; a signal-to-noise ratio (S / N ratio) of precursors; a medium scan resolution; an isolation window for the total ion current, TIC; and a total ion current, TIC, for fragmented scans. Paragraph 4. Method according to one of the preceding paragraphs, wherein a data size of the plurality of properties is at least 10 times smaller than a data size of the raw data, preferably 100 times smaller, more preferably 1000 times smaller. Paragraph 5. Procedure according to one of the preceding paragraphs, wherein the machine learning algorithm only processes the multitude of properties and does not directly process the raw data. Paragraph 6. Method according to any of the preceding paragraphs, wherein the time required by the machine learning algorithm to estimate the QC metric is shorter than the time required by a proteome discoverer algorithm to determine the QC metric based on the raw data. Paragraph 7. Method according to one of the preceding paragraphs, wherein the time required to determine the multitude of properties from the raw data is shorter than the time required to obtain the raw data by analyzing the QC sample, preferably at least ten times shorter. Paragraph 8. Method according to one of the preceding paragraphs, wherein one or more of the multitude of properties are determined from the raw data, while the LC-MS obtains the raw data by analyzing the QC sample. Paragraph 9. Procedure according to any of the preceding paragraphs, furthermore including: Determining one or more parameters from an ultra-fast database search based on raw data, Estimating the QC metric using a machine learning algorithm involves providing the multitude of properties and one or more parameters for the machine learning algorithm. Paragraph 10. Method for training a machine learning algorithm to estimate a quality control metric, QC metric, wherein the method comprises: Determining a variety of properties from raw training data obtained by analyzing a training sample in an LC-MS; Processing the raw training data via a database search to determine the QC metric for the QC training sample; and Training the machine learning algorithm using the multitude of properties and the QC metric for the QC training sample determined via database search, so that the machine learning algorithm is configured to estimate the QC metric for a subsequent QC sample based on a multitude of properties determined from raw data obtained by analyzing the subsequent QC sample in an LC-MS. Paragraph 11. Procedure according to paragraph 10, further comprising the validation of the machine learning algorithm by: Determining a variety of properties from crude validation data obtained by analyzing a QC validation sample in an LC-MS; and Estimating a QC metric for the QC validation sample by processing the multitude of properties using the machine learning algorithm; Processing the raw validation data via a database search to determine the QC metric for the QC validation sample; and Comparing the QC metric determined by processing the raw validation data via database search with the QC metric estimated by processing the multitude of properties using the machine learning algorithm. Paragraph 12. Method for determining a quality control result, QC result, for a liquid chromatography-mass spectrometry process, LC-MS process, wherein the method comprises: in a QC run, estimating a QC metric using the method of one of paragraphs 1 to 8; and Based on the QC metric, determine whether the QC result is "pass" or "fail". Paragraph 13. Method for performing a liquid chromatography-mass spectrometry sequence, wherein the sequence comprises one or more cycles, each cycle comprising the following steps: Performing a quality control run (QC run) and determining a QC result using the procedure in paragraph 12, If the QC result is "fail", end the sequence without further cycles; and If the QC result is "pass", perform one or more sample analysis runs and proceed to the next cycle. Paragraph 14. Liquid chromatography-mass spectrometer, LC-MS, configured to carry out the procedure according to any of the preceding paragraphs. Paragraph 15. Computer software comprising instructions which, when executed by the processor of a computer, cause the computer to carry out the procedure referred to in any one of paragraphs 1 to 13. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates a schematic diagram of a mass spectrometer. The invention is described with reference to the non-limiting examples illustrated in the following figures. Fig.Figure 2 illustrates, for each QC run in a test dataset, the number of normalized protein groups predicted using the proposed methods, as well as the number of protein groups determined by processing the raw data via a database search. The same LC-MS instrument was used to collect both the training and test datasets. Fig. Figure 3 illustrates, for each QC run in a test dataset, the number of normalized protein groups predicted by the algorithm of the proposed methods, as well as the number of protein groups determined by processing the raw data via a database search. Two different LC-MS instruments were used to collect the training and test datasets. Fig.Figure 4 illustrates, for each raw file in a test dataset, the number of normalized protein groups predicted by three proposed machine learning models, as well as the number of normalized protein groups determined by processing the raw data via a database search. Different LC-MS instrument combinations were used to collect the training and test datasets. Fig. Figure 5 illustrates, for each raw file in a test dataset, the number of normalized protein groups predicted by an average of three machine learning models, as well as the number of normalized protein groups determined by processing the raw data via a database search. Different LC-MS instrument combinations were used to collect the training and test datasets. Fig.Figure 6 illustrates the distribution of the relative error between the predicted number of protein groups and the number of protein groups determined by processing the raw data via a database search. The error has been rounded to one digit. Fig. Figure 7 illustrates, for each raw file in a test dataset, the number of normalized peptide spectrum matches (PSMs) predicted by averaging three machine learning models, as well as the number of normalized peptide spectrum matches (PSMs) determined by processing the raw data via a database search. Different LC-MS instrument combinations were used to collect the training and test datasets. Fig.Figure 8 illustrates the distribution of the relative error between the predicted number of peptide spectrum matches and the number of peptide spectrum matches determined by processing the raw data via a database search. The error has been rounded to one digit. Fig. Figure 9 illustrates, for a single high-quality raw dataset in a test dataset, the number of normalized peptide spectrum matches predicted by a machine learning model (multiple linear regression), the number of normalized peptide spectrum matches predicted by an average of three machine learning models, and the number of normalized peptide spectrum matches determined by processing the raw data via a database search (Proteome Discoverer) for specific time intervals (1 minute). Various LC-MS instrument combinations were used to collect the training and test datasets. Fig. Figure 10 illustrates, for a single, very low-quality raw dataset in a test dataset, the number of normalized peptide spectrum matches predicted by a machine learning model (multiple linear regression), the number of normalized peptide spectrum matches predicted by an average of three machine learning models, and the number of normalized peptide spectrum matches determined by processing the raw data via a database search (Proteome Discoverer) for specific time intervals (1 minute). Various LC-MS instrument combinations were used to collect the training and test datasets. Fig.Figures 11A to 11D illustrate data obtained during four different sequences on the same LC-MS setup. These figures illustrate, for each QC run in the sequence, the (normalized) number of protein groups obtained by processing the raw data with a database search. Fig. Figure 12 illustrates, for each QC run in a validation dataset, the (normalized) number of protein groups predicted by the proposed methods as well as the (normalized) number of protein groups determined by processing the raw data via a database search. Fig.Figure 13 illustrates, for each day during which data were collected in the validation dataset, the (normalized) average number of protein groups predicted by the machine learning model compared to the (normalized) average number of protein groups determined by processing the raw data via a database search. Fig. Figure 14 illustrates the performance of the proposed methods on a validation dataset obtained during a sequence run on a different LC-MS setup. Fig. Figure 14 illustrates the (normalized) number of protein groups predicted by machine learning and the (normalized) number of protein groups determined by processing the raw data via a database search. Here, as in Fig. 13, several results from a single day averaged and shown as a single data point for a single day. Fig.Figure 15 illustrates the performance of the proposed methods on a validation dataset obtained during a sequence run on a third LC-MS setup. Fig. Figure 14 illustrates the (normalized) number of protein groups predicted by machine learning and the (normalized) number of protein groups determined by processing the raw data via a database search. Here, as in Fig. 13 and Fig. 14, several results from a single day averaged and shown as a single data point for a single day. Fig.Figure 16 shows a graph of the (normalized) number of protein groups predicted / determined from validation data. The predicted (normalized) number of protein groups from two machine learning algorithms is compared to the (normalized) number of protein groups determined by processing the raw data via a database search (a thorough database search). In the first machine learning model, the inputs to the ML model are functions from a fast database search. In the second machine learning model, the inputs to the ML model are functions from a fast database search and properties extracted from the raw data. DETAILED DESCRIPTION

[0101] The most commonly used sample metric for evaluating sample runs (including QC runs) in proteomics labs is the sample metric "Total Number of (Identified) Protein Groups." The value of this sample metric is a combination of various other sample metrics (peptide groups, peptide spectrum matches (PSMs)) and filter options, which makes its value relatively robust (and therefore a suitable sample metric for QC application). Although this sample metric is a combination of several functions, a decrease in the value of the sample metric for the number of protein groups indicates that the LC-MS system (including samples and solvent) is not performing as expected when analyzing a known sample using a standardized LC-MS procedure.

[0102] A sample metric value for a number of protein groups can be obtained by performing a database search using software such as Proteome Discoverer (“PD”). Since running Proteome Discoverer is a time-consuming and computationally intensive process, an alternative solution is proposed that utilizes supervised machine learning to predict one or more sample metric values ​​(such as the number of protein groups) using only properties (features) extracted from the raw file. These raw-file-associated features can be extracted from the raw file in a relatively short time (preferably less than 5 minutes) compared to the total runtime of the LC-MS run (in some cases, approximately 80 minutes). The process of extracting the properties from the raw file is not hardware-intensive and can be performed while the LC-MS system continues to acquire raw files.

[0103] Ideally, the multitude of properties are extracted while the proteomics sample (for example, a QC sample) is still being analyzed (on-the-fly processing). The final calculation of the median, relative, or aggregated values, which are used as input for the machine learning model, is typically a fast process. Predicting protein groups (or other commonly used sample metrics) using machine learning is also very fast.

[0104] In an ideal situation (with the QC run processing in progress), the software controlling when a new sample is injected for LC-MS analysis would (i) only need to wait a very short time (e.g., < 1 minute) to obtain a predicted sample metric (e.g., number of protein groups) reported by a trained machine learning model. Alternatively (or additionally), (ii) the number of peptide spectrum matches (PSMs) for one or more specific retention time ranges would be predicted by machine learning.

[0105] Depending on the value of the sample metric predicted by machine learning, this can be interpreted as a positive or negative event. This event is provided to the instrument control software. No further sample is injected before this response is provided (although sample injection is software-controlled, which is not strictly necessary).

[0106] Alternatively, in other proposed examples, the control software can either: (i) inject a new biological sample before the reaction takes place. This sample can be left unused if the run is aborted in the event of a negative reaction), or (ii) wait for user input (the user can manually check the result of the QC run and enter a result) indicating whether to continue with the analysis of ‘biological’ samples or to interrupt the sequence.

[0107] To train the machine learning model and validate its performance in predicting a sample metric value (e.g., the number of identified protein groups), at least two different datasets are required in these examples: 1. Training data set

[0108] This dataset contains the extracted information from raw files as well as the identification results from a database search software (such as Proteome Discoverer). It is used to train the machine learning model. 2. Test data set

[0109] This dataset contains the extracted information from raw files used to predict the number of protein groups (or other sample metrics) using previously trained machine learning models. The test dataset also includes information from the database search to evaluate the results obtained by the machine learning model. This is a useful step during development to prevent misinterpretation of sample metric values ​​predicted by machine learning. 3. (Optional) Validation data set

[0110] This dataset contains the extracted information from raw files as well as the identification results from a database search software (such as Proteome Discoverer). It is used to optimize the machine learning model. Often, this dataset is a subset of the training dataset (and is removed from the training dataset).

[0111] Once the machine learning algorithm has been trained and evaluated, it can be used to process "live" data (e.g., data from a QC run where an estimation of the QC metric is required) and to predict the number of protein groups (or another sample metric). In this case, no identification results from a database search software are currently available. 4. Forecast dataset

[0112] The information obtained from the database search is not needed in its final application.

[0113] Only the raw file-related information from an LC-MS run (using the same sample and the same LC-MS method) is used to predict one or more sample metric values ​​(e.g., the number of protein groups) by machine learning.

[0114] Raw proteomics files (e.g., obtained from QC runs) can contain MS1 ​​and MS2 spectra with hundreds of MS peaks per scan. The proposed methods reduce the complexity of the data to an absolute minimum before processing. This data reduction can be achieved by reporting median, absolute, and / or relative values, as described above.

[0115] Reducing the complexity of the raw data simplifies the creation of numerous metrics for the test and training data. In some examples, the data volume required to predict the sample metric value is reduced by eliminating the retention time domain. In other examples, the raw file can be decomposed into specific retention time windows, and properties can be extracted from each of these windows. The retention time domain information is also removed from this reduced dataset.

[0116] In some examples, there are hundreds of functions in the test and training datasets that can be extracted from the raw data (depending on the LC and MS used). The sample metrics to be predicted can be provided by a database search engine. All functions used to train and process the "unseen" / new sample can be grouped into subcategories: a. Identification results from a database search engine

[0117] Examples: Protein groups, peptide groups, number of peptide spectrum matches (PSMs), success rate (= number of peptide spectrum matches / number of MS2 scans) b. Absolute and relative values ​​from the QC run

[0118] Examples: Number of MS1 scans, number of MS2 scans, ratio of scans (MS1 / MS2) c. Median values ​​from ion transmission

[0119] Examples: TIC MS1, Ion injection time MS1, TIC MS2, Ion injection time MS2 d. Median values ​​from MSn scans

[0120] Examples: charge states, intensity, resolution, noise e. Median values ​​from the status log

[0121] Examples: Current and voltages of circuit boards installed in the MS f. Median, minimum and maximum values ​​for the LC. Examples: pressure values

[0122] Values ​​from subcategory a) are provided in the training validation and test datasets and are used to train, optimize, and test the predictions made by the machine learning model. For predicting sample metrics (such as the number of protein groups), these values ​​are not usually available because retrieving them from database search engines (as described above) can be time-consuming.

[0123] Many state-of-the-art search engines rely solely on MS2 peaks and intensities. Therefore, the proposed approach is superior to state-of-the-art methods, at least because it utilizes more information extracted from the raw data.

[0124] Some state-of-the-art methods evaluate QC runs using database search engines and validate the LC-MS setup based on the total number of protein groups. Database searching is a computationally and time-intensive approach. To reduce the required time and hardware resources, the proposed methods use machine learning to predict a sample metric value (such as the number of protein groups) based solely on properties extracted from the raw data. These properties can be extracted from the raw data in a relatively short time and with fewer hardware resources than a full database search. This allows the state of the LC-MS system to be determined by predicting a discrete value (e.g., the number of protein groups) shortly after the QC run is completed.Furthermore, the properties extracted from the raw file are orthogonal to the search engine results, since the properties extracted from the raw file relate to the entire LC-MS system, while the search engine results relate (mainly) to the MS2 data quality.

[0125] In state-of-the-art methods, a relatively small number of QC runs can be performed during the instrument's peak performance period to determine a lower limit on the number of protein groups before a QC run is marked as "failed" (a threshold). The difference between QC runs can be calculated as the Median Absolute Deviation (MAD). This MAD value can be multiplied by a factor (e.g., a factor of 2 or 3) to determine the lower warning level. However, this approach is only effective over a short timeframe.

[0126] In contrast to the state-of-the-art approach, examples of the proposed approach provide a training dataset that covers various phases (such as time-series trials and / or different instruments) of the LC-MS setup to train the ML algorithm. In other words, data is obtained while the setup is operating at peak performance, but also when the LC-MS setup (1) deviates from expectations or (2) differs in type (i.e., different MS from the same model family) to make a good prediction about the quality of the raw data. Nevertheless, after training, the ML model can provide the results of a QC run faster than state-of-the-art approaches (in some examples, in real time / on-the-fly).

[0127] In some examples, two data reduction steps are performed to reduce data complexity: i. Reducing a parameter to a single value

[0128] Raw LC-MS data contain an m / z domain and a (retention) time domain because the samples are separated by LC and the peptides are elutated based on their physicochemical properties. If all parts of the LC-MS setup perform as expected, representing a function by its mean or median value is a valid approach, as this value should not change if the LC-MS setup remains unchanged. Any change represents a deviation that can negatively affect the LC-MS setup.

[0129] Alternatively, the raw file can be split into smaller parts based on the retention time, and representative values ​​could be extracted for these parts. ii. Reduction of the number of functions

[0130] The total number of features is reduced from up to several hundred (available via the raw file) to just a subset of relevant properties (approximately 30 properties in the specific examples described below) that can effectively represent the performance of an LC-MS setup. During machine learning model training, it turns out that only this subset of features is needed to accurately predict the sample or QC metric. In fact, using this subset of features as machine learning inputs resulted in particularly good model performance (both in terms of accurate prediction and fast processing time). Adding more features can improve prediction results, but it doesn't necessarily have to. In some cases, adding more features can negatively affect the predicted results (e.g., by reducing the sample size).B. if readback or calibration values ​​do not correlate with the number of protein groups).

[0131] As explained above, the number of protein groups is a useful parameter for evaluating QC runs. In state-of-the-art database methods, the number of protein groups (including some downstream processing to construct protein groups from PSMs via peptide groups via proteins) is based on MS2 data quality. The process of matching experimental and theoretical peptide spectra to obtain peptide spectrum matches is a hardware- and time-intensive process.

[0132] In contrast, predicting protein groups (or other sample-related metrics) for proteomic LC-MS data in a controlled environment (i.e., using a standardized LC-MS procedure) via machine learning is based on only a few key features extracted from a raw data file. Typically, between 30 and 40 features are used (depending on the mass spectrometer employed), rather than the hundreds of features available in the raw data that could be considered key. This reduction in data complexity improves data processing. In terms of quality control, this allows the instrument control software to quickly decide whether (or not) the sequence needs to be interrupted, based on the sample metric value predicted by machine learning.Since the next run might not begin until the sample metric value has been determined (the sample could remain unused if injected before the "fail" result from the immediately preceding QC run has been determined), the time spent processing the data from the QC run can delay the workflow for the remainder of the sequence (in other words, processing the data from the QC run can create a bottleneck). Preferably, the QC result can be determined before the system is ready to start the next run (which may be a production run). In some examples, the data from the QC run is analyzed while further data from the same run is still being collected (on-the-fly processing). The proposed methods offer clear advantages over classical methods (e.g., database search engines) because the sample metric value(s) can be predicted very quickly. Test data

[0133] This section uses two different datasets to demonstrate three example applications of using machine learning to predict sample metric values ​​(such as the number of protein groups). The input to the machine learning models includes extracted properties and representative functions from raw data files. These functions are significantly reduced from the raw data to keep the approach lightweight and enable fast prediction of sample metric values. I. QExactive HF-X data from long-term measurements

[0134] The first dataset consists of a long-term measurement performed on a QExactive HF-X coupled to an Ultimate 3000. Standardized measurements (i.e., using the same sample and the same LC-MS method) were injected as a quality control sample between real biological samples over an extended period.

[0135] The training dataset (designated "Instrument 06") consisted of 400 raw files. Each raw file provided a single sample metric for prediction, which in this example is the number of protein groups. Additionally, 77 functions (i.e., 77 representative values ​​for a single raw file) were extracted from the raw data, relating to (and corresponding variations of) the following: number of spectra, number of peaks per spectrum, ion injection times, ion fluxes and intensities, charge states, signal-to-noise ratio, precursor information, and information about whether enough ions were present to collect a scan. Variations of these functions could be specific to MS1 or MS2, or relative values, by relating one function to another. Approximately half of these functions were obtained by further grouping the functions into a subcategory.In this example, the ion injection time (below or at maximum ion injection time) was used for grouping.

[0136] The first test dataset was acquired using the same LC-MS setup (Instrument 06) and the same environment (i.e., the LC-MS procedure was kept constant). Consumables such as the sample and the LC column were kept as constant as possible. This dataset is referred to as "Instrument 06b" in these examples and consists of 404 raw files. The functions were extracted in a similar manner to the training dataset, as described above.

[0137] For visualization purposes, the protein groups were normalized to the highest value obtained by Proteome Discoverer. The predictions of a trained machine learning model are shown in Fig.Figure 2 shows that the prediction was very accurate for the first (approximately) 100 raw files. For raw files 100 and 150 (approximately), the prediction was less accurate (marked with "A").

[0138] In the approximate range of raw files 200 and 250 (labeled "B"), the prediction was again more accurate. Remarkably, between raw files 245 and 400 (approximately), a drop in performance to 60% normalized protein groups was observed, as predicted – although the prediction indicated lower values ​​compared to the true values ​​(labeled "C").

[0139] A second test dataset was collected using the same environment (i.e., the same LC-MS procedure and consumables kept as constant as possible), but a different LC-MS (Instrument 08). Using the same trained machine learning model, the predicted data were very accurate up to raw file index 250 (see Fig.3) Even minor performance drops (e.g. around raw file index 110; see label “A”) could be accurately predicted.

[0140] The general curve of decreasing performance could also be predicted for the raw files with raw file indexes of 250 to 350. For some specific regions, the prediction was much lower than the true value obtained by Proteome Discoverer (marked with "B"). Towards the end of the dataset, the prediction was again closer to the true value. II. Orbitrap Exploris 480 data from various LC-MS combinations

[0141] The above-described application of machine learning to predict a sample metric (e.g., number of protein groups) was based on training a machine learning algorithm using raw data (the training dataset) collected over a long period using the same LC-MS setup (Instrument 06). It has already been shown that trained models can also be used to predict sample metrics from different instruments using the same LC-MS procedure.

[0142] In another example, described below, the training data were obtained from various LC-MS combinations. The LC-MS method used to acquire the raw data was kept constant for all data sets.

[0143] The entire dataset comprises approximately 2500 raw files. 80% of these raw files were used to train the machine learning algorithms, and 20% were used for testing.

[0144] This approach differs from the previously described approach in how the input / training data were generated (using different LC-MS setups) and also in the machine learning approach used for prediction. In this example, the sample metrics "number of proteins," "number of protein groups," "number of peptide groups," "number of peptide spectrum matches," and the "success rate [number of peptide spectrum matches / number of MSIMS scans]" were predicted simultaneously by three different machine learning algorithms (multiple linear regression, XGBoost, and Random Forest). This differs from the previous example (long-term test with two QE HF-X instruments), where only the number of protein groups was predicted by machine learning using multiple linear regression.

[0145] The prediction for three trained machine learning models (Multiple Linear Regression, XGBoost, and Random Forest) compared to the results obtained from Proteome Discoverer is shown in Fig. Figure 4 illustrates this. In this diagram, the "number of proteins" is normalized to 100% using the highest value in the dataset. To reduce variations from individual machine learning predictions, the results obtained were averaged ("AveragedML" in Figure 4). Fig. 5).

[0146] Although most of the data were in the 80-90% range (relative scale to highest value), values ​​at 60%, 30%, or close to 0% could also be accurately predicted (in Fig. 5 labelled A, B and C).

[0147] The mean absolute error for all predictions relative to the true value reported by Proteome Discoverer was approximately 5%. The median was calculated to be 1.3%. As in Fig.As stated in Figure 6, the prediction was very accurate for 95% of the data (up to and including 5% relative error). Consequently, 5% of the predicted values ​​showed a very high deviation (25 predictions were off by more than 5%). In some rare cases, the relative mean absolute error was much higher (> 100%). An example is given in Figure 6. Fig. 5 marked with D'. Here the relative deviation was approximately 400%.

[0148] For the “number of peptide spectrum matches” (PSMs), in Fig. Figure 7 shows the average predictions from three machine learning models.

[0149] Overall, the sample metric values ​​predicted by the machine learning models (as illustrated by the predictions labeled A, B, and C) are close to the "true" values ​​obtained by Proteome Discoverer. Again, the data point labeled "D" represents an example where the prediction deviates significantly from the "true" value.

[0150] As previously shown, the average relative error is skewed by some "extreme" outliers. In this case, a mean absolute error of 2% was calculated. Compared to the number of proteins, this is somewhat higher. As described above in the prediction of the number of proteins, individual predictions can differ and exhibit a higher degree of deviation (51 predictions were off by more than 5%, as shown in a comparison of Fig. 8 with Fig.6). In particular, for runs of low quality (i.e. runs with a relative PSM value of 0.1% or less), the deviation between the true and the predicted value was very high. III. Orbitrap Exploris 480 data from different LC-MS combinations for predicting peptide spectrum matches per retention time interval

[0151] In the previous section, several sample metric values ​​(such as the number of protein groups) were predicted by multiple trained machine learning models. Some example procedures involve extracting a representative value for each property or “function” from a raw file. In other words, if 30 functions were extracted from a raw file, the raw file was described by those 30 functions.

[0152] To benefit from other data available in the raw file (but still significantly reduce data size and processing time compared to Proteome Discoverer), an alternative approach is presented below.

[0153] In this approach, the raw data file is divided into specific retention time windows, and two representative values ​​are extracted for each slice and function. In the case of 30 extracted functions (such as the number of MS1 scans) for a single raw data file, the raw data file is described by 60 functions per slice. These consist of 30 "local" functions (related to a specific retention time slice) and 30 "global" functions (related to the period from the start up to and including that retention time slice). Additionally, the number of input data points is increased by a factor calculated as "LC-MS runtime divided by the size of the retention time window." The retention time information is removed from the training and test datasets.

[0154] The sample metric “peptide spectrum matches” is well suited for use in this retention time slicing approach because the sample metric PSM is directly linked to a specific MS / MS spectrum.

[0155] In Fig. Figure 9 shows the PSMs predicted and obtained by Proteome Discoverer at 1-minute intervals. The PSMs were normalized to the highest value in the training dataset. The PSMs predicted for each time interval agree very well with the "true" value obtained by Proteome Discoverer. This holds true for a single machine learning model using multiple linear regression and also for the averaged machine learning result from three models (referred to as "AveragedML").

[0156] This example represents a commonly encountered curve shape, as can be seen from the final relative PSM value of approximately 80%.

[0157] Prediction accuracy is lower when "outliers" are predicted relative to the total training data. This is due to the fact that... Fig. Figure 10 illustrates the normalized PSM versus retention time; it forms a curve with a final relative PSM value of 3.5% compared to the highest PSM in the training dataset. The multiple linear regression machine learning model predicted a final relative PSM value of 20%, which was much higher than the true value (3.5%).

[0158] By averaging the three trained machine learning models (“AveragedML”), the final prediction was 13%. Although the difference between the true PSM value obtained by Proteome Discoverer and by machine learning is very high, the results still indicate that this raw data set is of lower quality. This can be particularly useful when analyzing QC samples to obtain rapid feedback on the performance of the LC-MS setup. IV. Variations of the extracted functions used for machine learning and the optimal training dataset

[0159] Finally, variations in the training set and the corresponding functions extracted from the raw data to predict sample-specific metrics (e.g., the number of protein groups or peptide spectrum matches) can also include functions that cannot be directly read from the raw data. This can cover QC data from classical evaluation processes (a priori). Such data may be related to LC (e.g., as effective gradient length, FWHM, retention time stability of added peptides such as thermo PRTC, or pump pressure profiles) or may originate from another (but faster) search algorithm / search engine. This makes the overall approach of using machine learning to predict a sample metric in a controlled environment (i.e., by keeping the LC-MS procedure constant) flexible. As with any machine learning project, the training data is of paramount importance.Ideally, the data set is collected using various LC-MS combinations run over a long period of time to obtain training data at peak performance as well as at lower performance levels. Further test data

[0160] To further demonstrate the feasibility of using properties extracted from the raw data file to predict a sample metric value via machine learning, a selection of data was obtained using three LC-MS systems (OE 480 MS and Ultimate 3000) named M01, M02, and M03. The systems were operated under controlled and defined conditions using the same LC-MS procedures.

[0161] All LC-MS setups were operated 24 hours a day, seven days a week. The sequences consisted of plasma samples to contaminate the LC-MS setup, while HeLa samples were measured as QC runs to demonstrate how long such a setup can be operated before cleaning of the LC-MS system is required.

[0162] Each LC-MS setup was operated at least once before cleaning / maintenance of the MS was required. This time period is referred to as a "sequence".

[0163] The first of the three LC-MS instrument setups (M01) was used to perform four data acquisition sequences (M01a, M01b, M01c and M01d).

[0164] In each of the four experiments described below, the machine learning algorithm was trained using data from the first three of the four sequences from M01: M01a, M01b and M01c.

[0165] In an initial experiment, data from the first three sequences of the first MS (M01a-M01c) were used to train a machine learning algorithm, which was then validated with data from the fourth sequence of the first MS (MOld). For each QC run in the fourth dataset, a variety of features were extracted from the raw data and provided as inputs to the ML algorithm. These were used to predict the number of protein groups for each QC run in the fourth dataset. The results were compared with the number of protein groups determined by analyzing the raw data using the Proteome Discoverer software.

[0166] In other words, for the first attempt, the training dataset consists of three sequences of M01 (namely: M01a, M01b, M01c). The data obtained during the last sequence (MOld) were used as a test dataset to evaluate how well the trained machine learning algorithm can predict protein groups from the same LC-MS setup by using only properties extracted from the raw data (rather than processing the entire raw data).

[0167] Details regarding the length of each sequence and the corresponding number of protein groups identified by Proteome Discoverer are available in Fig.Figures 11A to 11D illustrate this. The number of protein groups is normalized to a relative scale. The data (M01a - MOld) were obtained when the four different sequences were run on the same LC-MS setup (M01) using the same LC-MS procedure, the same MS, and the same LC. M01a, M01b, and M01c were used as training datasets, while M01d was used as a validation dataset to predict the number of protein groups using machine learning. These predictions are compared to the number of protein groups obtained by processing the validation data with the Proteome Discoverer software.

[0168] To demonstrate the effectiveness of the proposed methods, which use machine learning to predict the number of protein groups, the number of protein groups identified by Proteome Discoverer is compared with the number of protein groups predicted by the proposed machine learning methods. The machine learning algorithm was trained by feeding it 37 parameters (properties) extracted from the raw data (the raw files). These parameters can be obtained quickly from the raw data (< 5 minutes per raw file) because they represent absolute, relative, or median values ​​from the MS1 and MS2 spectra.

[0169] Table 1 below provides an overview of the parameters (and variations) used to train the machine learning algorithm. parameter MS-Level Additionally grouped by AGC type AGC filling 2 no median AGC fill rate [AGF fill = 1 / AGC fill ALL] 2 Yes Relative value Baseline 2 no median Fill quantity 2 no median intensity 2 no median Ion injection time [s] 1,2 no, yes median ITMS2 theoretically 2 Yes median Ion injection time [s] 1 no median Interference signal 2 no median Number of spectra 1,2 no Absolute value Number of peaks MS 1 1,2 no median precursor m / z [Da] 1 Yes median Signal-to-noise ratio 1 Yes median PrOSA Comp (Proportional Orbitrap Signal Adjustment) 1 no median MS1 / MS2 spectra ratio 1&2 Relative value RawOvFtT (and variations) 2 Yes median Res. Dep. Intens 2 Yes median resolution 2 median TIC (Isolation Window) 1 Yes median Total ion current (and variations) 1,2 Yes median

[0170] Collecting sufficient precursor ions for MS2 fragmentation requires a specific duration that depends on the precursor ion intensity in the MS1 spectrum. If the precursor ion intensity is low, the collection time is long. To limit the time required to collect precursor ions for MS2 fragmentation, a maximum ion injection time (Max-IT) is defined in the LC-MS procedure. If the actual precursor ion collection time is less than the Max-IT, the AGC filter value is 1. If the actual precursor ion collection time reaches the Max-IT, the AGC filter value is less than 1.If the preceding parameters specify that [AGC-Filter = 1], this property is based on data collected in cases where the actual time to collect precursor ions is below the Max-IT (data collected in cases where the actual time to collect precursor ions reaches the Max-IT are excluded). “RawOvFtT MS2” is a non-scaled filling (number of ions) of the Orbitrap FT cell. “Res. Dep. Intens” stands for the resolution-dependent intensity, i.e., the TIC of the high-resolution scan in relation to the TIC of the MS1 pre-scan.

[0171] In the specific example of the first experiment, the characteristic parameters listed above were extracted from the raw data and used to train the machine learning algorithm.

[0172] However, the machine learning algorithm is not limited to using these specific properties as inputs. The ML algorithm can be trained to predict the sample metric value based on any combination of the following: Properties extracted from the raw file data; Parameters from the LC; Parameters from calibration files; and / or Parameters from another ultra-fast database search engine.

[0173] This makes the proposed procedures very flexible and adaptable to different requirements for evaluating QC runs.

[0174] For the MOld validation dataset, 638 separate raw files were obtained (one raw file for each QC run), corresponding to 213 days of LC-MS measurements. For all measurements, the number of protein groups was obtained by analyzing the raw file using the Proteome Discoverer software. These results are used to demonstrate the feasibility of applying machine learning to predict the number of protein groups based solely on the properties extracted from the raw data. Fig. Figure 12 illustrates, for each QC run in the validation dataset, the number of protein groups predicted by extracting the properties from the raw data file and providing these properties to the ML algorithm, as well as the number of protein groups determined by feeding the raw data file to the Proteome Discoverer software. The number of protein groups is normalized to a relative scale.

[0175] To simplify data presentation, it illustrates Fig. 13 the average number of protein groups for each day instead of the number of protein groups determined for each QC run. Fig. Figure 13 therefore shows, for each day on which data were collected in the validation dataset, the average number of protein groups predicted by machine learning compared to the average number of protein groups obtained by the Proteome Discoverer software. The number of protein groups is normalized to a relative scale.

[0176] As in Fig.As shown in Figure 13, the results demonstrate that a sample metric value can be effectively predicted over a complete sequence (the MOld validation dataset) using machine learning and only the provision of raw file-related properties. The predicted sample metric value (number of protein groups) had an absolute error of 117 protein groups.

[0177] It is important that the performance degradation of the LC-MS setup, observed towards the end of the sequence, is accurately predicted by the machine learning algorithm. Therefore, the sample metric value predicted by the machine learning algorithm can be reliably used to determine a quality control result for the LC-MS.

[0178] Further specific points / areas are in Fig. 13 are marked and are explained in more detail below.

[0179] The prediction of protein groups for the first 50 days was quite accurate, with an average discrepancy of -3%.

[0180] On day 47 (label (A) in Fig. 13) The HeLa sample was replaced, which directly affects the prediction of protein groups by machine learning (average change of -6% for day 50 - day 150).

[0181] As the LC-MS setup runs for an increasing period, the performance of the LC-MS setup is negatively affected from day 180 onwards (label (B) in Fig. 13). From this day onwards, the number of protein groups continues to decrease, from a value of approximately 2200 to 800 protein groups at the end of the sequence on day 213.

[0182] These negative effects on LC-MS performance can also be determined by the machine learning algorithm. This is in Fig. 13 with field (C) highlighted.

[0183] On days 154 and 192, some significant outliers / failures in LC-MS performance are observed. These are accurately predicted by the ML algorithm due to a strong correlation with the PD software (label (D) in Fig. 13).

[0184] In a second experiment, data were acquired during a sequence using the second OE480 LC-MS instrument setup (M02). This dataset is designated "M02d2". The data from this sequence were used to test how well a machine learning algorithm trained on data from a first LC-MS setup, M01, could predict the number of protein groups for a different LC-MS setup, M02. Furthermore, this sequence ran for 257 days, which is approximately 40 days longer than the final sequence shown in the first experiment (MOld).

[0185] Overall, the mean absolute error for predicting the entire sequence M02d2 was 126 protein groups. Fig. Figure 14 illustrates the number of protein groups predicted by machine learning and the number of protein groups determined by Proteome Discoverer. Here, as in Fig. 13. Several results from a single day are averaged and shown as a single data point for that day. The number of protein groups is normalized to a relative scale.

[0186] For the first 90 days or so, the forecast was significantly better. From day 92 onwards (in Fig. 14 (labeled (A)) showed LC pressure instabilities associated with the MS, requiring manual intervention (including column replacement). This new column lasted only a few days. On day 112, another column replacement was necessary (in Fig.14 marked (B)). From day 112 to day 134, the predictions made by ML agreed well with the data obtained by PD. From day 134 (in Fig. 14 (labeled (C)) the instrument was moved to another laboratory, which directly affected the number of protein groups identified (the average change is about -8%).

[0187] For the third LC-MS instrument setup (M03), data from only one sequence were recorded. As with the first and second trials, the machine learning algorithm was trained using data obtained during the first three sequences (M01a–M01c) performed on the first instrument. The machine learning algorithm was then used to predict the sample metric value for the third instrument (M03).

[0188] Fig.Figure 15 illustrates the prediction of the number of protein groups by machine learning compared to the number of protein groups obtained by Proteome Discoverer. Here, the results for a single day are aggregated and shown as a single data point for that day. The number of protein groups is normalized to a relative scale.

[0189] Overall, the prediction by the ML algorithm was very precise, with a mean absolute error of 60 protein groups. This small discrepancy between the results predicted by ML and those of Proteome Discoverer is observed on almost every day. It is noteworthy that there is a small discrepancy around day 150.

[0190] In a fourth attempt, the number of protein groups for the M01d validation dataset was predicted by a machine learning algorithm trained on the M01aM01c datasets, supported by results from a fast database search engine.

[0191] As described above, QC runs can also be evaluated using search engines designed for "ultrafast" performance. One such ultrafast search engine, "Morpheus," was also used to analyze the raw files from the M01d validation dataset. Morpheus reports (among other things) the number of protein groups, peptides, and PSMs. These three parameters were used to train the machine learning algorithm in two different ways: a) using only the results obtained by Morpheus and b) using the results obtained from Morpheus together with properties extracted from the raw file.

[0192] The use of Morpheus search results (alone or in combination with properties extracted from the raw file) as input parameters for the ML algorithm was tested with the MOld validation dataset from the first trial. For both training methods, the mean absolute error was very small. This demonstrates in general that an external ultrafast search engine can also be used to predict the number of protein groups by machine learning without performing a database search with Proteome Discoverer. The mean absolute error was: a) Morpheus only: 71 protein groups b) Morpheus + raw file-related properties: 61 protein groups

[0193] This is an improvement of 40 protein groups compared to the first attempt, which only used parameters related to raw files (the mean absolute error of the first attempt was 117 protein groups).

[0194] The number of protein groups for the M01d validation dataset is in Fig. Figure 16 is shown. In these data, the results of a single day are aggregated and shown as a single data point for each day. The number of protein groups is normalized to a relative scale. In the Fig. Figure 16 illustrates the comparison between the number of protein groups predicted by the machine learning algorithm and the number of protein groups obtained by Proteome Discoverer.

[0195] Accordingly, the graphic shows in Fig. 16 three data series illustrated: A) the number of protein groups obtained by Proteome Discoverer B) the number of protein groups predicted by a ML algorithm trained using only Morpheus results C) the number of protein groups predicted by a ML algorithm trained on Morpheus results, combined with raw file-related properties

[0196] While the ML algorithm was trained exclusively with identification results from the ultrafast database search (Morpheus), the training dataset for the ML algorithm contained only the number of protein groups, individual peptides and PSMs from Morpheus.

[0197] Since the machine learning algorithm was trained solely on the results of the ultrafast database search, the input for predicting the protein groups from PD is also based on the results of the ultrafast database search engine. Therefore, the parameters from the ultrafast search engine are used to predict the number of protein groups (and compared to the number of protein groups reported by Proteome Discoverer). In this example, no variety of properties extracted from the raw file are used as inputs for the machine learning algorithm (neither during training nor prediction).

[0198] Overall, the prediction was slightly better than using only properties extracted from the raw file. This is not surprising, as both search engines preprocess MS2 spectra to map theoretical peptides to this spectrum. Furthermore, three identification results from Morpheus are used to predict a result from Proteome Discoverer (the number of protein groups).

[0199] As described above, both Morpheus and Proteome Discoverer are search engines that map MS2 spectra to peptides. The properties extracted from the raw data differ from this. They are absolute or mean values ​​of very specific and selected parameters, where only very small changes are expected if the LC-MS setup performs as expected. Therefore, it is advantageous to train the machine learning algorithm with both types of data. • Results from another search engine (Morpheus) that are data of the same type as the output sample metric value; and • Properties extracted from the raw file that contain orthogonal data.

[0200] In this example, the results of the ultrafast database search are combined with the properties extracted from the raw file to train the machine learning algorithm and serve as inputs for the machine learning algorithm for prediction.

[0201] The ultrafast database search engine performs similar operations to the Proteome Discoverer software. However, Proteome Discoverer uses significantly more sophisticated approaches to identify proteins in raw data than the ultrafast database search. Because these sophisticated approaches require more processing time, the overall processing time of a raw file by Proteome Discoverer is longer, but it contains more useful information for the user, such as quantification values.

[0202] Generally, the sample metric provided by Proteome Discoverer is considered to be closest to the true value (since this is the most sophisticated method for deriving the value). Parameters derived via an ultrafast database search can provide a good approximation. However, these parameters may not be accurate enough to determine a QC result. In particular, the combination of ultrafast database search parameters and properties extracted from raw files is more powerful than using an ultrafast search engine alone.

[0203] This combination of parameters for training the machine learning algorithm has experimentally shown the best performance with an average absolute mean of 61 protein groups.

[0204] As demonstrated by the two different approaches to training the ML algorithm in the fourth attempt above, various input parameters can be used to train the ML algorithm. As experts will recognize, different combinations or subcombinations of the properties and parameters mentioned in this application can be used as inputs for the ML algorithm.

[0205] This applies to the first 50 QC runs, which were difficult to predict using only Morpheus-related parameters for the training dataset (see Fig. 16, Label (A)).

[0206] The period from day 50 to 150 was difficult to predict in the first attempt due to changes in the HeLa samples. By combining Morpheus with the parameters associated with the raw file and using this combination to train the ML algorithm, the agreement between the sample metric predicted by the ML algorithm and the sample metric obtained via Protein Discoverer is very good.

[0207] For the evaluation of QC runs, it is important to monitor whether (or when) an LC-MS system is affected. These negative effects begin on day 180 for MOld, with the number of identified protein groups tending to decrease. This decrease is also observed in the predicted number of protein groups (illustrated with label (B) in Fig.16) Towards the end of the sequence (around day 190), when the LC-MS setup was significantly affected by a lower number of identifiable protein groups (1500 or less compared to a starting value of 2200), the values ​​predicted by an ML algorithm trained on a combination of raw file-related properties and the ultrafast search engine identifications appear to be somewhat less accurate.

[0208] In several other examples, the predictions of the number of protein groups generated by the machine learning algorithm are combined (or compared) with the results of an ultrafast database search. This increases the level of certainty in the data predictions. The prediction can also be accompanied by classical QC data evaluation processes, such as the effective gradient length, FWHM, and retention time stability of added peptides (e.g., Thermo PRTC) or the sensitivity (e.g., Thermo System Suitability Standard). All of these evaluation results could be incorporated into a final assessment.

[0209] When this document refers to the determination or prediction of “a number of protein groups”, this is synonymous with the determination or prediction of the parameter “total number of (identified) protein groups” from the Proteome Discoverer software.

[0210] The unit Dalton (Da) is used in this description as an alternative name for the unit atomic mass. Hydrogen has a mass of 1 Da. A hydrogen ion is seen at m / z = 1. The m / z ratio of hydrogen is 1.

[0211] In the description of the present invention, it should be clear that a word used in the singular also includes its plural counterpart, and vice versa, unless otherwise implicitly or explicitly stated. Furthermore, it should be clear that for any given component or embodiment described herein, any of the possible candidates or alternatives listed for that component may generally be used individually or in combination, unless otherwise implicitly or explicitly stated. It should also be understood that the figures shown herein are not necessarily to scale, and some elements may be included for illustrative purposes only.Reference symbols may also be repeated within the various figures to represent corresponding or analogous elements. Furthermore, it should be understood that any list of such candidates or alternatives serves only for illustration and does not constitute a limitation unless otherwise implicitly or explicitly stated.

[0212] Unless otherwise defined, all technical and scientific terms used herein have the meanings generally understood by a person skilled in the art in the field to which this invention relates. In case of conflict, the present specification, including definitions, shall prevail. It is understood that the quantitative terms mentioned in this description are preceded by an implicit "approximately," so that minor and immaterial deviations are within the scope of the present teachings. In this application, the use of the singular includes the plural unless expressly stated otherwise. The use of "comprise," "includes," "comprehensive," "Contains," "contains," "containing," "includes," "includes," and "inclusive" are not to be understood as limitations. As used here, "a" or "an" can also refer to "at least one" or "one or more." Furthermore, the use of "or" is inclusive, such that the expression "A or B" is true if "A" is true, "B" is true, or both "A" and "B" are true.

[0213] The term "scan" as a noun, as used in this document, refers to a mass spectrum, regardless of the type of mass analyzer used to generate and acquire the mass spectrum. When the term "scan" is used here as a verb, it refers to the generation and acquisition of a mass spectrum by a mass analysis procedure, regardless of the type of mass analyzer or mass analysis used to generate and acquire the mass spectrum. The term "full scan" as used herein refers to a mass spectrum that encompasses a range of mass-to-charge (m / z) values ​​containing a variety of mass spectral peaks.

[0214] As used in this document, the respective terms “liquid chromatograph” and “liquid chromatography” (both abbreviated “LC”) and the term “liquid chromatography-mass spectrometry” (abbreviated either “LCMS” or “LC-MS”) refer to any type of liquid separation system capable of separating a liquid sample containing multiple analytes into different “fractions” or “separations”, wherein the chemical composition of each of these “fractions” or “separations” differs from the chemical composition of each of these other fractions or separations, and the term “chemical composition” refers to the number, concentrations and / or identities of the different analytes in a fraction or separation.Therefore, the terms “liquid chromatograph”, “liquid chromatography”, “liquid chromatography-mass spectrometry”, “LC”, “LCMS” and “LC-MS” should, without limitation, include and refer to liquid chromatographs, high-performance liquid chromatographs, ultra-high-performance liquid chromatographs, size exclusion chromatographs and capillary electrophoresis devices.

[0215] Instead of the LC apparatus, any other separation device, including an ion mobility device, HPLC, GC, or ion chromatography, could be connected to the mass spectrometer. Furthermore, any known fragmentation method (including collision-activated dissociation, photon-induced dissociation, electron capture, or electron transfer dissociation) produces data suitable for use with the invention. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] WO 2012 / 160001

[0003] EP 3410463

[0009] Cited non-patent literature

[0000] Eng JK, McCormack AL and Yates JR 3rd, “An Approach to Correlate Tandem Mass Spectral Data of Peptides with Amino Acid Sequences in a Protein Database,” J. Am. Soc. Mass. Spectrom., 1994, 5: 976-989

[0017]

Claims

[1] Method for predicting one or more sample metric values ​​by machine learning, wherein the method comprises: Determining a variety of properties from raw data obtained by analyzing a proteomics sample in a liquid chromatography-mass spectrometer (LC-MS); and Prediction of one or more sample metric values ​​by processing the multitude of properties using one or more trained machine learning models. [2] Method according to claim 1, wherein the machine learning model is trained using data obtained from a plurality of LC-MS combinations and / or from one or more LC-MS instruments exhibiting a plurality of different LC-MS performance states. [3] Method according to claim 1 or claim 2, wherein a data size of the plurality of properties is at least 10 times smaller than a data size of the raw data, preferably 100 times smaller, more preferably 1000 times smaller. [4] Method according to any of the preceding claims, wherein the machine learning algorithm only processes the plurality of properties and does not directly process the raw data. [5] Method according to one of the preceding claims, wherein the time required to determine the plurality of properties from the raw data is shorter than the time required to obtain the raw data by analyzing the QC sample, preferably at least ten times shorter. [6] Method according to any of the preceding claims, wherein analyzing the proteomics sample in LC-MS comprises performing an LC-MS run, wherein the raw data comprise data relating to a retention time during the LC-MS run, wherein the method further comprises: Dividing the retention time into a large number of retention time slices; and Manipulating the raw data to provide data for each of the multitude of retention time slices, where the multitude of properties includes a multitude of properties for each retention time slice. [7] Method for training a machine learning model to predict a sample metric value, wherein the method comprises: Determining a variety of properties from raw training data obtained by analyzing a training sample in an LC-MS; Processing the raw training data via a database search to determine the sample metric value for the training sample; and Training the machine learning model using the multitude of properties and the sample metric value for the training sample, which was determined via database search, so that the machine learning model is trained to predict the sample metric value for a subsequent proteomics sample, based on a multitude of properties determined from raw data obtained by analyzing the subsequent proteomics sample in an LC-MS. [8] The method of claim 7, further comprising training the machine learning model using one or more of the following functions: Features from the LC, Functions from one or more calibration files and / or Features from another database search engine. [9] Method according to claim 7 or claim 8, further comprising validating the machine learning model by: Determining a variety of properties from raw validation data obtained by analyzing a validation sample in an LC-MS; Predicting a sample metric value for the validation sample by processing the multitude of properties using the machine learning model; Processing the raw validation data via a database search to determine the sample metric value for the validation sample; and Comparing the sample metric value determined by processing the raw validation data via database search with the sample metric value predicted by processing the multitude of properties using the machine learning model. [10] Method for determining the data quality of raw data obtained by analyzing a proteomics sample in a liquid chromatography-mass spectrometer, LC-MS, the method comprising: Predictions of one or more sample metric values ​​using the method according to any one of claims 1 to 6; Based on one or more sample metric values, determine whether the raw data quality is above a threshold data quality; If the raw data quality exceeds a threshold data quality, perform further data analysis on the raw data; and If the data quality of the raw data does not exceed the threshold data quality, further data analysis for the raw data is skipped. [11] Method for determining a quality control result, QC result, for a liquid chromatography-mass spectrometry process, LC-MS process, wherein the method comprises: in a QC run, prediction of one or more sample metric values ​​by analyzing a proteomic sample in a liquid chromatography-mass spectrometer, LC-MS, using the method according to any one of claims 1 to 6, wherein the proteomic sample is a quality control sample, QC sample, wherein the one or more sample metric values ​​comprise a QC sample metric value; and Based on the QC sample metric value, determine whether the QC result is "pass" or "fail". [12] Method for performing a liquid chromatography-mass spectrometry sequence, wherein the sequence comprises one or more cycles, each cycle comprising the following steps: Performing a quality control run, QC run, and determining a QC result using the method of claim 11, If the QC result is "fail", end the sequence without further cycles; and If the QC result is "pass", perform one or more sample analysis runs and proceed to the next cycle. [13] Liquid chromatography-mass spectrometer, LC-MS, configured to carry out the method according to any of the preceding claims. [14] Computer software comprising instructions which, when executed by the processor of a computer, cause the computer to perform the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Hybrid mass spectrometer

    EP3410463A1

  • Method and apparatus for mass analysis

    WO2012160001A1