Methods, media, and systems for comparing data within and between groups

By receiving a list of target molecules and utilizing statistical analysis and shared library technology, the difficulty of comparing mass spectrometry data across laboratories and equipment is resolved, enabling more efficient data processing and identification, and improving data consistency and accuracy.

CN115380212BActive Publication Date: 2025-09-19WATERS TECH IRELAND LIMITED IE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180030485.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-24
Filing Date
2021-04-23
Publication Date
2025-09-19
Estimated Expiration
2041-04-23

AI Technical Summary

Technical Problem

The mass spectrometry data generated by existing mass spectrometry equipment is difficult to compare across laboratories and equipment due to differences in equipment characteristics, laboratory conditions and user operations. Existing software is limited by sample type, instrument platform and acquisition method, making it difficult to effectively process and analyze data.

Method used

By receiving a list of target molecules to be identified, using statistical analysis and multi-parameter selection, combined with data processing of the mass spectrometry device, a score is generated to represent the identification possibility of the target molecule, and through common library and normalization technology, data comparison and verification across devices and laboratories are achieved.

Benefits of technology

It enables data consistency comparison between different mass spectrometry devices and laboratories, improves the recognition rate and reduces the false discovery rate, and enhances the versatility and accuracy of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115380212B_ABST
    Figure CN115380212B_ABST
Patent Text Reader

Abstract

Exemplary embodiments provide methods, media, and systems for analyzing spectral and / or chromatographic data, and particularly relate to techniques for improving the reproducibility of the results of spectral and / or chromatographic experiments. For example, some embodiments provide techniques for normalizing mass spectrometry (MS) and / or liquid chromatography (LC) data across different experimental equipment, thereby allowing direct comparison of data from different groups. To this end, exemplary embodiments provide reliable, reproducible target libraries that can be used across different platforms, laboratories, and users. One embodiment utilizes statistical techniques to select experimental parameters that are configured to reduce or minimize the chance of misidentification of target molecules. Another embodiment utilizes the law of large numbers to generate composite product ion spectra that can be used across different experiments. The composite product ion spectra allow the generation of a regression curve, wherein the regression curve can be used to normalize the experimental mass spectra.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is related to and claims priority to U.S. Provisional Patent Application No. 63 / 014,858, filed on April 24, 2020, entitled “Universal Workflow: Software to Enable Global Collaboration of Large Cohorts of LC-MS and LC-MS / MS Data.” Background Art

[0003] In a mass spectrometry (MS) or liquid chromatography / mass spectrometry (LC-MS) experiment, a precursor ion breaks down into product ions. The precursor and product ions are then analyzed in an attempt to identify them. The result of an MS or LC-MS experiment is typically a mass spectrum (a plot of a molecule's intensity versus its mass-to-charge ratio, as measured by an ion detector).

[0004] The problem is that mass spectra generated by MS or LC-MS instruments vary based on instrument characteristics such as age, calibration settings, and environmental conditions. These characteristics can also vary from laboratory to laboratory, user to user, and even experiment to experiment. Therefore, comparing data across different acquisition events (even when generated on a single instrument) can be a daunting task. Due to this challenge, existing MS / LC-MS data processing software is often limited by sample type, instrument platform, and acquisition method. Summary of the Invention

[0005] The following is a brief overview of exemplary embodiments. These embodiments can be embodied as methods, instructions stored on non-transitory computer-readable storage media, devices for performing the actions, etc. Unless otherwise specified, it is expected that the embodiments can be used alone or in any combination to achieve synergistic effects.

[0006] According to a first embodiment, a computing device can receive a list of two or more target molecules to be identified in a sample via a user interface, and analyze the list by a mass spectrometer according to an experimental method defined by a plurality of parameters. The device can perform a statistical analysis based on the list of the two or more target molecules, which is configured to determine the probability of misidentifying one or more target molecules in the target molecule in the list. The device can select a set of values ​​for a plurality of parameters, which reduces the probability to below a predetermined threshold, and can present the set of selected values ​​on the user interface.

[0007] According to a second embodiment, which may be used together with the first embodiment, the statistical analysis may be performed based on a subset of the target molecules in the list, the subset representing a relatively small set of common markers from among the target molecules.

[0008] According to a third embodiment, which can be used with any of the first to second embodiments, the plurality of parameters of the experimental method can include elution position, collision cross section, drift position, and / or mass accuracy.

[0009] According to a fourth embodiment, which may be used with any of the first to third embodiments, values ​​for multiple parameters may be selected based on a known three-dimensional positioning of one of the target molecules, a known fragmentation pattern of one of the target molecules, and / or a known relationship between two of the target molecules.

[0010] According to a fifth embodiment, which may be used with any of the first to fourth embodiments, a computing device may present a score on a user interface to represent the likelihood that the presence or absence of two or more target molecules will be correctly identified in a sample given a set of selected values.

[0011] According to a sixth embodiment that can be used together with the fifth embodiment, the score can be calculated by: (a) querying a common library of immutable attributes including spectral features; (b) counting the frequency of presence of one of the target molecules in the common library; (c) converting the frequency count into a probability score; repeating steps (a) to (c) for multiple target molecules; and multiplying the probability scores for the multiple target molecules together.

[0012] According to a seventh embodiment, which can be used together with the fifth or sixth embodiment, the score can be configured to increase as a function of the number of available separation dimensions determined by the applied acquisition method.

[0013] According to an eighth embodiment, which can be used together with the fifth to seventh embodiments, the score may be configured to increase according to the resolving power of the mass spectrometry device.

[0014] According to a ninth embodiment, which can be used together with the fifth to eighth embodiments, the score can be calculated based on a combination of mass-to-charge ratios, retention times, and drift times of two or more target molecules.

[0015] According to a tenth embodiment, which may be used together with the fifth to ninth embodiments, a score may be calculated based on fluctuations in measurements of a spectral standard across a plurality of spectral devices.

[0016] According to an eleventh embodiment, which can be used separately or in conjunction with any of the first through tenth embodiments, a computing device can apply an acquisition method to receive a mass spectrum generated by a mass spectrometer that determines the number of separation dimensions available in the mass spectrum. The computing device can define a putative product ion spectrum for the mass spectrum and can access a repository storing composite product ion spectra that match the putative product ion spectrum. For each of the separation dimensions, the computing device can retrieve a regression curve generated based on the composite product ion spectrum and can generate a normalized mass spectrum by applying the corresponding regression curve to normalize the values ​​of the corresponding separation dimension.

[0017] According to a twelfth embodiment, which may be used together with the eleventh embodiment, a corresponding regression curve may be applied to raw peak detections in the mass spectrum to correct for changes in at least one of time, mass-to-charge ratio, or drift.

[0018] According to a thirteenth embodiment, which can be used with any of the eleventh to twelfth embodiments, the mass spectrum can be a first mass spectrum, and the mass spectrometry device can be a first mass spectrometry device. The computing device can further receive a second mass spectrum generated by a second mass spectrometry device different from the first mass spectrometry device, normalize the second mass spectrum using the composite product ion spectrum, and verify that the second mass spectrum reproduces the first mass spectrum by comparing the normalized second mass spectrum with the first normalized mass spectrum.

[0019] According to a fourteenth embodiment, which can be used with any of the eleventh to thirteenth embodiments, the first mass spectrometry device may be an instrument platform of a different type than the first mass spectrometry device.

[0020] According to a fifteenth embodiment, which can be used with any of the eleventh to fourteenth embodiments, the first mass spectrometry device may be an instrument platform of a different type than the first mass spectrometry device.

[0021] According to a sixteenth embodiment, which may be used with any of the eleventh to fifteenth embodiments, putative product ion spectra may be clustered into ion clusters, and the computing device may re-cluster the putative product ion spectra based on normalization values.

[0022] According to a seventeenth embodiment, which can be used together with the sixteenth embodiment, putative production ion spectra can be clustered by: calculating a theoretical isotope distribution for putative product ion spectra; clustering the putative product ion spectra based on the theoretical isotope distribution; determining that the intensity of an isotope present in the mass spectrum exceeds a predetermined threshold amount; forming virtual ions from the excess isotope; and clustering the virtual ions into new isotope groups.

[0023] According to an eighteenth embodiment, which can be used together with any one of the eleventh to seventeenth embodiments, a computing device can receive a list of two or more target molecules to be identified in a sample for which a mass spectrum is generated. The computing device can further select a target library customized for the sample, the target library comprising a set of precursor ions and product ions for identifying the two or more target molecules, wherein the target library is composed of a subset of the precursor ions and product ions represented in the mass spectrum. The computing device can use the target library to determine whether the target molecule is present in the sample.

[0024] According to a nineteenth embodiment, which can be used with any of the eleventh to eighteenth embodiments, the subset of precursor ions and product ions can be the minimum subset necessary to identify the target molecule in the mass spectrum.

[0025] According to a twentieth embodiment, which may be used with any of the eleventh to nineteenth embodiments, the normalized values ​​for the separation dimensions may each be associated with a matching tolerance that defines a window through which the normalized mass spectrum will be considered to match the target spectrum. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To easily identify the discussion of any particular element or act, the single most significant digit or digits in a reference numeral refer to the figure numeral that first introduces that element.

[0027] Figure 1 An example of a mass spectrometry system according to an exemplary embodiment is shown.

[0028] Figure 2 An example of a workflow for operating on acquired data according to one embodiment is shown.

[0029] Figure 3 A data flow diagram according to one embodiment is shown.

[0030] Figure 4 An exemplary target discovery and verification environment is shown according to one embodiment.

[0031] Figure 5 An exemplary target verification environment is shown according to one embodiment.

[0032] Figure 6 A flow diagram of exemplary value selection logic 600 is shown, according to one embodiment.

[0033] Figure 7 A flow chart of exemplary spectral normalization and validation logic 700 is shown, according to one embodiment.

[0034] Figure 8An exemplary computer system architecture is shown that may be used to operate the exemplary embodiments described herein. DETAILED DESCRIPTION

[0035] Traditionally, the lifecycle of a cohort of liquid chromatography-mass spectrometry (LC-MS) experiments consists of processing, interpretation, and conclusion drawing, followed by storage of the information on offline storage media. Data processing software is often limited by sample type, instrument platform, and acquisition method. Comparing and contrasting similar datasets within and between laboratories using different sample preparation methods, instrument platforms, acquisition methods, and gradients is challenging.

[0036] However, regardless of sample preparation or the instrument platform and collection method used, molecular mass and hydrophobicity are immutable properties. Furthermore, when the compared instruments employ similar fragmentation mechanisms (e.g., collision cells) and are operated according to the manufacturer's instructions, the fragmentation patterns of the molecules are very similar across the instruments.

[0037] By leveraging these insights, exemplary embodiments provide methods, media, and systems that can compare spectral data acquired within and across cohorts, across multiple laboratories and / or across multiple experimental instruments of different types; embodiments are agnostic to sample type, instrument platform, and acquisition method. These embodiments can increase identification rates while reducing false discovery rates.

[0038] To achieve these improvements, exemplary embodiments can detect ions, create spectra, perform cross-sample and cross-group clustering, perform spectral validation, search multiple databases, extract matching precursor and product ions, create offline and online putative libraries from matching and non-matching spectra, validate identified spectra, and use validated molecular ion spectra to screen normalized peak detection for all target compounds. Using the principles of successive approximations, iterative analysis, and the law of large numbers, exemplary embodiments convert data from multiple instrument platforms into a stream of validated target molecular ion spectra.

[0039] In order to facilitate understanding, a series of examples will be provided before describing the specific implementation of the following embodiment. It should be noted that these examples are only for illustration and the present invention is not limited to the embodiments shown.

[0040] Reference is now made to the accompanying drawings, in which similar reference numerals are used throughout to refer to similar elements. In the following detailed description, for the purpose of explanation, many specific details are set forth to provide a thorough understanding thereof. However, the novel embodiment can operate without these specific details. In other cases, well-known structures and devices are shown in block diagram form to facilitate their description. It is intended to cover all modifications, equivalents, and alternatives that conform to the claimed subject matter.

[0041] In the drawings and accompanying detailed description, the designations "a," "b," and "c" (and similar designations) are variables representing any positive integer. Thus, for example, if an embodiment sets a value of 5, the complete set of components 122 shown as components 122-1 through 122-a may include components 122-1, 122-2, 122-3, 122-4, and 122-5. The embodiments are not limited in this context.

[0042] For illustration purposes, Figure 1 is a schematic diagram of a system that can be used in conjunction with the techniques herein. Figure 1 A specific type of device is depicted in a specific LCMS configuration, but one of ordinary skill in the art will understand that different types of chromatography devices (eg, MS, tandem MS, etc.) may also be used in conjunction with the present disclosure.

[0043] A sample 102 is injected into a liquid chromatograph 104 via an injector 106. A pump 108 pumps the sample through a column 110 to separate the mixture into its components according to their retention time across the column.

[0044] The output of the column is input into a mass spectrometer 112 for analysis. Initially, the sample is desolvated and ionized by a desolvation / ionization device 114. Desolvation can be any desolvation technique, including, for example, a heater, a gas, a heater combined with a gas or other desolvation technique. Ionization can be performed using any ionization technique, including, for example, electrospray ionization (ESI), atmospheric pressure chemical ionization (APCI), matrix-assisted laser desorption (MALDI), or other ionization techniques. By applying a voltage gradient to an ion guide 116, the ions resulting from the ionization are fed to a collision cell 118. The collision cell 118 can be used to deliver ions (low energy) or fragment ions (high energy).

[0045] Various techniques can be used, including those described in U.S. Pat. No. 6,717,130 to Bateman et al., which is incorporated herein by reference, in which an alternating voltage can be applied across the collision cell 118 to induce fragmentation. Spectra are collected of the precursor (no collision) at low energy and the fragments (products of collision) at high energy.

[0046] The output of the collision cell 118 is input to a mass analyzer 120. The mass analyzer 120 can be any mass analyzer, including a quadrupole, a time-of-flight (TOF), an ion trap, a magnetic sector mass analyzer, and combinations thereof. A detector 122 detects ions emitted from the mass analyzer 122. The detector 122 can be integral to the mass analyzer 120. For example, in the case of a TOF mass analyzer, the detector 122 can be a microchannel plate detector that counts ion intensities (i.e., counts the ions that strike it).

[0047] Raw data repository 124 can provide permanent storage for storing ion counts for analysis. For example, raw data repository 124 can be an internal or external computer data storage device, such as a disk, flash memory storage, etc. Analysis device 126 analyzes the stored data. Data can also be analyzed in real time without being stored in storage medium 124. In real-time analysis, detector 122 passes the data to be analyzed directly to computer 126 without first storing it in permanent storage.

[0048] Collision cell 118 carries out the fragmentation of precursor ion. Fragmentation can be used to determine the primary sequence of peptide and subsequently identify the origin protein. Collision cell 118 comprises gas, such as helium, argon, nitrogen, air or methane. When charged precursor interacts with gas atoms, the collision of resulting can make precursor fragmentation by decomposing precursor into the fragmentation ion of resulting. Such fragmentation can use the technology described in Bateman, by the voltage in the collision cell switching between low voltage state (for example, low energy, <5V) and high voltage state (for example, high energy or rising energy, >15V) to realize, wherein the low voltage state is used to obtain the MS spectrum of peptide precursor, and the high voltage state is used to obtain the MS spectrum of the collision-induced fragmentation of precursor. High voltage and low voltage can be referred to as high energy and low energy, because high voltage or low voltage is used respectively to give kinetic energy to ion.

[0049] Various protocols can be used to determine when and how to switch the voltage for this type of MS / MS acquisition. For example, conventional methods trigger the voltage in a targeted or data-dependent mode (data-dependent analysis, or DDA). These methods also include coupled gas-phase isolation (or preselection) of the target precursor. Low-energy spectra are acquired and reviewed in real time by the software. When the desired mass reaches a specified intensity value in the low-energy spectrum, the voltage in the collision cell is switched to a high-energy state. High-energy spectra are then acquired for the preselected precursor ion. These spectra contain fragments of the precursor peptide seen at low energy. After sufficient high-energy spectra have been collected, data acquisition is returned to low energy to continue searching for a precursor mass of suitable intensity for high-energy collision analysis.

[0050] Different suitable methods can be used together with the system as described herein to obtain ion information, such as precursor ions and product ions in conjunction with the mass spectrum for analyzing the sample. Although traditional switching techniques can be adopted, embodiments can use the technology described in Bateman, which can be characterized as a fragmentation procedure that switches voltage with a simple alternating cycle. This switching is completed at a sufficiently high frequency to accommodate multiple high-energy spectra and multiple low-energy spectra in a single chromatographic peak. Unlike traditional switching procedures, the cycle is independent of the content of the data. Such switching techniques described in Bateman provide effective simultaneous mass analysis of both precursor ions and product ions. In Bateman, the use of high-energy and low-energy switching procedures can be applied as a part of LC / MS analysis of a single injection of a peptide mixture. In the data collected from a single injection or experimental run, the low-energy spectrum includes ions primarily from unfragmented precursors, while the high-energy spectrum includes ions primarily from fragmented precursors. For example, a portion of the precursor ions can be fragmented to form product ions, and the precursor ions and product ions can be analyzed substantially simultaneously, or the fragmentation can be regulated by rapidly switching or alternating the voltage of the collision cell of the MS module between a low voltage (e.g., to generate primarily precursors) and a high or elevated voltage (e.g., to generate primarily fragments) either simultaneously or, for example, continuously. In accordance with the Bateman technique described above, MS operations employing rapid alternations of alternating between high (or elevated) energy and low energy may also be referred to herein as the Bateman technique and the high-low protocol.

[0051] The data collected by the high-low protocol allows the retention time, mass-to-charge ratio and intensity of all ions collected under both low-energy mode and high-energy mode to be accurately determined. Generally speaking, different ions are seen in two different modes, and the spectrum collected under each mode can be further analyzed individually or in combination. As seen in one or both modes, ions from a common precursor will have essentially the same retention time (and therefore essentially the same scan time) and peak shape. The high-low protocol allows for meaningful comparison of the different features of ions within a single mode and between modes. This comparison can then be used to group the ions seen in the low-energy spectrum and the high-energy spectrum.

[0052] In summary, a sample 102 is injected into an LC / MS system, such as when the Bateman technique is used to operate the system. The LC / MS system generates two sets of spectra: a set of low-energy spectra and a set of high-energy spectra. The set of low-energy spectra includes the primary ions associated with the precursor. The set of high-energy spectra includes the primary ions associated with the fragmentation. These spectra are stored in a raw data repository 124. After data acquisition, these spectra can be extracted from the raw data repository 124 and displayed and processed by a post-acquisition algorithm in an analysis device 126.

[0053] Metadata describing various parameters related to data acquisition can be generated along with the raw data. This information can include the configuration of the liquid chromatograph 104 or mass spectrometer 112 (or other chromatographic device used to acquire the data), which can define the data type. The identification of the codec configured to decode the data (e.g., a key) can also be stored as part of the metadata and / or with the raw data. The metadata can be stored in a metadata catalog 130 within the document repository 128.

[0054] Analysis device 126 can provide analysts with data visualization at each workflow step based on workflow operations and allow analysts to generate output data by performing processing specific to the workflow steps. Workflows can be generated and retrieved via client browser 132. As analysis device 126 executes workflow steps, it can read raw data from the data stream located in raw data repository 124. As analysis device 126 executes workflow steps, it can generate processed data that is stored in metadata catalog 130 in document repository 128; alternatively or in addition, the processed data can be stored in a different location specified by the user of analysis device 126. It can also generate audit records that can be stored in audit log 134.

[0055] The example embodiments described herein may be executed at multiple locations, such as client browser 132 and analysis device 126 . Figure 8 An embodiment of a device suitable for use as analysis device 126 and / or client browser 132, as well as various data storage devices, is depicted in FIG.

[0056] For context, Figure 2 Describes the Figure 1 FIG2 is a simplified embodiment of a workflow 202 applied by an analysis device 126 of FIG2. The workflow 202 is designed to take a set of inputs 204, apply multiple workflow steps or stages to the inputs to generate outputs at each stage, and continue processing the outputs at subsequent stages to generate experimental results. Note that the workflow 202 is a specific example of a workflow and includes specific stages executed in a specific order. However, the present invention is not limited to Figure 2 Other suitable workflows may have more, fewer, or different stages executed in a different order.

[0057] The initial set of inputs 204 may include a sample set 206 comprising raw (unprocessed) data received from a chromatography experimental setup. This may include measurements or readings (such as mass-to-charge ratios). The measurements initially present in the sample set 206 may be measurements that have not yet been processed, such as to perform peak detection or other analytical techniques. The sample set 206 may comprise data in the form of a stream (e.g., a sequential list of data values ​​received from the experimental setup in a steady, continuous stream).

[0058] In the context of the present application, sample set 206 can represent raw data stored in raw data repository 124 and returned by endpoint interface. Sample set 206 can be represented as a model of a data flow (e.g., including data structures corresponding to data points collected by a chromatographic device). Workflow 202 can be executed on sample set 206 by an application running on analytical device 126 and / or within a data ecosystem.

[0059] The initial input set 204 may also include a processing method 208, which may be a template method (as described above) that is applied to (and thereby embedded in) the workflow 202. The processing method 208 may include settings to be applied to various levels of the workflow 202.

[0060] The initial set of inputs 204 may also include a result set 210. When created, the result set 210 may include information from the sample set 206. In some cases, the sample set 206 may be processed in some initial manner when copied into the result set 210, e.g., the MS data may require extraction and smoothing filtering before being provided to the workflow 202. The processing applied to the initial result set 210 may be determined on a case-by-case basis based on the workflow 202 being used. Once the raw data is copied from the sample set 206 to create the result set 210, the result set 210 may be completely independent of the sample set 206 for the remainder of its usage cycle.

[0061] The workflow 202 can be divided into groups of stages. Each stage can be associated with one or more stage processors that perform the execution calculations associated with the stage. Each stage processor can be associated with a stage setting that affects how the processor generates output from a given input.

[0062] Levels can be separated from each other by step boundaries 238. Step boundaries 238 can represent the point at which output has been generated by a level and stored in a result set, at which point processing can proceed to the next level. Some level boundaries may require specific types of input in order to be crossed (for example, data generated at a given level may need to be reviewed by one or more reviewers, who may need to provide their authorization to cross step boundary 238 to the next level). Step boundaries 238 can apply at any time and in any direction as a user moves from one level to a different level. For example, step boundaries 238 exist when a user moves from initialization level 212 to channel processing level 214, and also when a user attempts to move back from quantification level 222 to accumulation level 216. Step boundaries 238 can be ungated, meaning that once a user confirms they want to move to the next level, no further input is required (or only a cursory input is required), or they can be gated, meaning that the user must provide some confirmation indicating their desire to proceed to the selected level (perhaps in response to an alert from analysis device 126), the reason for moving to the level, or credentials authorizing workflow 202 to proceed to the selected level.

[0063] In the initialization stage 212, each stage processor can respond by clearing the results it generated. For example, the stage processor for the channel processing stage 214 can clear all of its resulting channels and peak tables (see below). At any point in time, the clear stage setting can clear the stage tracking for the current stage and any subsequent stages. In this embodiment, the initialization stage 212 does not generate any output.

[0064] After crossing step boundary 238, processing may proceed to channel processing stage 214. As described above, a chromatogram detector may be associated with one or more channels from which data may be collected. At channel processing stage 214, analysis device 126 may derive a set of processing channels present in the data in result set 210 and may output a list of processed channels 226. The list of processed channels 226 may be stored in a version subdocument associated with channel processing stage 214, which may be included in result set 210.

[0065] After crossing step boundary 238, processing may proceed to accumulation stage 216, which identifies peaks in the data in result set 210 based on the list of processed channels 226. Accumulation stage 216 may identify peaks using techniques specified in the settings of accumulation stage 216, which may be defined in processing method 208. Accumulation stage 216 may output a peak table 228 and store the peak table 228 in a version subdocument associated with accumulation stage 216. The subdocument may be included in result set 210.

[0066] After crossing step boundary 238, processing may proceed to identification stage 218. In this stage, analysis device 126 may identify the components in the mixture analyzed by the chromatographic device based on the information in peak table 228. Identification stage 218 may output component table 230, which includes a list of components present in the mixture. Component table 230 may be stored in a versioned subdocument associated with identification stage 218. The subdocument may be included in result set 210.

[0067] After crossing step boundary 238, processing can enter calibration stage 220. During the chromatography experiment, calibration compounds can be injected into the chromatography device. This process allows the analyst to take into account slight changes in electronic components, surface cleanliness, environmental conditions in the laboratory, etc. throughout the experiment. In calibration stage 220, the data obtained about these calibration compounds is analyzed and used to generate calibration table 232, which allows the analytical equipment 126 to calibrate the data to ensure that it is reliable and reproducible. Calibration table 232 can be stored in a version subdocument associated with calibration stage 220. The subdocument can be included in result set 210.

[0068] After crossing step boundary 238, processing can enter quantification stage 222. Quantification refers to the process of determining the numerical value of the amount of analyte in the sample. Analysis device 126 can use the results from the previous stage to quantify the components included in component table 230. Quantification stage 222 can update 234 component table 230 stored in result set 210 with the quantitative results. The updated component table 230 can be stored in a version subdocument associated with quantification stage 222. The subdocument can be included in result set 210.

[0069] The full width at half height of the two masses is included in the metadata for each ion, and in the case of ion rate mobility separation. These half heights are used to calculate mass and drift resolution. These resolutions are then divided into m / z bins. For each bin, the average resolution is calculated as well as the standard deviation and coefficient of variation for each ion in that bin. This allows the ion detection algorithm to calculate the purity score for each ion. Essentially, this process identifies deconvoluted interferences. This ensures high-precision quantitative measurements. This technique can obtain highly accurate quantitative precursor ion regions as well as highly accurate normalized product ion spectra.

[0070] After crossing step boundary 238, processing may proceed to summary stage 224. In summary stage 224, the results of each of the previous stages may be analyzed and combined into a record of summary results 236. Summary results 236 may be stored in a version subdocument associated with summary stage 224. The subdocument may be included in result set 210.

[0071] As used herein, a step can correspond to a level as described above. Alternatively, a single level can include multiple steps, or multiple levels can be organized into a single step. In any case, all activities performed in a given step should be performed by the same user or user group, and each step is associated with one or more pages of a group of configuration options that describe the step (e.g., visualization options, review options, step configuration settings, etc.).

[0072] There may be transitions at some or all of the step boundaries 238, although not every step boundary 238 requires a transition. A transition may represent a change in data group responsibility from a first user or group of users to a second, different user or group of users.

[0073] Overview of molecular ion reservoirs ( MIR )

[0074] The mass spectrometer acquires and centers ions within a preset acquisition time. The acquisition time is usually a function of the width of the chromatographic peak. A balance is struck between having enough scans across the peak to accurately determine its area and too many scans, which can adversely affect file size. Each scan consists of three numbers, or four numbers if ion mobility separation (IMS) is available. These numbers refer to the scan number, mass-to-charge ratio ( m / z ), intensity, and drift time. Qualitative and quantitative results are derived from a series of algorithmic interpretations of these numbers, including peak detection, deisotoping, charge state reduction, precursor and product ion alignment, database searching, and quantification. Additionally, algorithms exist to model isotope assignments and calculate peak widths (m / z, time, and drift) to correct for interferences.

[0075] Algorithms have an inherent degree of error known as the "95% rule." Here, assuming a 100% "correct" result, the error can be equated to the efficiency percentage. For example, to illustrate the 95% rule, if two consecutive algorithms are applied to data, and each algorithm produces a 95% correct result, the cumulative efficiency is: 0.95 * 0.95 * 100 = 90.25% correct. The more algorithms added to the sequence of processing steps, the greater the error and the lower the specificity. The serial application of multiple algorithms to the same data will reflect cumulative error. The more algorithms applied, the less accurate the results.

[0076] Sample complexity dictates the resolution requirements of the available separation dimensions needed to successfully identify and quantify the maximum number of compounds across the widest dynamic range. If the number of available separation dimensions and their corresponding resolving power is not commensurate with the sample complexity, interferences will increase along with the combined algorithm errors, resulting in significantly impaired results. As an example, there are many isotopic lipids that need to be separated chromatographically, providing the correct gradient. The gradient length has a profound impact on sample throughput. Increasing the gradient slope increases the throughput, which is detrimental to isotopic separation. Isotopic lipids have the same m / z, and if they are not chromatographically separated, the recorded intensities are composite. If IMS is available and the collision cross sections (CCS) of the two isoforms are unique and within the IMS resolution, then increasing the gradient slope, and by extending the sample throughput, will have little effect on the calculated intensity of each isoform. The added dimension of IMS provides a means to increase throughput without compromising the ability to detect and accurately quantify isotopic species.

[0077] Interferents can be quite common, but they can also be handled correctly through replication and validation across independent samples and / or datasets. Consider the case of trypsin and lipids. The 20 naturally occurring amino acids comprise only six elements and are similarly distributed across all proteins. Lysine and arginine are the preferred cleavage sites for trypsin, both present at approximately 6%. If all amino acids were evenly distributed, then digestion with trypsin would, on average, produce a string of 10 residue peptides, each with a similar composition. Molecules of similar length and composition tend to have similar hydrophobicity. Due to co-elution, similar m / z High concentrations of hydrophobic and molecular ions can lead to increased ionic interference.

[0078] The exemplary embodiment creates a list of identified chemical components, with their relative intensities, precursors, and products aligned across the sample set. This is achieved through continuous refinement. Scan spectra are correlated with sample-by-sample complexes. Complexes from each sample are then correlated with consensus spectra. These consensus spectra are used to create libraries, which are then used to screen normalized peak lists.

[0079] The consensus is not generated from a single laboratory or a single sample, but rather from multiple different datasets. Furthermore, the implementation not only creates a target library on identified content, but also builds a consensus spectrum on unidentified content through continuous refinement.

[0080] Furthermore, measurement accuracy is a function of the signal-to-noise ratio. Selecting the apex scan as a reference allows the exemplary embodiment to monitor ion peak shape during elution. m / z The algorithm identifies interferences based on the rate of change and scan-to-scan fluctuations in the ion concentration. In LC-MS or LC-MS / MS analysis, each ion has a small probability of being an interference in each scan.

[0081] In discovery mode (discussed in more detail below), exemplary embodiments can perform identification with as little as one unburdened low and high energy scan. There will always be some impaired product ion spectra; the conversion from product ion spectra to consensus spectra depends largely on n (in nis an integer representing the number of samples. According to the law of large numbers, the average of the product ion spectra (estimated) obtained from a large number of samples should be close to the expected value (composite), and will tend to be even closer as the number of groups increases (common). In practice, the exemplary embodiment utilizes replication to eliminate false positives and aggregate true positives into a common spectrum.

[0082] In some embodiments, a computing device can retrieve publicly available LC-MS and LC-MS / MS data from numerous laboratories, running different instrument platforms and acquisition methods. This data can be used to create and validate a molecular ion repository 318 (MIR; see Figure 3 ). MIR can be separated into:

[0083] • Presumptive Section 320

[0084] • Composite Section 322

[0085] • Common portion 324

[0086] • Target section 326.

[0087] The putative portion 320 includes all matching and non-matching product ion spectra from each sample. The composite portion 322 includes the related scan-by-scan spectra acquired on the same precursor ion during discovery. The common portion 324 contains the related composite spectra that have been reduced to only the most specific product ions or associated features (e.g., lipids of a class or pathway, or peptides of a protein) in the LC-MS dataset. The target portion 326 includes the minimum number of product ions required to correctly identify the common spectrum.

[0088] The molecules residing in the MIRs can be grouped into:

[0089] • Peptide to protein

[0090] • Protein to pathway

[0091] • Lipid to

[0092] • Lipid to pathway

[0093] • Metabolites to pathways

[0094] • Metabolites to drugs.

[0095] The following combination Figure 3 A more detailed explanation of MIR and its operation is provided. Figure 3The data path through an exemplary embodiment is depicted, in which MS data is validated using an exemplary MIR. The data path is divided into a discovery loop 330, an optional database search, a transformation loop 332, and target validation. Note that these subdivisions are primarily organizational, and the various actions performed within each grouping are described below (although other groupings may be used in other embodiments).

[0096] Discovery Ring

[0097] The following section describes an exemplary discovery ring 330 suitable for use with the exemplary embodiments.

[0098] Peak detection logic 302 scans the center ion events one by one in all available separation dimensions. m / z and drift resolution.

[0099] The ion screening logic 304 can be based on the calculated m / z Resolution applies an ion screening filter (ISF) to bin raw peak detections. Each bin m / z Values ​​can be tracked across time. Peaks are selected by finding a local maximum, ensuring minimum scan continuity over time in the neighborhood of the maximum, ensuring a minimum rate of change between scans within that neighborhood, and confirming that a minimum number of scans have a center intensity at least equal to the average of all neighboring scans.

[0100] The ions that pass through the ISF can be provided to z determination logic 306, which performs charge determination and deisotoping. m / z Supreme m / z Classification, assuming the lowest m / z The isotope cluster is A0. m / z Starting and using the same charge determination algorithm as previously described, each ion can be assigned to a charge state.

[0101] Once the charge is assigned, the molecular weight (M r ) can be calculated by adduct and variant logic 308. Some molecules can support multiple charge states, with multiple adducts supporting each charge. r The precursors can be divided into component groups, with the strongest member of the group being labeled the primary member.

[0102] Once the ions have been assembled into isotope groups, clustering logic 310 can apply the average value (the average elemental composition of all amino acids) as the isotope model. Clustering logic 310 can calculate and compare the calculated theoretical isotope distribution with the experimental one. If the experimental intensity of any isotope in the isotope group exceeds 125% of the theoretical value, the algorithm can create a "virtual" ion from the excess. This process is repeated until all virtual ions are clustered into a new isotope group or discarded.

[0103] Once the peaks have been verified and all relevant charge states and adducts have been assembled into component groups, product ions showing the same metadata can be paired with their corresponding precursors on a scan-by-scan basis. Clustering logic 310 can end up with a product ion spectrum for each precursor ion cluster in each scan. The normalized area intensity ratio (AR3) for each aligned product ion to precursor can be calculated:

[0104] AR 3 = Product ion strength / S Product ion strength.

[0105] At the conclusion of the focusing logic 310 , the single scan product ion spectra may be stored directly in the putative portion 320 of the molecular ion repository 318 , sent to the demultiplexing spectrum generation logic 312 , or sent directly to the database search engine application database search logic 314 .

[0106] Optional compound generation and database searching

[0107] A composite product ion spectrum can be generated by summing the product ion spectra for the entire peak using demultiplexed spectrum generation logic 312. Only those product ions with a minimum n / 2+1 The scans show similar rates of change because the precursor can remain as a complex ion, where n is an integer representing the total number of scans. Similarly, the normalized intensity ratio AR 3 Can be calculated for each product ion that passes the minimum match criteria. Can be calculated for all single scan product ions across the peak. AR 3 The standard error is calculated from the values. The demultiplexed spectrum generation logic 312 can also calculate the standard error on the scan-by-scan isotope distribution of both the precursor ions and the product ions that make up the composite spectrum. These statistics provide the algorithm with a means to identify interferents. Product ions with a coefficient of variation less than 30% can be retained and stored in the estimated portion 320 of the molecular ion repository 318, or, if searchable, sent to a database search engine.

[0108] In samples where an optional database search is performed, the database search logic 314 may retain all matching ions and calculate new AR 3 The matched spectra are then deposited into the putative portion 320 of the molecular ion reservoir 318. In data independent acquisition (DIA) acquisition, product ions are often shared with more than one precursor. To this end, the matched product ions are removed from the unmatched spectra and the remaining product ions are used to calculate the new AR 3 ratio. The unmatched spectra are then uploaded to the putative section 320 of the molecular ion repository 318 where they are aggregated and then re-searched while processing continues until there are no new identifications. Removing the matched product ions can increase the probability of each AR 3 The accuracy of the value.

[0109] Transformation Ring

[0110] The transition loop 332 applies the demultiplexed common validation logic 316 to continuously query the putative portion 320 of the molecular ion repository 318 for sequences that exhibit a minimum match count that meets or exceeds a predetermined threshold. m (e.g., 50) product ion spectra. If IMS is used, spectra exceeding the minimum count can be extracted ( Figure 4 Block 402), their fragmentation patterns can be m / z 、 AR 3 Ratio and drift are correlated ( Figure 4 After the initial correlation, the standard error can be calculated based on all matching product ions. m / z value, AR 3 Ratio, retention and drift time calculations ( Figure 4 If the product ion match count reaches or exceeds a predetermined threshold (e.g., 38), and the AR3 standard error reaches or falls below a predetermined value (e.g., 0.35), a composite spectrum is generated ( Figure 4 The metadata associated with each transition composite spectrum includes: theoretical m / z value (used to identify the target); the average of each product ion AR 3 ratio; mean retention time; mean drift time, if IMS is available; coefficient of variation for each product ion; and mean product ion intensity.

[0111] Extracted spectra that failed to transform were restored for future analysis by ( Figure 4When the match rate of the initial failed group of putative spectra increases by 20%, the trigger event will cause a different random m Each product ion spectrum is extracted, and the process is repeated. The transformation loop is continuous, recirculating each time the putative portion of the MIR is refreshed. Depending on the processing rate, this can be done hourly, daily, weekly, or monthly. Molecular properties such as hydrophobicity signatures, isoelectric point, and elemental composition can be extracted and used to update the retention time and CCS prediction algorithms.

[0112] Once a minimum number (i.e., above a predetermined minimum threshold) of composite spectra are retained ( Figure 4 The demultiplexing common verification logic 316 may extract the composite spectrum from the composite portion 322 ( Figure 4 The composite spectrum can be associated with reducing the number of only those product ions that are most specific for the molecule, ie, those that have passed through the transition from the composite portion 322 to the common portion 324 of the molecular ion reservoir 318.

[0113] Molecular ion reservoir ( MIR )

[0114] There are five different product ion spectra and LC / MS signatures in the molecular ion reservoir 318 .

[0115] • Estimates (initial scan-by-scan alignment of spectra),

[0116] • Estimation (correlating scan-by-scan spectra),

[0117] • complexes (correlated putative spectra across samples),

[0118] • shared (complexes related across groups), and

[0119] • Target (minimal highly selective shared spectrum).

[0120] At the end of the discovery loop 330, single scans and composite product ion spectra have two forward routes: first, into the putative portion 320 of the molecular ion repository 318, and then directly into the database search engine. For LC / MS processing, all features can be transferred directly to the molecular ion repository 318. Instead of product ion spectra, pseudo spectra are made from all ions in the scan. As long as the column matrix and buffer compositions are similar, the elution order can be preserved, resulting in a set of linked features similar to the putative product ion spectra. These linked features are then processed similarly to the MS / MS spectra described previously. Once in the molecular ion repository 318, the putative product ion spectra or LC / MS features for each sample are sorted by intensity. In proteomics samples, the identified peptides can be grouped into proteins and sorted in descending order of intensity. The peptide intensities are then summed, and the identified proteins are sorted in descending order of intensity. A similar grouping and sorting process can be applied to small molecules. While peptides are grouped into proteins, lipids are grouped by metabolite class or pathway, grouped by pathway or drug. Thus, regardless of sample type, the molecular ion repository 318 contains all putative product ion spectra or LC / MS features for each sample in order of their intensity.

[0121] Discovery Normalization

[0122] The composite product ion spectrum or LC / MS signature generated from each cohort can be matched to the putative product ion spectrum or LC / MS signature of each individual sample from that cohort. Once matched, the discovery normalization logic can generate a regression curve for each of the available separation dimensions. Each regression is applied to the raw ion detections from the ion screening. m / z , retention time, and drift time can be used to regroup putative product ion spectra or LC / MS features across groups. The discovery normalization logic also calculates matching tolerances for each separation dimension across groups of subsequent target rings.

[0123] Target Verification

[0124] Target validation logic 328 continuously queries the common portion 324 of the molecular ion repository 318 for new common spectra. Target validation logic 328 performs a correlation analysis of the normalized product ion intensities of the matching product ions. The same correlation analysis is performed on the linkage features in the LCMS data. All composite product ion spectra and linkage features that demonstrate a correlation coefficient > 0.7 and provide a pass count rate > n / 2 + 1 are retained, and an initial common spectrum or linkage feature list is generated and stored in the MIR.

[0125] The number of target product ions can be determined by the following precursor properties:

[0126] • Molecular weightM r ,

[0127] • Sorting strength, and

[0128] • Charge z .

[0129] Criteria for product ion selection include:

[0130] • Match rate,

[0131] • Strength

[0132] • Frequency and

[0133] • AR3 changes.

[0134] Points can be assigned to each of the four selection criteria ( Figure 4 412 in (block 412):

[0135] • Matching rate / maximum rate of signature ID,

[0136] • Strength / Maximum Strength of the Feature ID, and

[0137] • Frequency – The number of times the product ion is found in the consensus library.

[0138] The four scores are multiplied and the product ions are sorted in descending order of score ( Figure 4 The number of product ions that can be verified is a function of the mass, intensity, and dissociation kinetics of the molecule. Small molecules are typically singly charged and have low M r。 Furthermore, fragmentation kinetics result in many product ions having intensities much lower than those of the primary fragments. The number of product ions that can be identified is primarily a function of the precursor ion intensity. For example, with trypsin titanium, the number of target product ions can range from a minimum of (e.g.) 5 to a maximum of (e.g.) 10. As the resolution and / or separation dimensionality values ​​increase, the frequency of each product ion decreases as the fraction increases, making frequency the primary factor in product ion selection. Conversely, as the dimensionality and / or resolution values ​​decrease, the matching rate, intensity, and number of associated features have a greater impact.

[0139] Each frequency count can be converted into a probability score ( Figure 4 The probability score reflects the chance of finding each target ion as a random event (e.g., the probability of misidentifying the target ion). Multiplying the probabilities provides the likelihood of finding all target ions and precursors as random events. This likelihood can be converted to a PPM FDR, and if the FDR is less than a predetermined threshold ( Figure 4"PASS" at block 416), the target product ions in the analysis can be selected as a group.

[0140] Once target product ions have been selected for a population, the target spectra can be converted into the target portion 326 of the molecular ion repository 318. Target selection is dynamic in that it can automatically adjust based on changes in sample complexity, resolution, and / or the number of separation dimensions available for the population being analyzed. This process ultimately results in a target list consisting of the most reproducible, lowest frequency, lowest probability product ions or linkage features for each common product ion spectrum present in the molecular ion repository 318.

[0141] Figure 4 The paths of target product ion spectra or linkage feature groups (pathways, classes, proteins) are shown in FIG and have been described in detail above. In short, the successive refinement loops of the exemplary embodiment ultimately form a target library that is uniquely selected for that sample. The target library includes the minimum number of product ions and / or linkage features necessary to identify all target compounds residing in the library in any sample, group, or series of groups with an ultra-low FDR rate.

[0142] Figure 5 It shows how the target portion 326 of the molecular ion reservoir 318 can be applied to incoming data in a target ring.

[0143] After data collection, a normalized peak list 502 can be generated by post-discovery normalization. A unique number of target ions selected for the number of separation dimensions and resolutions applied in the experiment can be input into target identification logic 504 along with the normalized peak list 502. Matching normalized peak detections can be clustered and scored by scan. For example, the normalization and target ions can be assembled into a series of independent grids or cubes. The separation method used is a function of the acquisition method and the number of separation dimensions available. The time, drift, and m / z The width is determined by the discovery normalization logic.

[0144] When pairing grid or cube matching, the target ion must be matched not only to a single scan, but also to a series of scans that define the elution composition. The number of consecutive matching scans must also be commensurate with the intensity of the apex scan. More complex or lower resolution samples may require more dimensions, such as AR 3, Linking features and additional product ions. The number of dimensions required to ensure highly accurate results is a direct result of the required depth of coverage and the resolution of each available separation dimension. As the number of dimensions and / or resolution increases, the probability of a target being misidentified or as a random event decreases. Target rings are calculated for each grid or cube for precursor ions and product ions. m / zThe target precursor ions and product ions in the raw data are m / z The calculated probabilities (frequencies / counts) for the values ​​(+ / - 10 ppm) are multiplied together to give two comparative probabilities of finding all six ions in the exact same scan. The two comparative probabilities are derived from the normalized peak detections for the group and the consensus library. Each ion is constrained by using retention time and / or drift time (CCS) tolerances. m / z The specificity can be further improved by counting the frequency of the values. Therefore, the probability of randomly aligning these ions is m / z , retention time, and drift time.

[0145] In proteomics, once a target peptide is identified, the targeted algorithm generates all other possible product ions. m / z values ​​and screened for matching scans. Additional product ion matches increase selectivity by allowing the targeting algorithm to remove them from the normalized peak detections. As mentioned previously, the mass analyzer is an excellent measure of isotope distribution. Knowing the elemental composition and A0 intensity of each molecule allows the targeting algorithm to correct for any interferences. The fitted region intensities for each isotope of the precursor and product ions are removed. The remaining intensities are used to create virtual ions and are passed through a second round of focusing along with the unmatched peak detections. The newly generated spectrum is investigated through the same discovery loop as before, and the investigation is terminated when no new identifications are found.

[0146] As with other putative identifications, target identifications can be validated by demultiplexing target validation logic 506, correlating using the aforementioned transformation loop 332 to ensure reproducibility across all samples in the cohort. Validated target identifications are then grouped as previously described and transferred for analysis by multivariate statistics / machine learning 508.

[0147] After MIR and target libraries are generated as discussed above, they can be used for various purposes, such as verifying whether mass spectra generated by different devices, different users, different laboratories, etc. are repeatable to each other, normalizing data for future comparisons, determining whether target molecules are present in a sample, and other applications. Figure 6 is a flow chart depicting exemplary logic for one such application: designing an experiment to maximize the chance of finding target molecules if they are present in a sample. For example, given a particular combination of target molecules, the logic may output that the device should be configured with X Minute gradient, Y psi and Z Indication of PPM mass accuracy.

[0148] Among other possibilities, Figure 6The logic blocks may be implemented as instructions stored on a non-transitory computer-readable medium, as methods executed by a computing device, or as a computing device programmed with the instructions to perform the actions.

[0149] At block 602, a computing device may receive, via a user interface, a list of two or more target molecules to be identified in a sample for analysis by a mass spectrometer according to an experimental method defined by a plurality of parameters, such as a gradient slope, a pressure value, and / or a mass accuracy of the experimental method.

[0150] At block 604 , the device may perform a statistical analysis based on the list of two or more target molecules, the statistical analysis configured to determine a probability of misidentifying one or more target molecules in the list.

[0151] The statistical analysis can be based on data collected by running a known standard multiple times on different machines as stored in the MIR. Because different machines will inherently have some differences, the hold-up time between different machines will likely be different.

[0152] Thus, a stream of spectra can be retrieved from each machine, where the spectra include both precursor (LCMS and LCMSMS) and product ions (LCMSMS). Ion detections can be sorted by intensity so that n The most intense ion detection should be at the top of each sample. n The next most intense detections are sorted in ascending order of mass. Optionally, the masses can be binned (e.g., in 10 ppm).

[0153] The computing device determines how many times each mass is found across the population of samples. The goal is to compile the data into n group, which would indicate that one ion detection occurred per sample and that the ion detections could therefore be aligned with each other.

[0154] If the number is greater than n (number of samples), meaning that masses exist in the data at different retention times. The computing device can sort the ion detections by binning the masses and then the retention times. Retention times can also be binned (e.g., within 1 minute, 5 minutes, etc.).

[0155] If the ion detection is when sorting by mass and time n If the sample is a single set, then one detection occurs per sample and can therefore be aligned. For example, if n =15 and 15 detections of a given mass occur within 5 minutes, 15 detections occur within 10 minutes and 15 detections occur within 15 minutes, then one detection occurs per sample and can be aligned.

[0156] At each step, it can be determined whether the data is aligned based on the current set of dimensions (quality, time, drift, etc.). If the number of detections at a given step cannot be arranged into n Groups, another dimension can be added and detections can be sorted based on the new dimension.

[0157] The ion detection is arranged as n After grouping, the change in time for each detection in the detection can be plotted (because the same molecule should appear in the data at the same time point). Fluctuations in drift time can be explained by mapping the differences on the regression line and aligning them. By measuring the change in normalized time, fluctuations can be measured. Fluctuations can then be used to define matching tolerances. The tolerances provided for each of the dimensions can be determined because the original sample is the standard, so the experimenter knows the composition of the sample. For example, if the sample can be matched within X ppm, Y minutes, and Z drift bins, these values ​​can define the parameters of the experimental setup required to detect the identified target molecule with high accuracy.

[0158] In some embodiments, the statistical analysis is performed based on a subset of target molecules in the list, which subset represents a relatively small set of common markers from the target molecules. Because many molecules contain a large number of common markers, finding these common markers may not provide much information about the molecules in the sample (because many different molecules can form the common marker). By selecting a subset of target molecules based on which markers are relatively rare in common, more accurate identification can be performed more quickly.

[0159] At block 606 , the device may select a set of values ​​for the plurality of parameters that reduces the probability below a predetermined threshold and present the selected set of values ​​on a user interface.

[0160] In some embodiments, values ​​for multiple parameters can be selected based on known multidimensional positioning of one or more of the target molecules and / or known relationships between the target molecules.

[0161] At block 608 , the computing device may present a score on a user interface to represent the likelihood that the presence or absence of two or more target molecules will be correctly identified in the sample given the selected set of values.

[0162] Based on the above changes, the computing device can determine the likelihood that the target molecule will be identified for a given parameter set given these parameters. This likelihood can be normalized and converted into the above score.

[0163] In some embodiments, the score can be calculated by: (a) querying a common library of immutable attributes including spectral features; (b) counting the frequency at which one of the target molecules is present in the common library; (c) converting the frequency count to a probability score; repeating steps (a) to (c) for multiple target molecules; and multiplying the probability scores for the multiple target molecules together. The score can be configured to increase as a function of the number of available separation dimensions determined by the acquisition method applied and / or according to the resolving power of all available separation systems employed. The score can be calculated based on a combination of the mass-to-charge ratio, retention time, and drift time of two or more target molecules and / or based on fluctuations in measurements of spectral standards across multiple spectroscopic devices. In other words, the score reflects the frequency at which a given mass is reflected with a given normalized time and drift, which is then converted to a probability. Combining the probabilities across all molecules of interest limits the likelihood that the identified feature will be detected by chance or as a random event.

[0164] Figure 7 is a flow chart depicting exemplary logic for another application of the above-described MIR, which normalizes acquired data for various purposes.

[0165] Among other possibilities, Figure 7 The logic blocks may be implemented as instructions stored on a non-transitory computer-readable medium, as methods executed by a computing device, or as a computing device programmed with the instructions to perform the described actions.

[0166] At block 702 , a computing device may apply an acquisition method to receive a mass spectrum generated by a mass spectrometry apparatus, the acquisition method determining a number of separation dimensions available in the mass spectrum.

[0167] At block 704 , the computing device may define putative product ion spectra for the mass spectrum using techniques similar to those described above in connection with generating the putative portion of the MIR.

[0168] At block 706, the computing device may access a repository storing composite product ion spectra that match the putative product ion spectra. An example of such a repository is the MIR described above.

[0169] At block 708, the computing device may retrieve the next separation dimension to be considered. For each separation dimension in the separation dimension, the computing device may retrieve (at block 710) a regression curve generated based on the composite product ion spectrum and may generate a normalized mass spectrum by applying the corresponding regression curve (at block 712) to normalize the values ​​of the corresponding separation dimension. Figure 3The normalization process for finding the regression curve has been described above. The corresponding regression curve can be applied to the raw peak detections in the mass spectrum to correct for changes in at least one of time, mass-to-charge ratio, or drift.

[0170] As described above, normalizing the spectra can include clustering the putative product ion spectra into ion clusters and then re-clustering the putative product ion spectra based on the normalized values. Clustering can be performed by calculating a theoretical isotope distribution for the putative product ion spectra; clustering the putative product ion spectra based on the theoretical isotope distribution; determining that the intensity of an isotope present in the mass spectrum exceeds a predetermined threshold amount; forming virtual ions from the excess isotope; and clustering the virtual ions into new isotope groups.

[0171] At block 714, the system determines whether more separation dimensions are to be considered. If so, processing returns to block 708 and the next separation dimension is retrieved. Otherwise, processing proceeds to decision block 716.

[0172] In some embodiments, the computing device may retrieve multiple mass spectra and may verify the mass spectra to determine whether the mass spectra are reproducible with each other. To this end, at decision block 716, the device may determine whether more spectra are to be analyzed. If so, processing returns to block 702 and the next spectrum is selected and normalized as described above.

[0173] After all spectra have been analyzed ("None" at block 716), processing may proceed to block 718, where the device may verify that the spectra reproduce one another. To this end, the normalized spectra, as created at each iteration of block 712, may be compared; because the spectra have been normalized using the techniques described herein, the spectra should now be directly comparable (e.g., within a predetermined acceptable tolerance). To this end, the normalized values ​​for the separate dimensions may each be compared to a matching tolerance (e.g., in combination with the normalized values ​​for the dimensions). Figure 6 The match tolerance defines the window within which the normalized mass spectrum is considered to match the target spectrum (as described above). This allows spectra generated by different mass spectrometry devices of different types (e.g., different platforms using different acquisition methods) to be compared with each other.

[0174] In some embodiments, the computing device can further determine whether a target molecule is present in the sample from which the mass spectrum is generated. To this end, at box 720, the computing device can receive a list of two or more target molecules to be identified in the sample from which the mass spectrum is generated. At box 722, the computing device can select a target library customized for the sample, the target library comprising a set of precursor ions and product ions for identifying two or more target molecules, wherein the target library consists of a subset of the precursor ions and product ions represented in the mass spectrum. In some embodiments, the subset of precursor ions and product ions is the minimum subset necessary to identify the target molecule in the mass spectrum. A target library can be generated for MIR according to the above-mentioned techniques. At box 724, the computing device can use the target library to determine whether a target molecule is present in the sample (using the above-mentioned "target positioning" technique) and can output an indication of the presence of the molecule on a user interface.

[0175] Figure 8 An example of a system architecture and data processing device that can be used to implement one or more illustrative aspects described herein in a standalone and / or networked environment is shown. Various network nodes (such as data server 810, web server 806, computer 804, and laptop 802) can be interconnected via a wide area network (WAN) 808, such as the Internet. Other networks, including private intranets, corporate networks, LANs, metropolitan area networks (MANs), wireless networks, personal area networks (PANs), and the like, can also or alternatively be used. Network 808 is provided for illustrative purposes and can be replaced by fewer or additional computer networks. The local area network (LAN) can have one or more of any known LAN topologies and can use one or more of a variety of different schemes, such as Ethernet. Devices data server 810, web server 806, computer 804, laptop 802, and other devices (not shown) can be connected to one or more of the networks via twisted pair wiring, coaxial cable, fiber optic cables, radio waves, or other communication media.

[0176] The computer software, hardware, and networks may be used in a variety of system environments, including stand-alone environments, networked environments, remote access (also known as remote desktop) environments, virtualized environments, and / or cloud-based environments.

[0177] As used herein and as depicted in the accompanying drawings, the term "network" refers not only to a system in which remote storage devices are coupled together via one or more communication paths, but also to independent devices that may be coupled to such a system from time to time having storage capabilities. Thus, the term "network" includes not only a "physical network" but also a "content network" that includes data residing on various physical networks that are attributable to a single entity.

[0178] The components may include a data server 810, a web server 806, and client computers 804 and laptops 802. Data server 810 provides overall access, control, and management of the database and control software used to implement one or more illustrative aspects described herein. Data server 810 may be connected to web server 806, through which users interact with and obtain requested data. Alternatively, data server 810 may act as a web server itself and be directly connected to the Internet. Data server 810 may be connected to web server 806 via a network 808 (e.g., the Internet) via a direct or indirect connection or via some other network. Users may interact with data server 810 using remote computers 804 and laptops 802, for example, by using a web browser to connect to data server 810 via one or more externally published websites hosted on web server 806. Client computers 804 and laptops 802 may be used with data server 810 to access data stored therein or for other purposes. For example, from a client computer 804, a user may access the web server 806 using an internet browser as is known in the art, or by executing a software application over a computer network (such as the internet) that communicates with the web server 806 and / or data server 810.

[0179] The server and application can be combined on the same physical machine and retain separate virtual or logical addresses, or they can reside on separate physical machines. Figure 8 Only one example of a network architecture that may be used is shown, and those skilled in the art will appreciate that the specific network architecture and data processing equipment used may be different, and the functionality provided is secondary, as further described herein. For example, the services provided by web server 806 and data server 810 may be combined on a single server.

[0180] Each component—data server 810, web server 806, computer 804, laptop 802—can be any type of known computer, server, or data processing device. Data server 810 may, for example, include a processor 812 that controls the overall operation of data server 810. Data server 810 may also include RAM 816, ROM 818, a network interface 814, input / output interfaces 820 (e.g., a keyboard, mouse, display, printer, etc.), and memory 822. Input / output interfaces 820 may include various interface units and drivers for reading, writing, displaying, and / or printing data or files. Memory 822 may also store operating system software 824 for controlling the overall operation of data server 810, control logic 826 for instructing data server 810 to perform aspects described herein, and other application software 828 that provides secondary, support, and other functionality that may or may not be used in conjunction with the aspects described herein. This control logic may also be referred to herein as data server software control logic 826. The functionality of the data server software may refer to a combination of operations or decisions made automatically based on rules coded into the control logic, manually made by a user providing input into the system, and / or automated processing based on user input (e.g., queries, data updates, etc.).

[0181] Memory 1122 may also store data used to perform one or more aspects described herein, including a first database 832 and a second database 830. In some embodiments, the first database may include the second database (e.g., as separate tables, reports, etc.). That is, depending on the system design, information may be stored in a single database or separated into different logical, virtual, or physical databases. Web server 806, computer 804, and laptop 802 may have similar or different architectures as described with respect to data server 810. Those skilled in the art will appreciate that the functionality of data server 810 (or web server 806, computer 804, and laptop 802) as described herein can be extended across multiple data processing devices, for example, to distribute processing load across multiple computers, to separate transactions based on geographic location, user access level, quality of service (QoS), and the like.

[0182] One or more aspects may be implemented in computer-readable or readable data and / or computer-executable instructions (such as in one or more program modules) executed by one or more computers or other devices as described herein. Generally, program modules include routines, programs, objects, components, data structures, and the like that, when executed by a processor in a computer or other device, perform specific tasks or implement specific abstract data types. Modules may be written in a source code programming language that is subsequently compiled for execution, or in a scripting language such as (but not limited to) HTML or XML. Computer-executable instructions may be stored on a computer-readable medium, such as a non-volatile storage device. Any suitable computer-readable storage medium may be utilized, including hard disks, CD-ROMs, optical storage devices, magnetic storage devices, and / or any combination thereof. Furthermore, various transmission (non-storage) media representing data or events as described herein may be transmitted between a source and a destination in the form of electromagnetic waves via a signal-conducting medium (such as metal wires, optical fibers) and / or wireless transmission media (e.g., air and / or space). Various aspects described herein may be implemented as methods, data processing systems, or computer program products. Thus, various functions may be implemented in whole or in part in software, firmware, and / or hardware or hardware equivalents, such as integrated circuits, field programmable gate arrays (FPGAs), etc. Certain data structures may be used to more efficiently implement one or more aspects described herein, and such data structures are contemplated as within the scope of the computer-executable instructions and computer-usable data described herein.

[0183] The components and features of the above-described devices may be implemented using discrete circuits, application-specific integrated circuits (ASICs), logic gates, and / or single-chip architectures. Additionally, features of the devices may be implemented using microcontrollers, programmable logic arrays, and / or microprocessors, or any combination thereof. It should be noted that hardware elements, firmware elements, and / or software elements may be collectively or individually referred to herein as "logic" or "circuitry."

[0184] It should be understood that the exemplary devices shown in the above block diagrams may represent functional descriptive examples of specific implementations of the license. Therefore, omitting or including block functions depicted in the drawings does not mean that hardware components, circuits, software and / or elements for implementing these functions must be separated, omitted or included in the implementation.

[0185] At least one computer-readable storage medium may include instructions that, when executed, cause a system to perform any of the computer-implemented methods described herein.

[0186] Some embodiments may be described using the expressions "one embodiment" and "an embodiment" and their derivatives. These terms indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearance of the phrase "in one embodiment" in various places in the specification does not necessarily refer to the same embodiment. Furthermore, unless otherwise indicated, the above-described features are considered to be usable together in any combination. Therefore, any features discussed individually may be used in combination with each other unless it is noted that the features are incompatible with each other.

[0187] The detailed description herein may be presented in terms of a program executed on a computer or network of computers, with general reference to the annotations and symbols used herein. These program descriptions and representations are used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art.

[0188] A procedure is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. These operations are those requiring physical manipulations of physical quantities. Usually, but not necessarily, these quantities take the form of electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It proves convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. It should be noted, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to those quantities.

[0189] Additionally, the operations performed are often expressed in terms such as addition or comparison, which are often associated with mental operations performed by an operator. In most cases, no such capabilities of an operator are required in any of the operations described herein that form part of one or more embodiments. Instead, the operations are machine operations. Useful machines for performing the operations of the various embodiments include general-purpose digital computers or similar devices.

[0190] The expressions "coupled" and "connected," as well as their derivatives, may be used to describe some embodiments. These terms are not necessarily intended to be synonymous with each other. For example, some embodiments may be described using the terms "connected" and / or "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0191] Various embodiments also relate to devices or systems for performing these operations. This device can be constructed specifically for the desired purpose or it can include a general-purpose computer that is selectively activated or reconfigured by a computer program stored in the computer. The programs presented herein are not inherently related to a specific computer or other device. Various general-purpose machines can be used together with the programs written according to the teachings of this article or it may prove convenient to build more specialized devices to perform the required method steps. The required structure for various of these machines will appear from the description given.

[0192] It should be emphasized that the Abstract of the present disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It should be understood that the submitted Abstract will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features may be grouped together in a single embodiment for the purpose of simplifying the present disclosure. This approach of the present disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than those expressly recited in each claim. On the contrary, as reflected in the following claims, the subject matter of the present invention has fewer than all the features of a single disclosed embodiment. Therefore, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate embodiment. In the appended claims, the terms "including" and "in which" are used as the plain English equivalents of the corresponding terms "comprising" and "wherein," respectively. In addition, the terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements on their objects.

[0193] What has been described above includes examples of the architecture. It is, of course, not possible to describe every conceivable combination of components and / or methodologies, but one of ordinary skill in the art will appreciate that many additional combinations and permutations are possible. Accordingly, the novel architecture is intended to embrace all such changes, modifications, and variations that fall within the spirit and scope of the appended claims.

Claims

1. A computer-implemented method comprising: receiving, via a user interface, a list of two or more target molecules to be identified in a sample for analysis by a mass spectrometry device according to an experimental method defined by a plurality of parameters; performing a statistical analysis based on the list of two or more target molecules, the statistical analysis configured to determine a probability of misidentifying one or more of the target molecules in the list; selecting a set of values ​​for the plurality of parameters that reduces the probability below a predetermined threshold; presenting the selected set of values ​​on the user interface; as well as The sample is analyzed using the mass spectrometry device according to an experimental method defined by the set of selected values ​​for the plurality of parameters. 2 . The computer-implemented method of claim 1 , wherein the statistical analysis is performed based on a subset of the target molecules in the list, the subset representing a relatively small set of common markers from among the target molecules. 3 . The computer-implemented method of claim 1 , wherein the plurality of parameters of the experimental method include one or more of elution position, collision cross section, drift position, or mass accuracy.

4. The computer-implemented method of claim 1 , wherein the values ​​for the plurality of parameters are selected based on at least one of: a known three-dimensional location of one of the target molecules, a known fragmentation pattern of one of the target molecules, or A known relationship between two of the target molecules.

5. The computer-implemented method of claim 1 , further comprising presenting on the user interface a score representing the likelihood that the presence or absence of the two or more target molecules will be correctly identified in the sample given the set of selected values.

6. The computer-implemented method of claim 5, wherein the score is calculated by: (a) querying a consensus library of immutable attributes including spectral features; (b) counting the frequency of one of the target molecules in the common library; (c) converting said frequency counts into probability scores; Repeating steps (a) to (c) for a plurality of said target molecules; and The probability scores for the plurality of target molecules are multiplied together.

7. The computer-implemented method of claim 5, wherein the score increases as a function of the number of available separation dimensions as determined by the acquisition method applied.

8. The computer-implemented method of claim 5, wherein the score increases according to a resolving power of the mass spectrometry device.

9. The computer-implemented method of claim 5, wherein the score is calculated based on a combination of mass-to-charge ratio, retention time, and drift time of the two or more target molecules.

10. The computer-implemented method of claim 5, wherein the score is calculated based on fluctuations in measurements of a spectral standard across a plurality of spectral devices.

11. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to: receiving, via a user interface, a list of two or more target molecules to be identified in a sample for analysis by a mass spectrometry device according to an experimental method defined by a plurality of parameters; performing a statistical analysis based on the list of two or more target molecules, the statistical analysis configured to determine a probability of misidentifying one or more of the target molecules in the list; selecting a set of values ​​for the plurality of parameters that reduces the probability below a predetermined threshold; presenting the selected set of values ​​on the user interface; as well as The sample is analyzed using the mass spectrometry device according to an experimental method defined by the set of selected values ​​for the plurality of parameters.

12. The computer-readable storage medium of claim 11, wherein the statistical analysis is performed based on a subset of target molecules in the list, the subset representing a relatively small set of common markers from among the target molecules.

13. The computer-readable storage medium of claim 11, wherein the plurality of parameters of the experimental method include one or more of elution position, collision cross section, drift position, or mass accuracy.

14. The computer-readable storage medium of claim 11 , wherein the values ​​for the plurality of parameters are selected based on at least one of: a known three-dimensional location of one of the target molecules, a known fragmentation pattern of one of the target molecules, or A known relationship between two of the target molecules.

15. The computer-readable storage medium of claim 11, wherein the instructions further configure the computer to present on the user interface a score representing the likelihood that the presence or absence of the two or more target molecules will be correctly identified in the sample given a selected set of values.

16. The computer-readable storage medium of claim 15, wherein the score is calculated by: (a) querying a consensus library of immutable attributes including spectral features; (b) counting the frequency of one of the target molecules in the common library; (c) converting said frequency counts into probability scores; Repeating steps (a) to (c) for a plurality of said target molecules; and The probability scores for the plurality of target molecules are multiplied together.

17. The computer-readable storage medium of claim 15, wherein the score increases as a function of the number of available separation dimensions as determined by the applied acquisition method.

18. The computer-readable storage medium of claim 15, wherein the score increases according to a resolving power of the mass spectrometry device.

19. The computer-readable storage medium of claim 15, wherein the score is calculated based on a combination of mass-to-charge ratio, retention time, and drift time of the two or more target molecules.

20. The computer-readable storage medium of claim 15, wherein the score is calculated based on fluctuations in measurements of a spectral standard across a plurality of spectral devices.

21. A computing device comprising: processor; and a memory storing instructions that, when executed by the processor, configure the apparatus to: receiving, via a user interface, a list of two or more target molecules to be identified in a sample for analysis by a mass spectrometry device according to an experimental method defined by a plurality of parameters; performing a statistical analysis based on the list of two or more target molecules, the statistical analysis configured to determine a probability of misidentifying one or more of the target molecules in the list; selecting a set of values ​​for the plurality of parameters that reduces the probability below a predetermined threshold; presenting the selected set of values ​​on the user interface; as well as The sample is analyzed using the mass spectrometry device according to an experimental method defined by the set of selected values ​​for the plurality of parameters.

22. The computing device of claim 21, wherein the statistical analysis is performed based on a subset of target molecules in the list, the subset representing a relatively small set of common markers from among the target molecules.

23. The computing device of claim 21, wherein the plurality of parameters of the experimental method include one or more of elution position, collision cross section, drift position, or mass accuracy.

24. The computing device of claim 21 , wherein the values ​​of the plurality of parameters are selected based on at least one of: a known three-dimensional location of one of the target molecules, a known fragmentation pattern of one of the target molecules, or A known relationship between two of the target molecules.

25. The computing device of claim 21, wherein the instructions further configure the device to present on the user interface a score representing the likelihood that the presence or absence of the two or more target molecules will be correctly identified in the sample given a selected set of values.

26. The computing device of claim 25, wherein the score is calculated by: (a) querying a consensus library of immutable attributes including spectral features; (b) counting the frequency of one of the target molecules in the common library; (c) converting said frequency counts into probability scores; Repeating steps (a) to (c) for a plurality of said target molecules; and The probability scores for the plurality of target molecules are multiplied together.

27. The computing device of claim 25, wherein the score increases as a function of the number of available separation dimensions as determined by the acquisition method applied.

28. The computing device of claim 25, wherein the score increases according to the resolving power of the mass spectrometry device.

29. The computing device of claim 25, wherein the score is calculated based on a combination of mass-to-charge ratios, retention times, and drift times of the two or more target molecules.

30. The computing device of claim 25, wherein the score is calculated based on fluctuations in measurements of a spectral standard across a plurality of spectral devices.

Citation Information

Patent Citations

  • Methods and apparatus for mass spectrometry

    US6717130B2

  • Techniques for processing of mass spectral data

    US20180166265A1