Post-detection chromatographic peak verification
By using a machine learning classifier to verify chromatographic peaks, the problem of resource waste caused by false positive results in existing technologies is solved, and the efficiency of chromatographic and mass spectrometric analysis is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THERMO FINNIGAN LLC
- Filing Date
- 2025-12-15
- Publication Date
- 2026-06-19
AI Technical Summary
Existing chromatographic peak detection techniques are prone to false positive results, resulting in wasted time and computational resources on analyzing mass spectra that are of no value.
A machine learning classifier was used to verify suspected chromatographic peaks, separating the valid peak set from the invalid peak set, and mass spectrometry analysis was performed only on the valid peak set.
This reduces the need for mass spectrometry analysis of false positive peaks, saving time and computational resources and improving analytical efficiency.
Smart Images

Figure CN122238548A_ABST
Abstract
Description
Background Technology
[0001] The fields of chromatography and mass spectrometry may involve executing analytical algorithms that consume a significant amount of time. Summary of the Invention
[0002] The following summary is presented to provide a basic understanding of one or more embodiments. This summary is not intended to identify key or essential elements, or to depict any scope of a particular embodiment or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a prelude to a more detailed description that follows. In one or more embodiments described herein, devices, systems, computer-implemented methods, apparatuses, or computer program products that facilitate post-detection chromatographic peak validation are described.
[0003] According to one or more embodiments, a system is provided. In various aspects, the system may include a processor capable of executing computer-executable components stored in a non-transitory computer-readable storage memory. In various instances, the computer-executable components may include a scanning component that enables a chromatographic device coupled to a mass spectrometer to scan a sample, thereby generating a chromatogram and a mass spectrum. In various cases, the computer-executable components may include a peak component capable of identifying multiple suspected peaks in the chromatogram via a peak detection algorithm. In various aspects, the computer-executable components may include a model component capable of separating multiple suspected peaks into a set of valid peaks and a set of invalid peaks by performing a machine learning classifier on the corresponding suspected peaks among the multiple suspected peaks. In various instances, the computer-executable components may include an execution component capable of performing mass spectrometric analysis on a first portion of the mass spectrum corresponding to the set of valid peaks, but not on a second portion of the mass spectrum corresponding to the set of invalid peaks.
[0004] According to one or more embodiments, a computer-implemented method is provided. In various embodiments, the computer-implemented method may include, via a device operatively coupled to a processor, causing a chromatographic device coupled to a mass spectrometer to scan a sample, thereby generating a chromatogram and a mass spectrum. In various aspects, the computer-implemented method may include, via the device and via a peak detection algorithm, identifying multiple suspected peaks in the chromatogram. In various instances, the computer-implemented method may include, via the device and via performing a machine learning classifier on the respective suspected peaks among the multiple suspected peaks, separating the multiple suspected peaks into a set of valid peaks and a set of invalid peaks. In various cases, the computer-implemented method may include, via the device, performing mass spectrometric analysis on a first portion of the mass spectrum corresponding to the set of valid peaks, but not performing mass spectrometric analysis on a second portion of the mass spectrum corresponding to the set of invalid peaks.
[0005] According to one or more embodiments, a computer program product is provided for facilitating post-detection chromatographic peak validation. In various embodiments, the computer program product may include a non-transitory computer-readable storage device embodying program instructions. In various aspects, the program instructions are executable by a processor to cause the processor to scan a sample using a chromatographic device coupled to a mass spectrometer, thereby generating a chromatogram and a mass spectrum. In various instances, the program instructions may further execute to cause the processor to identify multiple suspected peaks in the chromatogram via a peak detection algorithm. In various cases, the program instructions may further execute to cause the processor to separate the multiple suspected peaks into a set of valid peaks and a set of invalid peaks by performing a machine learning classifier on the corresponding suspected peaks among the multiple suspected peaks. In various aspects, the program instructions may further execute to cause the processor to perform mass spectrometry analysis on a first portion of the mass spectrum corresponding to the set of valid peaks, but not on a second portion of the mass spectrum corresponding to the set of invalid peaks. Attached Figure Description
[0006] Various embodiments will be readily understood through the following detailed description taken in conjunction with the accompanying drawings. For ease of description, the same reference numerals indicate the same structural elements. The embodiments are shown in the figures by way of example rather than limitation. The drawings are not necessarily drawn to scale.
[0007] Figure 1 shows an exemplary non-limiting block diagram of a scientific instrument module according to various embodiments described herein.
[0008] Figure 2 shows an exemplary non-limiting flowchart of a computer-implemented method according to various embodiments described herein.
[0009] Figure 3 shows a block diagram of an exemplary non-limiting system for facilitating post-detection chromatographic peak validation according to one or more embodiments described herein.
[0010] Figure 4 shows a block diagram of an exemplary non-limiting system comprising chromatograms and multiple mass spectra for facilitating post-detection chromatographic peak validation according to one or more embodiments described herein.
[0011] Figure 5 shows an exemplary non-limiting block diagram illustrating chromatograms and multiple mass spectra according to one or more embodiments described herein.
[0012] Figure 6 shows a block diagram of an exemplary non-limiting system comprising multiple suspected peaks that facilitates post-detection chromatographic peak validation according to one or more embodiments described herein.
[0013] Figure 7 shows an exemplary non-limiting block diagram illustrating how multiple potential peaks can be obtained according to one or more embodiments described herein.
[0014] Figure 8 shows a block diagram of an exemplary non-limiting system for facilitating post-detection chromatographic peak validation, comprising a machine learning classifier, a set of valid peaks, and a set of invalid peaks, according to one or more embodiments described herein.
[0015] Figure 9 shows an exemplary non-limiting block diagram illustrating how a machine learning classifier according to one or more embodiments described herein can separate multiple suspected peaks into a set of valid peaks and a set of invalid peaks.
[0016] Figure 10 shows a block diagram of an exemplary non-limiting system, including a training component and a training dataset, for facilitating post-detection chromatographic peak validation according to one or more embodiments described herein.
[0017] Figure 11 shows an exemplary non-limiting block diagram of a training dataset according to one or more embodiments described herein.
[0018] Figure 12 shows an exemplary non-limiting block diagram illustrating how a machine learning classifier can be trained according to one or more embodiments described herein.
[0019] Figure 13 shows exemplary non-limiting experimental results according to one or more embodiments described herein.
[0020] Figure 14 shows a block diagram of an exemplary non-limiting operating environment that can facilitate one or more embodiments described herein.
[0021] Figure 15 illustrates an exemplary networking environment operable to perform the various specific implementations described herein. Detailed Implementation
[0022] The following detailed descriptions are illustrative only and are not intended to limit the implementation or application / use of the embodiments. Furthermore, there is no intention to be bound by any express or implied information presented in the foregoing background section, invention summary section, or detailed description section.
[0023] One or more embodiments will now be described with reference to the accompanying drawings, wherein the same reference numerals are used throughout to refer to the same elements. In the following description, numerous specific details are set forth for purposes of explanation, intended to provide a more thorough understanding of the one or more embodiments. However, it will be apparent, in various cases, that one or more embodiments may be implemented without these specific details.
[0024] Various operations can be described as multiple independent actions or operations in a manner most conducive to understanding the subject matter disclosed herein. However, the order of description should not be construed as implying that these operations necessarily depend on a specific order. Specifically, these operations may be performed in an order different from the order presented. The described operations may be performed in an order different from the described embodiments. Various additional operations may be performed, or the described operations may be omitted in additional embodiments.
[0025] While some elements may be mentioned in the singular (e.g., "a processing device"), any suitable element can be represented by multiple instances of that element, and vice versa. For example, a set of operations described as being performed by a processing device can be implemented by different operations performed by different processing devices. As used herein, unless otherwise stated, the phrase "based on" should be understood to mean "at least partially based on".
[0026] A mass spectrometer coupled to a chromatographic apparatus can be considered a scientific instrument that can be deployed in research, laboratory, clinical, or clinical settings to determine the chemical composition or components of unknown samples or specimens. To facilitate such chemical composition determination, a mass spectrometer or chromatographic apparatus may include a complex combination of driven components (e.g., ion source, ion lens, heater, cooler, chromatographic column, incubator, injector, mass analyzer, fluid valve, fluid pump, circuit switch), sensors (e.g., ion detector, voltmeter, thermistor, potentiometer, pressure gauge), or consumables (e.g., carrier fluid, calibrator, filter).
[0027] During a scan, a portion of any given sample can be injected into the chromatographic apparatus and thus pass through any hardware components that make up the apparatus (e.g., an oven-heated column). The chromatographic apparatus can be constructed or designed to elute different chemicals (e.g., molecules, compounds, analytes) within the injection portion of the sample at different times (e.g., physically separated or isolated from the rest of the injection portion of the sample). The time required to elute any particular chemical substance can be referred to as the retention time of that particular chemical substance (denoted as “RT” along the horizontal axis in Figure 13). The chromatographic apparatus can be configured to generate a chromatogram, which can be viewed as a graph of the intensity (e.g., the amplitude of the detector signal recorded by the chromatographic apparatus) changing with retention time. As the corresponding chemical substances elute in or from the chromatographic apparatus, these substances can enter a mass spectrometer and thus pass through any hardware components that make up the mass spectrometer (e.g., a mass analyzer). The mass spectrometer can be constructed or designed to separate (or in some cases, measure) these ions based on the mass-to-charge ratio of the individual ions constituting any particular chemical substance. Such separations or measurements can produce mass spectra of a specific chemical substance, which can be viewed as a graph of intensity (e.g., the amplitude of a detector signal recorded by a mass spectrometer) as a function of mass-to-charge ratio.
[0028] In various instances, a mass spectrometer can be viewed as generating a corresponding mass spectrum for each point in a chromatogram (e.g., each time-intensity tuple). However, it may be that not all mass spectra recorded by a mass spectrometer are analytically valuable. In fact, in all respects, only the mass spectrum of any point in the chromatogram that forms an identifiable peak is analytically valuable. After all, a peak in a chromatogram may correspond to elution and therefore a high concentration of the corresponding chemical substance, while a valley in a chromatogram may correspond to elution without the chemical substance (e.g., a valley may simply be a product of noise captured by the detector of the chromatographic apparatus). Therefore, these corresponding chemical substances can be identified or determined by analyzing any mass spectrum corresponding to a peak in the chromatogram in any suitable manner (e.g., via metabolomics algorithms, via proteomics algorithms, via statistical algorithms, via library retrieval or library scoring calculations).
[0029] Typically, any analysis of mass spectra corresponding to chromatographic peaks is extremely time-consuming (e.g., it could take minutes or hours, depending on the complexity or resolution of the mass spectrum). Therefore, avoiding such analyses of mass spectra corresponding to valleys or other non-peak regions of the chromatogram is highly advisable. In fact, performing such analyses on these mass spectra not only fails to yield valuable information about the sample composition but is also considered a significant waste of time and resources that could be better used to analyze mass spectra corresponding to chromatographic peaks (e.g., analyzing non-peak mass spectra yields no compositional information and incurs a substantial opportunity cost).
[0030] Unfortunately, existing techniques for identifying chromatographic peaks, such as thresholding (e.g., labeling any chromatographic point above a threshold intensity level as a peak), local maximum detection (e.g., labeling chromatographic points with an intensity higher than a threshold number of adjacent points by a threshold number of times as peaks), derivative calculation (e.g., peak labeling based on the first or second derivative of intensity with respect to retention time), curve fitting (e.g., labeling a sequence of chromatographic points with a bell-shaped curve as a peak), machine learning segmentation (e.g., a machine learning segmenter receives a chromatogram as input and determines which sequence of points in the chromatogram qualifies as a peak), or combinations thereof, frequently produce false positive results (e.g., the probability of misidentifying valleys or other non-peak regions as chromatographic peaks is unacceptably high). Therefore, when implementing existing techniques, excessive time and computational resources are wasted on analyzing mass spectra that do not actually correspond to chromatographic peaks. In other words, existing techniques can be considered to be plagued by a variety of technical problems.
[0031] Therefore, systems or technical solutions that can alleviate such technical problems can be considered to have application value.
[0032] The various embodiments described herein address one or more of these technical problems. These embodiments may include systems, computer-implemented methods, apparatuses, or computer program products that facilitate post-detection peak validation. Specifically, the inventors of the various embodiments described herein recognize that false positive results in the prior art can be eliminated, reduced, filtered, or otherwise addressed by training a machine learning classifier to act as a post-detection peak validator. In other words, regardless of the specific technique or combination of techniques chosen to detect or identify peaks in a given chromatogram, the machine learning classifier can then mark any suspected peaks detected or identified by such techniques as valid or invalid. Furthermore, regardless of the specific technique or combination of techniques used to detect peaks, the machine learning classifier can be considered as a redundancy, backup, filter, or sanity checker to distinguish between true positive peak detection results and false positive peak detection results. Therefore, the machine learning classifier can be considered an additional or supplementary safety layer that eliminates falsely detected peaks, eliminating the need to spend time and computational resources analyzing any mass spectra corresponding to those falsely detected peaks.
[0033] The various embodiments described herein can be considered as a computerized tool (e.g., any suitable combination of computer-executable hardware and computer-executable software) that can be electronically mounted on or otherwise mounted to a mass spectrometer equipped with a chromatograph, and that facilitates post-detection peak validation. In various aspects, such a computerized tool may include scanning components, peak components, modeling components, or execution components.
[0034] In various implementations, a mass spectrometer equipped with a chromatograph can be considered as including a mass spectrometer operatively coupled to a chromatographic apparatus in any suitable manner. In various aspects, a mass spectrometer can include any suitable component hardware. As some non-limiting examples, a mass spectrometer may include: any suitable ion beam emitter (e.g., matrix-assisted laser desorption / ionization [MALDI] source, electrospray ionization [ESI] source, atmospheric pressure chemical ionization [APCI] source, atmospheric pressure photoionization [APPI] source, inductively coupled plasma [ICP] source, electron ionization source, chemical ionization source, photoionization source, glow discharge ionization source, thermal spray ionization source, combined ion source); any suitable mass analyzer (e.g., quadrupole mass filter analyzer, ion trap analyzer, quadrupole ion trap analyzer, time-of-flight [TOF] analyzer, electrostatic trap [e.g., orbital trap] mass analyzer, Fourier transform ion cyclotron resonance [FT-ICR] mass analyzer); any suitable ion detector (e.g., electron multiplier detector, microchannel plate detector, mirror charge detector, Faraday cup detector); or any suitable ion optics (e.g., ion focusing lens, ion guide, ion deflector). Similarly, in various instances, chromatographic equipment may include any suitable constituent hardware (e.g., gas chromatography hardware, liquid chromatography hardware, ion chromatography hardware). As some non-limiting examples, chromatographic equipment may include: any suitable sample injector (e.g., hot injectors, such as split / splitless injectors, direct injectors, or gas sampling valves [GSV]; cold injectors, such as cold head injection [COC] or temperature programmed vaporization injection [PTV]; syringes; infusion sets; vaporizers; nebulizers); any suitable chromatographic column (e.g., any suitable capillary column containing any suitable adsorption packing material or having a different stationary phase film); any suitable column oven or heater; or any suitable fluid-carrying device (e.g., fluid valves, fluid pumps). In all respects, any suitable autosampler or auxiliary sampling device can be used in conjunction with chromatographic hardware for sample preparation and introduction (e.g., gas / liquid sampling valves, headspace autosamplers, solid-phase microextraction [SPME], headspace-solid-phase microextraction [HS-SPME]; in-tube extraction-dynamic headspace [ITEX-DHS], thermal desorbers [TD], purge-and-trap samplers [P&T], pyrolysis devices). In various instances, the carrier gas can be, but is not limited to, helium, hydrogen, nitrogen, argon, methane, or any suitable combination thereof. In all cases, given any sample, the sample can be injected into the chromatographic apparatus, separated into various components by the chromatographic apparatus, and these components can be ionized and subsequently analyzed by a mass spectrometer (e.g., a mass spectrometer can record the relative abundance of sample ions as a function of mass-to-charge ratio). In all respects, a mass spectrometer equipped with a chromatograph can be loaded with samples.
[0035] In various implementations, the computerized tool can electronically access the mass spectrometer equipped with the chromatograph. That is, the computerized tool can electronically connect to or communicate with the mass spectrometer equipped with the chromatograph, enabling any component of the computerized tool to electronically interact with the mass spectrometer equipped with the chromatograph (e.g., send electronic commands to it, read electronic signals from it).
[0036] In various implementations, the scanning component of a computerized tool can electronically cause a mass spectrometer equipped with a chromatograph to inject a portion of the loaded sample. The scanning component can then correspondingly cause the mass spectrometer equipped with the chromatograph to scan that portion of the loaded sample. In various aspects, the mass spectrometer equipped with the chromatograph can perform such scans using any suitable scanning protocol (e.g., full scan, selected ion monitoring, split injection, splitless injection). In any case, such scans can produce chromatograms and multiple mass spectra. A chromatogram can be multiple time-intensity tuples (e.g., a graph or curve showing the measured intensity as a function of retention time), while each of the multiple mass spectra can be multiple ratio-intensity tuples (e.g., a graph or curve showing the measured intensity as a function of mass-to-charge ratio). Each intensity value of the chromatogram can correspond to a specific mass spectrum among the multiple mass spectra.
[0037] In various implementations, the peak component of a computerized tool can electronically identify multiple potential peaks within a chromatogram. In various contexts, the peak component can accomplish this identification by applying any suitable peak detection technique to the chromatogram (such as thresholding, local maximum detection, derivative calculation, curve fitting, machine learning segmentation, or any suitable combination thereof). In various instances, each of the multiple potential peaks can be a corresponding consecutive string or sequence of time-intensity tuples that is inferred, predicted, or otherwise determined to collectively form a corresponding peak within the chromatogram, thereby representing a corresponding chemical substance within the loaded sample. It should be noted that one or more of the multiple potential peaks may be incorrect. In other words, any peak detection technique applied to the chromatogram by the peak component has the potential to inadvertently characterize certain non-peak sequences of time-intensity tuples in the chromatogram as peaks. In other words, these peak detection techniques may incorrectly determine that one or more given regions of the chromatogram correspond to a corresponding chemical substance in the loaded sample, when in fact, such one or more given regions do not actually correspond to any chemical substance in the loaded sample. Such erroneous characterization may be due to noise fluctuations in the detector signal of a mass spectrometer equipped with a chromatograph (e.g., chromatographic noise may interfere with or otherwise hinder the peak detection technique employed by the peak assembly).
[0038] In various implementations, the model components of a computerized tool can electronically store, maintain, control, or otherwise access the machine learning classifier. In various aspects, the machine learning classifier can employ any suitable artificial intelligence architecture. For example, in some cases, the machine learning classifier can employ any suitable deep learning internal architecture. For instance, a machine learning classifier can include any suitable number of layers of any suitable type (e.g., input layer, one or more hidden layers, output layer, any of which can be a convolutional layer, dense layer, long short-term memory (LSTM) layer, transformer layer, nonlinear layer, pooling layer, batch normalization layer, or padding layer). As another example, a machine learning classifier can include any suitable number of neurons in various layers (e.g., different layers can have the same or different numbers of neurons). As yet another example, a machine learning classifier can include any suitable activation function (e.g., softmax, sigmoid, hyperbolic tangent, rectified linear unit) among various neurons (e.g., different neurons can have the same or different activation functions). As yet another example, a machine learning classifier can include any suitable inter-neuron or inter-layer connections (e.g., forward connections, skip connections, recurrent connections). In other instances, machine learning classifiers can employ any other suitable artificial intelligence architecture, such as support vector machines, linear or logistic regression models, Naive Bayes models, or decision trees.
[0039] Regardless of its specific internal architecture, a machine learning classifier can be configured to label suspected chromatographic peaks as valid or invalid using a binary or dichotomous approach. That is, a machine learning classifier can be configured to take any given chromatographic peak (or any suitable property or characteristic thereof, such as height or width) as input and produce a classification label as output indicating whether the given chromatographic peak is valid (e.g., has been correctly identified as a peak) or invalid (e.g., has been incorrectly identified as a peak).
[0040] Therefore, in various implementations, the model component can perform a machine learning classifier on each of the multiple suspected peaks identified by the peak component, and such execution can produce multiple validity classification labels. For example, suppose the machine learning classifier employs a deep learning architecture. In such a case, for any particular suspected peak among the multiple suspected peaks, the model component can feed that particular suspected peak into the input layer of the machine learning classifier, which can complete a forward pass through one or more hidden layers, and the output layer of the machine learning classifier can compute a corresponding validity classification label for the particular suspected peak based on the activations provided by the one or more hidden layers. It should be noted that such validity classification labels can be any suitable electronic data indicating that: the particular suspected peak is valid, or has otherwise been correctly identified as a chromatographic peak; or the particular suspected peak is invalid, or has otherwise been incorrectly identified as a chromatographic peak. In other words, the machine learning classifier can be considered as determining whether any numerical patterns (which may be subtle, or not visually apparent or prominent) exhibited by the time-intensity tuples that make up the particular suspected peak characterize or indicate a true chromatographic peak.
[0041] By performing a machine learning classifier on each of the multiple suspected peaks in this manner, the model component can be considered to separate or divide the multiple suspected peaks into: a set of valid peaks; and a set of invalid peaks. In all respects, the set of valid peaks can be any peak among the multiple suspected peaks that has been labeled as valid by the machine learning classifier. Conversely, the set of invalid peaks can be any peak among the multiple suspected peaks that has been labeled as invalid by the machine learning classifier. In all cases, the set of valid peaks can be considered, or otherwise referred to as, a true positive result produced by any peak detection technique used by the peak component, while the set of invalid peaks can be considered, or otherwise referred to as a false positive result produced by any peak detection technique used by the peak component.
[0042] In various implementations, the execution component of the computerized tool can electronically perform any suitable type of statistical, numerical, or computational analysis on any mass spectrum corresponding to a valid set of peaks across multiple mass spectra. Conversely, the execution component can avoid performing such analysis on any mass spectra corresponding to a invalid set of peaks across multiple mass spectra. In other words, the execution component can ignore or discard any mass spectra corresponding to chromatographic regions that are incorrectly or erroneously identified as chromatographic peaks. In this way, valuable compositional information about the loaded sample can be obtained or derived (e.g., due to the analysis of mass spectra corresponding to a valid set of peaks) without excessively consuming time or computational resources (e.g., because time and resources are not spent or wasted on analyzing mass spectra corresponding to invalid sets of peaks).
[0043] Therefore, the various embodiments described herein can be considered as a software validator or evaluator that verifies any peaks detected in the chromatogram, allowing mass spectra of falsely detected peaks to be ignored during downstream analysis. Such embodiments can save significant time and resources compared to existing techniques that easily waste time and resources analyzing mass spectra of false positive peaks. Furthermore, such embodiments can be implemented regardless of the specific type of peak detection technology utilized or selected.
[0044] To enable the various implementations described herein to function properly, a machine learning classifier may first be trained. In various aspects, the computational tools may include training components that can facilitate such training (e.g., in a supervised manner), as described later herein.
[0045] The various implementation schemes described herein can be used to solve problems that are inherently technical (e.g., facilitating post-detection chromatographic peak validation), non-abstract, and cannot be performed by humans as a set of mental behaviors, using hardware or software. Furthermore, some procedures within the process can be performed by dedicated computers (e.g., chromatographic devices and mass spectrometers capable of injecting and scanning portions of samples; machine learning classifiers composed of specific types of neural network layers) to perform specific operations related to chromatographic and mass spectrometric analysis.
[0046] For example, such specific operations may include: scanning a sample with a chromatographic device coupled to a mass spectrometer via a device operably coupled to a processor, thereby generating a chromatogram and a mass spectrum; identifying multiple suspected peaks in the chromatogram via the device and via a peak detection algorithm; separating the multiple suspected peaks into a set of valid peaks and a set of invalid peaks via the device and via a machine learning classifier applied to the corresponding suspected peaks among the multiple suspected peaks; and performing mass spectrometric analysis on a first portion of the mass spectrum corresponding to the set of valid peaks, but not on a second portion of the mass spectrum corresponding to the set of invalid peaks. In various aspects, for a first suspected peak among the multiple suspected peaks, the device may feed the first suspected peak or one or more properties of the first suspected peak as input to the machine learning classifier, and the machine learning classifier may generate a classification label as output, the classification label indicating whether the first suspected peak is a valid peak or an invalid peak.
[0047] These specific operations are inherently computerized. In fact, chromatographic equipment and mass spectrometers are highly technical computerized devices, containing specific computerized hardware (e.g., temperature sensors, pressure sensors, voltage sensors, ion beam emitters, ion focusing lenses, mass analyzers, ion detectors, oven-heated columns, autosamplers). Without a computer, neither a mass spectrometer equipped with a chromatograph nor the operations it performs can be performed in any reasonable or feasible way by the human brain or by humans using only pen and paper (e.g., neither the human brain nor humans using only pen and paper can inject a portion of a sample or heat the column, ion source, mass analyzer, or detector in an oven-heated mass spectrometer equipped with a chromatograph). Furthermore, machine learning classifiers (e.g., artificial neural networks) are inherently computerized structures, including specific software-oriented architectures (e.g., input layers, hidden layers, or output layers, any of which can consist of trainable or untrainable internal parameters, such as convolutional layers or LSTM layers). Without a computer, machine learning classifiers cannot be trained or executed by human thought, nor can they be trained or executed in any reasonable or feasible way by humans using only pen and paper.
[0048] Furthermore, the various embodiments described herein can integrate various teachings related to chromatography and mass spectrometry analysis into practical applications. As explained above, chromatographic equipment generates a chromatogram, where peaks represent elution and thus determine the presence of the corresponding chemical substance, while a mass spectrometer can generate a mass spectrum for each point in the chromatogram (e.g., as the eluted chemical substance passes through the mass spectrometer). Mass spectra corresponding to peaks in the chromatogram can be considered analytically valuable (e.g., because they contain valuable information about the composition of the corresponding chemical substance). Conversely, mass spectra corresponding to non-peak regions in the chromatogram can be considered not analytically valuable. Therefore, any time or computational resources spent analyzing mass spectra corresponding to non-peak regions can be considered wasteful. Thus, it may be desirable to first selectively identify peaks in the chromatogram and then analyze only any mass spectra corresponding to those peaks. Unfortunately, existing techniques that facilitate peak detection (e.g., thresholding, local maximum detection, derivative calculation, curve fitting, machine learning segmentation) can be prone to false positives. That is, such peak detection techniques may misidentify more than an acceptable number of non-peak chromatographic regions as peaks. Therefore, when implementing existing technologies, it is often time-consuming or computationally resource-intensive to analyze mass spectra of non-peak regions that have been incorrectly detected as peaks, which may be undesirable. In other words, existing technologies can be considered to have one or more technical problems.
[0049] The various implementation schemes described herein can help improve such technical problems by facilitating post-detection chromatographic peak validation. Specifically, the various implementation schemes described herein can be viewed as automated validators that verify the work done by upstream peak detection techniques. More specifically, given a chromatogram, any suitable peak detection technique (e.g., thresholding, local maximum detection, derivative calculation, curve fitting, machine learning segmentation) can be applied to the chromatogram to identify or detect multiple suspected peaks within the chromatogram. As explained above, peak detection techniques can be erroneous, meaning that one or more of these multiple suspected peaks may actually be non-peak regions in the chromatogram. In various respects, those incorrectly identified non-peak regions can be found or filtered out using the machine learning classifier described herein. Specifically, this machine learning classifier can be trained to receive a suspected peak (or any suitable numerical property thereof, such as peak height or peak width) as input and produce a classification label for the suspected peak as output, where the classification label indicates whether the suspected peak is valid or invalid. In various respects, the machine learning classifier can be viewed as a redundancy or a second line of defense in verifying the work done by peak detection techniques. It should be noted that the machine learning classifier itself is not a peak detector (e.g., it does not accept chromatograms as input and identify one or more peaks in the chromatogram as output). Instead, the machine learning classifier can be viewed as a peak discriminator, trained to accept a string or sequence of time-intensity tuples predicted to form a chromatographic peak as input and determine whether that string or sequence of time-intensity tuples actually or indeed constitutes a chromatographic peak as output. In any case, regardless of how likely the peak detection technique is to misidentify non-peak regions as chromatographic peaks, the likelihood that the machine learning classifier will also misidentify such non-peak regions as valid is likely to be extremely low (e.g., because the machine learning classifier may be different from, independent of, or not a replica of the peak detection technique). Therefore, the machine learning classifier can be viewed as a filter, validator, discriminator, or sanity checker to remove false positives from the output of the peak detection technique. In this way, time or resources are not wasted on analyzing mass spectra corresponding to erroneously detected chromatographic peaks. Therefore, the various embodiments described herein can save time and computational resources compared to existing techniques.
[0050] Furthermore, the counterintuitive nature of the various implementations described herein must be emphasized. It is always desirable to reduce the number of false positives produced by peak detection techniques (e.g., thresholding, local maximum detection, derivative calculation, curve fitting, machine learning segmentation). The conventional wisdom for reducing such false positives tells us that the peak detection technique itself should be improved or upgraded. The inventors break with this conventional wisdom by recognizing that false positives can be reduced alternatively via a validator or discriminator employed downstream of the peak detection technique. In other words, those who wish to reduce the number of false positives produced by peak detection techniques would attempt to enhance the peak detection technique itself (e.g., try to set more precise thresholds or derive more refined formulas); they certainly wouldn't think of reducing the number of false positives produced by peak detection techniques by implementing a completely independent computational entity downstream of the peak detection algorithm (e.g., the machine learning classifier described herein). In other words, the various implementations described herein can be seen as ingenious, unusual, or counterintuitive solutions to the problem of false positives misidentified by chromatographic peak detection techniques. In other words, the various implementations described herein can be seen as ingenious, unusual, or counterintuitive uses of machine learning classifiers.
[0051] For at least these reasons, the various implementation schemes described herein can be considered a concrete and practical technical improvement in the field of chromatography and mass spectrometry. Therefore, the various implementation schemes described herein undoubtedly represent a useful and practical application of computers.
[0052] Furthermore, the various embodiments described herein can control real-world tangible devices based on the disclosed teachings. For example, the various embodiments described herein can electronically activate, deactivate, or otherwise manipulate the real hardware (e.g., sample injector, ion beam emitter, ion focusing lens, flow-carrying fluid valve / pump) of real scientific instruments (e.g., chromatography equipment, mass spectrometer, autosampler).
[0053] Figure 1 shows an exemplary non-limiting block diagram of a scientific instrument module 102 according to various embodiments described herein.
[0054] In various implementations, the scientific instrument module 102 can be implemented via a circuit system (e.g., including electrical or optical components) such as a programmed computing device. The logic components of the scientific instrument module 102 can be incorporated into a single computing device or, depending on the situation, distributed across multiple communicating computing devices. Examples of computing devices that can implement the scientific instrument module 102 individually or in combination are discussed herein with reference to Figure 14, and examples of systems or networks of interconnected computing devices in which the scientific instrument module 102 can be implemented across one or more computing devices are discussed herein with reference to Figure 15.
[0055] Scientific instrument module 102 may include a first logic unit 104, a second logic unit 106, a third logic unit 108, and a fourth logic unit 110. As used herein, a "logic unit" may include means for performing a set of operations associated with that logic unit. For example, any logic element included in scientific instrument module 102 may be implemented by one or more computing devices programmed with instructions to cause one or more processing devices of the computing device to perform an associated set of operations. In a particular embodiment, a logic element may include one or more non-transitory computer-readable media having instructions on them that, when executed by one or more processing devices of the one or more computing devices, cause the one or more computing devices to perform an associated set of operations. As used herein, the term "module" may refer to a collection of one or more logic elements that together perform the functions associated with the module. Different logic elements in a module may take the same form or may take different forms. For example, some logic elements in a module may be implemented by a programmed general-purpose processing device, while other logic elements in the module may be implemented by an application-specific integrated circuit (ASIC). In another example, different logic elements in a module may be associated with different sets of instructions executed by one or more processing devices. A module may omit one or more logic elements depicted in the relevant figures; for example, a module may include a subset of the logic elements depicted in the relevant figures when it is to perform a subset of the operations discussed herein with reference to the module.
[0056] In various embodiments, a scientific instrument corresponding to scientific instrument module 102 may be present. In various aspects, the scientific instrument can be any suitable computerized device capable of electronically measuring some scientifically relevant, clinically relevant, or research-related characteristics, properties, or attributes of an analytical sample (e.g., a mixture, compound, or collection of substances, known or unknown). As a non-limiting example, the scientific instrument can be a mass spectrometer operatively coupled to a chromatographic apparatus. In such cases, the scientific instrument can measure or capture a chromatogram (e.g., relative abundance of substances changing with retention time) or a mass spectrum (e.g., relative abundance of ions changing with mass-to-charge ratio) of the analytical sample.
[0057] In various implementations, the first logic component 104 can cause a scientific instrument to scan and analyze a sample. Such a scan can produce chromatograms and mass spectra corresponding to the analyzed sample.
[0058] In various implementations, the second logic component 106 can identify multiple suspected peaks in the chromatogram using any suitable peak detection technique, such as thresholding, curve fitting, or machine learning segmentation.
[0059] In various implementations, the third logic unit 108 can separate multiple suspected peaks into a set of valid peaks and a set of invalid peaks by executing a machine learning classifier. Specifically, the machine learning classifier can be configured to take any suspected peak or its properties as input and produce a binary label as output indicating whether the suspected peak is valid (e.g., a true positive result produced by logic unit 106) or invalid (e.g., a false positive result produced by logic unit 106). In some cases, the machine learning classifier can be considered as validating, evaluating, or reviewing the work done by peak detection techniques to identify instances where peak detection techniques misdetect non-peak portions of the chromatogram as peaks.
[0060] In various implementations, the fourth logic unit 110 can perform mass spectrometry analysis (e.g., any suitable type of statistical or numerical processing) on any mass spectrum corresponding to a valid set of peaks, and can avoid performing mass spectrometry analysis on any mass spectrum corresponding to a invalid set of peaks. Therefore, there is no need to waste time and resources analyzing mass spectra corresponding to falsely detected peaks in the chromatogram.
[0061] Therefore, the scientific instrument module 102 can facilitate the validation of chromatographic peaks after detection.
[0062] Figure 2 is an exemplary non-limiting flowchart of a computer-implemented method 200 according to various embodiments described herein. The operation of the computer-implemented method 200 can be used in any applicable scenario to perform any suitable operation (e.g., it can be performed or used in conjunction with any of the various modules, computing devices, or graphical user interfaces described with respect to Figures 1, 14, or 15). In Figure 2, the operations are each shown once in a specific order, but the operations can be reordered or repeated as needed and as appropriate (e.g., different operations performed can be performed in parallel where appropriate).
[0063] In various aspects, action 202 may include performing a first operation, namely, causing a chromatographic device coupled to a mass spectrometer to scan a sample via a device operably coupled to the processor, thereby generating a chromatogram and a mass spectrum. In various cases, the first logic unit 104 may perform or otherwise facilitate action 202.
[0064] In various instances, action 204 may include performing a second operation, namely, identifying multiple suspected peaks in the chromatogram via the device and a peak detection algorithm. In various cases, the second logic unit 106 may perform or otherwise facilitate action 204.
[0065] In various instances, action 206 may include performing a third operation, namely, separating the multiple suspected peaks into a set of valid peaks and a set of invalid peaks by performing a machine learning classifier on the respective suspected peaks among the multiple suspected peaks through the device. In various cases, the third logic component 108 may perform or otherwise facilitate action 206.
[0066] In various instances, action 208 may include performing a fourth operation, namely, performing mass spectrometry analysis on the first portion of the mass spectrum corresponding to the valid peak set, but not on the second portion of the mass spectrum corresponding to the invalid peak set. In various cases, the fourth logic unit 110 may perform or otherwise facilitate action 208.
[0067] Therefore, the computer-implemented method 200 can facilitate post-detection chromatographic peak validation.
[0068] Figure 3 shows a block diagram of an exemplary non-limiting system that can facilitate post-detection chromatographic peak validation according to one or more embodiments described herein.
[0069] In various embodiments, a mass spectrometer 302 equipped with a chromatograph may be included. In various aspects, the mass spectrometer 302 equipped with a chromatograph, as its name suggests, may consist of or otherwise include a chromatographic device operatively coupled to the mass spectrometer.
[0070] In various embodiments, the chromatographic apparatus of the mass spectrometer 302 equipped with the chromatograph can be any suitable chromatographic apparatus, such as a gas chromatograph or a liquid chromatograph. In various aspects, the chromatographic apparatus can include any suitable component hardware for separating an analytical sample into two or more components. As a non-limiting example, the component hardware can include an injector, an oven-heated column, and a carrier fluid valve or pump. In various aspects, the carrier fluid valve or pump can allow a carrier fluid (e.g., an inert gas or a water-organic solvent mixture) to flow through the chromatographic apparatus. In various instances, the injector can inject an analytical sample (e.g., a compound or solution to be measured or analyzed) into the flowing carrier fluid. In various cases, the injected analytical sample can be carried by the flowing carrier fluid through an oven-heated column, which can contain any suitable adsorbent packing material or stationary phase membrane. In various aspects, different components of the analytical sample (e.g., different chemical elements, molecules, analytes, or substances) can interact differently or uniquely with the adsorbent packing material or stationary phase membrane, resulting in different flow rates of the different components of the analytical sample through the oven-heated column. Due to these different flow rates, the different components can be considered physically separated from each other.
[0071] In various aspects, the mass spectrometer of the mass spectrometer 302 equipped with the chromatograph can be any suitable mass spectrometer. In various instances, the mass spectrometer can include any suitable component hardware for measuring the ion spectrum of an analytical sample. As a non-limiting example, the component hardware can include an ion beam emitter, ion optics, a mass analyzer, and an ion detector. In various cases, the ion beam emitter can receive a component of the analytical sample from the chromatographic apparatus and can ionize that component into an ion beam. The ion beam emitter can facilitate this via any suitable ionization technique, such as electron ionization, chemical ionization, matrix-assisted laser desorption / ionization, electrospray ionization, photoionization, or inductively coupled plasma ionization, any of which can be performed in a vacuum or at atmospheric pressure. In various aspects, the ion optics can guide or manipulate the ion beam generated by the ion beam emitter through the mass analyzer and to the ion detector. Non-limiting examples of such ion optics can include ion focusing lenses, ion guides, or ion deflectors. In various instances, the mass analyzer can separate or classify any ions present in the ion beam according to their mass-to-charge ratio. Non-limiting examples of mass analyzers may include quadrupole mass analyzers, time-of-flight mass analyzers, magnetic sector mass analyzers, electrostatic sector mass analyzers, quadrupole ion trap mass analyzers, or ion cyclotron resonance mass analyzers. In various cases, the ion detector can electronically detect or measure the relative abundance of any ions that collide with it. Non-limiting examples of ion detectors may include electron-multiplying ion detectors or Faraday cup ion detectors.
[0072] In various embodiments, the mass spectrometer 302 equipped with a chromatograph may currently or now be loaded with sample 304. In other words, sample 304 may be physically present within any suitable injector or autosampler of the mass spectrometer 302 equipped with a chromatograph, enabling the mass spectrometer 302 equipped with a chromatograph to inject a portion of sample 304 for analysis or scanning. In various aspects, sample 304 may be any suitable mixture, solution, or colloid for which mass spectrometry analysis is desired. As a non-limiting example, sample 304 may be a mixture, solution, or colloid of food or beverage. As another non-limiting example, sample 304 may be a mixture, solution, or colloid of pharmaceutical or medical products. As yet another non-limiting example, sample 304 may be a mixture, solution, or colloid of soil, water, air, excrement, or other environmental samples. In various instances, sample 304 may have or exhibit any suitable stability or shelf life. In fact, in some cases, sample 304 may be a stable mixture, solution, or colloid with a shelf life of several days, weeks, months, or even years. In other cases, sample 304 may be an unstable mixture, solution, or colloid with a shelf life of only a few hours or minutes.
[0073] In any case, it may be desirable to perform any suitable mass spectrometry analysis on or against sample 304. As described herein, system 306 can facilitate or otherwise accomplish such a goal.
[0074] In various aspects, system 306 may include processor 308 (e.g., computer processing unit, microprocessor) and nontransitory computer-readable storage 310 operatively or communicatively connected or coupled to processor 308. Nontransitory computer-readable storage 310 may store computer-executable instructions that, when executed by processor 308, cause processor 308 or other components of system 306 (e.g., scan component 312, peak component 314, model component 316, execution component 318) to perform one or more actions. In various embodiments, nontransitory computer-readable storage 310 may store computer-executable components (e.g., scan component 312, peak component 314, model component 316, execution component 318), and processor 308 may execute these computer-executable components.
[0075] In various embodiments, system 306 can be electronically coupled to or integrated with mass spectrometer 302 equipped with chromatograph via any suitable wired or wireless electronic connection. Therefore, system 3056 can electronically access mass spectrometer 302 equipped with chromatograph. That is, system 306 can electronically communicate with or otherwise electronically interact with mass spectrometer 302 equipped with chromatograph (e.g., transmit electronic instructions or commands to or receive electronic data from it). Therefore, any component of system 306 can interact with, communicate with, or otherwise manipulate mass spectrometer 302 equipped with chromatograph.
[0076] In various embodiments, system 306 may include scanning component 312. In various aspects, as described herein, scanning component 312 may enable mass spectrometer 302 equipped with a chromatograph to generate chromatograms and multiple mass spectra of sample 304.
[0077] In various implementations, system 306 may include peak component 314. In various instances, as described herein, peak component 314 may identify multiple potential peaks within a chromatogram.
[0078] In various implementations, system 306 may include model component 316. In various cases, model component 316 can separate multiple suspected peaks into a set of valid peaks and a set of invalid peaks by executing a machine learning classifier.
[0079] In various embodiments, system 306 may include execution component 318. In various respects, as described herein, execution component 318 may perform any suitable mass spectrometry analysis on any mass spectrum corresponding to a valid set of peaks, but not on any suitable mass spectrometry analysis on a mass spectrum corresponding to a invalid set of peaks.
[0080] It should be noted that in various instances, scanning component 312, peak component 314, model component 316, and execution component 318 can be collectively considered as one or more software components 311 of system 306. In all respects, it should be understood that, for ease of explanation and illustration, this document primarily describes one or more software components 311 as comprising four components (e.g., scanning component 312, peak component 314, model component 316, and execution component 318). However, one or more software components 311 are not limited to being implemented precisely as such four components in every embodiment. In fact, in some embodiments, the functionality of these four components described herein can be combined in any suitable manner so that it can be implemented by or by fewer than four components (e.g., in some cases, a single component can perform all the functionality described herein with respect to scanning component 312, peak component 314, model component 316, and execution component 318). In other implementations, the functionality of the four components described herein may (as an alternative) be distributed, separated, split, or fragmented in any suitable manner so as to be implemented by or by more than four components (e.g., two or more components may facilitate functionality that can be performed by the scanning component 312; two or more components may facilitate functionality that can be performed by the peak component 314; two or more components may facilitate functionality that can be performed by the model component 316; two or more components may facilitate functionality that can be performed by the execution component 318).
[0081] Figure 4 shows a block diagram of an exemplary non-limiting system comprising chromatograms and multiple mass spectra that can facilitate post-detection peak validation according to one or more embodiments described herein.
[0082] In various embodiments, scanning component 312 may electronically instruct, electronically command, or otherwise electronically cause mass spectrometer 302 equipped with a chromatograph to scan sample 304 according to any suitable chromatographic or spectroscopic scanning protocol. As a non-limiting example, mass spectrometer 302 equipped with a chromatograph may perform any suitable type of full scan or full-spectrum scan protocol on sample 304, wherein a wide scan is performed over a defined mass-to-charge ratio range or interval. As another non-limiting example, mass spectrometer 302 equipped with a chromatograph may perform any suitable type of selected ion monitoring (SIM) protocol on sample 304, wherein only a few specific mass-to-charge ratios are targeted. As yet another non-limiting example, mass spectrometer 302 equipped with a chromatograph may perform any suitable type of multiple reaction monitoring (MRM) protocol on sample 304, wherein precursor ions are fragmented by collision-induced dissociation and selectively monitored together with their fragment ions. As another non-limiting example, the mass spectrometer 302 equipped with a chromatograph can scan sample 304 with any suitable type of temperature rise or temperature control program for oven-heating the chromatographic column to separate or elute compounds of different volatiles. As yet another non-limiting example, the mass spectrometer 302 equipped with a chromatograph can perform any suitable type of splitless protocol on sample 304, wherein the entire injection portion of sample 304 is sent to the oven-heated chromatographic column. As yet another non-limiting example, the mass spectrometer 302 equipped with a chromatograph can perform any suitable type of split protocol on sample 304, wherein less than the entire injection portion of sample 304 is sent to the oven-heated chromatographic column. As yet another non-limiting example, the mass spectrometer 302 equipped with a chromatograph can perform any suitable type of isocratic elution protocol on sample 304, wherein the mobile phase composition of the mass spectrometer 302 equipped with a chromatograph remains constant throughout the scan. As another non-limiting example, the mass spectrometer 302 equipped with a chromatograph can perform any suitable type of gradient elution protocol on the sample 304, wherein the composition of the mobile phase changes throughout the scan.
[0083] Regardless of the scanning protocol employed, such scanning enables a mass spectrometer 302 equipped with a chromatograph to produce a chromatogram 402 and multiple mass spectra 404. In various instances, the chromatogram 402 may be a graph or curve showing the intensity versus retention time of sample 304, and each of the multiple mass spectra 404 may be a graph or curve showing the intensity versus mass-to-charge ratio. Non-limiting aspects are described with reference to Figure 5.
[0084] Figure 5 shows an exemplary non-limiting block diagram illustrating a chromatogram 402 and a plurality of mass spectra 404 according to one or more embodiments described herein.
[0085] In various aspects, chromatogram 402 may include multiple time-intensity tuples 502. In various instances, the total number of time-intensity tuples 502 can be [number missing]. n tuples, of which n For any suitable positive integer: time-intensity tuple 502(1) to time-intensity tuple 502( n In various cases, each of the multiple time-intensity tuples 502 can be a bi-element vector, the first element of which can be a scalar indicating the corresponding retention time, and the second element of which can be a scalar indicating how much intensity or concentration was measured at the corresponding retention time by the chromatographic detector of the mass spectrometer 302 equipped with the chromatogram. As a non-limiting example, time-intensity tuple 502(1) can be a vector indicating a first retention time and a first intensity value measured at that first retention time. As another non-limiting example, time-intensity tuple 502( n ) can be an instruction for the first n Retention time and in the [number]th n The retention time measured n A vector of intensity values. In various aspects, it is possible that two time-intensity tuples in multiple time-intensity tuples 502 do not have the same retention time as each other. Furthermore, in various instances, it is possible that multiple time-intensity tuples 502 are ordered chronologically from lowest retention time to highest retention time. Therefore, chromatogram 402 can be viewed as a time series of intensities measured by the chromatographic apparatus of mass spectrometer 302 equipped with a chromatograph.
[0086] In various aspects, multiple mass spectra 404 can each correspond to multiple time-intensity tuples 502 (e.g., in a one-to-one manner). Therefore, since multiple time-intensity tuples 502 can have... n If there are multiple tuples, then multiple mass spectra 404 can have n One spectrum: mass spectrum 404(1) to mass spectrum 404( n In various instances, each of the multiple mass spectra 404 can be a graph or curve of ion intensity versus mass-to-charge ratio measured by a spectral detector of a mass spectrometer 302 equipped with a chromatograph at a corresponding retention time. As a non-limiting example, mass spectrum 404(1) can correspond to a time-intensity tuple 502(1). Thus, mass spectrum 404(1) can be a first graph or curve of ion intensity versus mass-to-charge ratio produced by a mass spectrometer of a mass spectrometer 302 equipped with a chromatograph for any chemical substance (if any) eluted at a first retention time. As another non-limiting example, mass spectrum 404( nThis can correspond to the time-intensity tuple 502. n Therefore, mass spectrum 404 ( n It can be a mass spectrometer 302 equipped with a chromatograph, targeting the first... n The ionic strength of any chemical substances eluted during the retention time (if any) is compared to the mass-to-charge ratio. n Charts or graphs.
[0087] It should be noted that some of the mass spectra in the plurality of mass spectra 404 can be considered to contain valuable information about the composition of sample 304, while others in the plurality of mass spectra 404 can be considered not to contain such valuable information. Specifically, any chemical substance constituting sample 304 can elute from or within the mass spectrometer 302 equipped with the chromatograph at a corresponding retention time. Such elution can be displayed or manifested as a peak within chromatogram 402. In other words, the retention time of a peak displayed in chromatogram 402 can be considered to represent the time during which the corresponding chemical substance eluted, appeared, or otherwise existed at a considerable concentration, while the retention time of a peak not displayed in chromatogram 402 can be considered to represent the time during which no chemical substance eluted, appeared, or otherwise existed at a considerable concentration. Therefore, any mass spectrum in the plurality of mass spectra 404 corresponding to a peak in chromatogram 402 can be considered to contain valuable compositional information about the chemical substances in sample 304. Conversely, any mass spectrum in multiple mass spectra 404 that does not correspond to a peak in chromatogram 402 can be considered as not containing valuable compositional information about the chemical substances in sample 304.
[0088] Figure 6 shows a block diagram of an exemplary non-limiting system comprising multiple suspected peaks that can facilitate post-detection chromatographic peak validation according to one or more embodiments described herein.
[0089] In various implementations, peak assembly 314 can electronically identify or detect multiple suspected peaks 602 within chromatogram 402. Non-limiting aspects are described with reference to Figure 7.
[0090] Figure 7 shows an exemplary non-limiting block diagram illustrating how multiple suspected peaks 602 can be obtained according to one or more embodiments described herein.
[0091] In various implementations, peak component 314 can electronically apply or incorporate any suitable peak detection technique to or onto chromatogram 402.
[0092] As a non-limiting example, peak component 314 can be applied to or incorporated into chromatogram 402 using any suitable type of thresholding peak detection technique. In such cases, peak component 314 can be considered as searching for a continuous intensity string or sequence in chromatogram 402 that exceeds any suitable intensity threshold over time. In other words, any intensity value recorded in chromatogram 402 above the intensity threshold can be considered part of a peak, while conversely, any intensity value recorded in chromatogram 402 below the intensity threshold can be considered not part of a peak.
[0093] As another non-limiting example, peak component 314 can be applied to or incorporated into chromatogram 402 using any suitable type of local maximum peak detection technique. In such a case, peak component 314 can be considered as searching for intensities in chromatogram 402 that are greater than their neighboring intensities. In other words, any record in chromatogram 402 with an intensity value higher than both its preceding and following intensity values can be considered part of a peak (e.g., as the apex of a peak), while conversely, any record in chromatogram 402 with an intensity value not greater than either its preceding or following intensity value can be considered not part of a peak.
[0094] As yet another non-limiting example, peak component 314 can be applied to or incorporated into chromatogram 402 using any suitable type of derivative peak detection technique. In such cases, peak component 314 can be considered as a temporally continuous string or sequence of intensities in chromatogram 402 searching for its first or second derivative (e.g., its slope or concavity) that satisfies any suitably defined equality or inequality. For example, a string or sequence of intensities in chromatogram 402 with negative concavity and a slope passing through zero can be considered as forming a peak, while conversely, a string or sequence of intensities in chromatogram 402 with non-negative concavity or a slope not passing through zero can be considered as not forming a peak.
[0095] As another non-limiting example, peak component 314 can be applied to or incorporated into chromatogram 402 using any suitable type of curve-fitting peak detection technique. In such cases, peak component 314 can be considered as searching for time-continuous intensity strings or sequences in chromatogram 402, which can be approximated (e.g., by least squares) via a defined curve with adjustable parameters (such as a Gaussian curve (e.g., whose adjustable parameters can be amplitude, mean, or standard deviation), a Lorenz curve (e.g., whose adjustable parameters can be height, center, and half-width at half-height), or a polynomial curve (e.g., whose adjustable parameters can be the coefficients of the corresponding polynomial terms)). For example, intensity strings or sequences in chromatogram 402 that can fit to the defined curve with a deviation or error less than a threshold amount can be considered to form peaks, while intensity strings or sequences in chromatogram 402 that cannot fit to the defined curve with a deviation or error less than a threshold amount can be considered not to form peaks.
[0096] As another non-limiting example, peak component 314 can be applied to or applied to chromatogram 402 using any suitable type of machine learning peak detection technique. In such a case, peak component 314 can be considered as feeding chromatogram 402 as input to a machine learning segmenter configured to specify one or more temporally consecutive intensities that it deems eligible as or constitute a peak.
[0097] As another non-limiting example, peak component 314 can be applied to or applied to any suitable combination of the above techniques to chromatogram 402.
[0098] Regardless of the specific peak detection technique or combination of peak detection techniques chosen or selected, applying such peak detection to chromatogram 402 can produce multiple potential peaks 602. In various respects, multiple potential peaks 602 can include a total of... p One suspected peak, among which Any suitable positive integer: suspected peak 602(1) to suspected peak 602( p In various instances, each of the plurality of suspected peaks 602 may be a temporally consecutive string or sequence of intensity values from chromatogram 402, which peak component 314 has identified as a peak. In other words, each of the plurality of suspected peaks 602 may be a set of mutually adjacent time-intensity tuples from chromatogram 402, and peak component 314 has determined that it represents the elution of the corresponding chemical substance of sample 304. As a non-limiting example, suspected peak 602(1) may be composed of a total of Each tuple consists of: time-intensity tuple 602(1)(1) to time-intensity tuple 602(1)( In every respect, these Each tuple can be temporally continuous, meaning they can be adjacent to each other in chronological order or otherwise close in chromatogram 402. In other words, in chromatogram 402, there may not be time-intensity tuples 602(1)(1) and 602(1)(1) in chronological order. The time-intensity tuples between but not belonging to the suspected peak 602(1). As another non-limiting example, the suspected peak 602( p It can be composed of the total Composition of tuples: Time-Intensity tuple 602 ( p (1) To time-intensity tuple 602 ( p ()( As mentioned above, these Individual tuples can be temporally consecutive, meaning they can be adjacent to each other in chronological order or otherwise close in chromatogram 402. In other words, in chromatogram 402, there may not be a time-intensity tuple 602 that is chronologically adjacent to each other. p (1) with time-intensity tuple 602 ( p ()( Between but not belonging to the suspected peak 602 ( p The time-intensity tuple.
[0099] It should be noted that some of the suspected peaks may (e.g., due to noise) be false, inaccurate, incorrect, or not actually, or in fact, within chromatogram 402. In fact, the multiple suspected peaks 602 can be considered simply any part, fragment, or segment of chromatogram 402 that any peak detection technique implemented by peak component 314 considers or determines to be a peak. Because any peak detection technique implemented by peak component 314 may have some non-zero probability of producing false positive results, one or more of the multiple suspected peaks 602 may be false positives (e.g., not actually indicating the elution of the corresponding chemical substance of sample 304). At least for this reason, the term "suspected" may be considered appropriate.
[0100] Figure 8 shows a block diagram of an exemplary non-limiting system comprising a machine learning classifier, a set of valid peaks, and a set of invalid peaks, which can facilitate post-detection chromatographic peak validation according to one or more embodiments described herein.
[0101] In various implementations, model component 316 may electronically store, maintain, control, or otherwise electronically access machine learning classifier 802. In various instances, model component 316 may utilize machine learning classifier 802 to separate, partition, or assign multiple suspected peaks 602 into a set of valid peaks 804 and a set of invalid peaks 806. Various non-limiting aspects are described with reference to Figure 9.
[0102] Figure 9 shows an exemplary non-limiting block diagram illustrating how a machine learning classifier 802, according to one or more embodiments described herein, can separate multiple suspected peaks 602 into a set of valid peaks 804 and a set of invalid peaks 806.
[0103] In various implementations, the machine learning classifier 802 can employ any suitable type, style, construction, or design of internal architecture. For example, the machine learning classifier 802 can employ any suitable deep learning internal architecture.
[0104] In practice, a machine learning classifier 802 can have an input layer, one or more hidden layers, and an output layer in various scenarios. In various instances, any of these layers can be coupled together by any suitable inter-neuron or inter-layer connection, such as forward connections, skip connections, or recurrent connections. Furthermore, in various scenarios, any of these layers can be any suitable type of neural network layer with any applicable learnable or trainable internal parameters. For example, any of such an input layer, one or more hidden layers, or output layer can be a convolutional layer, whose learnable or trainable parameters can be convolutional kernels. As another example, any of such an input layer, one or more hidden layers, or output layer can be a dense layer, whose learnable or trainable parameters can be weight matrices or bias values. As yet another example, any of such an input layer, one or more hidden layers, or output layer can be a batch normalization layer, whose learnable or trainable parameters can be translation factors or scaling factors. As yet another example, any of such an input layer, one or more hidden layers, or output layer can be an LSTM layer, whose learnable or trainable parameters can be the input state weight matrix or the hidden state weight matrix. As another example, any of such an input layer, one or more hidden layers, or output layer can be a transformer layer, whose learnable or trainable parameters can be single-head or multi-head attention modules or other weight matrices. Furthermore, in various cases, any of these layers can be any suitable type of neural network layer with any applicable fixed or untrainable internal parameters. For example, any of such an input layer, one or more hidden layers, or output layer can be a non-linear layer, a padding layer, a pooling layer, or a concatenation layer.
[0105] However, these are merely non-limiting examples. In various implementations, the machine learning classifier 802 can employ any other suitable type of artificial intelligence architecture. As a non-limiting example, the machine learning classifier 802 can employ any suitable type of support vector machine architecture. As another non-limiting example, the machine learning classifier 802 can employ any suitable type of Naive Bayes architecture. As yet another non-limiting example, the machine learning classifier 802 can employ any suitable type of linear regression architecture. As yet another non-limiting example, the machine learning classifier 802 can employ any suitable type of logistic regression architecture. As yet another non-limiting example, the machine learning classifier 802 can employ any suitable type of decision tree or random forest architecture. As yet another non-limiting example, the machine learning classifier 802 can employ any suitable combination of the above architectures.
[0106] Regardless of the specific internal architecture implemented within the machine learning classifier 802 (e.g., the specific number, type, or organization of layers), the machine learning classifier 802 can be configured as a discriminator to distinguish between valid and invalid peaks. In other words, the machine learning classifier 802 can be configured to receive any given string or sequence of time-intensity tuples that are suspected to be chromatographic peaks as input and produce a classification label as output, indicating whether the string or sequence of time-intensity tuples truly or actually constitutes a chromatographic peak. Thus, in various respects, the model component 316 can electronically execute the machine learning classifier 802 on each of the multiple suspected peaks 602 to produce multiple validity classification labels 902.
[0107] As a non-limiting example, model component 316 can execute machine learning classifier 802 on the suspected peak 602 (1), and such execution can cause machine learning classifier 802 to produce a valid classification label 902 (1). More specifically, it is assumed that machine learning classifier 802 has a deep learning internal architecture. In various respects, model component 316 can convert the components constituting the suspected peak 602 (1) into... A time-intensity tuple is concatenated and can be fed into the input layer of a machine learning classifier 802. In various instances, the concatenation can perform a forward pass through one or more hidden layers of the machine learning classifier 802. In various cases, the output layer of the machine learning classifier 802 can compute or estimate the validity classification label 902(1) based on any activation map or feature map produced by one or more hidden layers. In various aspects, the validity classification label 902(1) can be any suitable electronic data (e.g., one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof) indicating in a binary or dichotomous manner: time-intensity tuple 602(1)(1) to time-intensity tuple 602(1)(1) ) constitutes a valid, accurate, or correct chromatographic peak; or time-intensity tuple 602(1)(1) to time-intensity tuple 602(1)( This does not constitute a valid, accurate, or correct chromatographic peak. In other words, the machine learning classifier 802 can be considered as evaluating the components of the suspected peak 602(1) (those that look, behave, or otherwise resemble those that the machine learning classifier 802 has learned). Whether any numerical pattern exhibited by a time-intensity tuple characterizes an actual or real chromatographic peak, and whether the validity classification label 902(1) can be considered as a piece of electronic data indicating the result of the evaluation. In other words, peak component 314 can be considered as having identified time-intensity tuple 602(1)(1) to time-intensity tuple 602(1)(1) Together they form a chromatographic peak, and the machine learning classifier 802 can be regarded as a verification of the determination result of peak component 314.
[0108] As another non-limiting example, model component 316 can handle suspected peak 602 ( p ) Execute machine learning classifier 802, and such execution may cause machine learning classifier 802 to produce effective classification labels 902. p More specifically, suppose the machine learning classifier 802 has a deep learning internal architecture. In various aspects, model component 316 can constitute the suspected peak 602 ( p ) of A time-intensity tuple is concatenated and can be fed into the input layer of machine learning classifier 802. In various instances, this concatenation can perform a forward pass through one or more hidden layers of machine learning classifier 802. In various cases, the output layer of machine learning classifier 802 can compute or estimate the effective classification label 902 based on any activation map or feature map produced by one or more hidden layers. p As shown above, the validity classification label is 902. p () can be any suitable electronic data (e.g., one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof), indicated in a binary or dichotomous manner: time-intensity tuple 602 ( p (1) To time-intensity tuple 602 ( p ()( ) constitutes a valid, accurate, or correct chromatographic peak; or time-intensity tuple 602 ( p (1) To time-intensity tuple 602 ( p ()( This does not constitute a valid, accurate, or correct chromatographic peak. In other words, the machine learning classifier 802 can be considered as evaluating the components of the suspected peak 602 ( p (those that look, behave, or otherwise resemble those learned by the machine learning classifier 802) Whether any numerical pattern exhibited by the time-intensity tuple characterizes an actual or real chromatographic peak, and the validity classification label 902 ( p This can be considered as a piece of electronic data indicating the result of the evaluation. In other words, peak component 314 can be considered as a determined time-intensity tuple 602 ( p (1) To time-intensity tuple 602 ( p ()( Together they form a chromatographic peak, and the machine learning classifier 802 can be regarded as a verification of the determination result of peak component 314.
[0109] In all respects, validity classification label 902(1) to validity classification label 902( p These can be considered as forming multiple validity classification labels 902.
[0110] It should be noted that in some implementations, the machine learning classifier 802 may be configured to receive additional, supplementary, auxiliary, or complementary inputs in addition to the suspected peak. In fact, for any given suspected peak, the machine learning classifier 802 may be configured to receive not only the given suspected peak as input, but also any suitable numerical properties or attributes associated with the given suspected peak as input. As a non-limiting example, consider suspected peak 602(1). In some cases, the machine learning classifier 802 may receive not only the components of suspected peak 602(1) A time-intensity tuple, and can also receive: the vertex of the suspected peak 602 (1) (e.g., any peak detection technique used by peak component 314 can...). One of the time-intensity tuples is marked, identified, or marked as the vertex or highest point of its identified suspected peak 602(1); the starting point of the suspected peak 602(1) (e.g., any peak detection technique used by peak component 314 can identify the peak). One of the time-intensity tuples is marked, identified, or labeled as indicating the leading edge of the suspected peak 602(1) it identifies; the endpoint of the suspected peak 602(1) (e.g., any peak detection technique used by peak component 314 can identify the leading edge of the suspected peak 602(1)). A time-intensity tuple in a time-intensity tuple is marked, identified, or indicated as indicating the trailing edge of its identified suspected peak 602(1); the width of the suspected peak 602(1) (e.g., the length of time spanned by any peak detection technique used by peak assembly 314); the height of the suspected peak 602(1) (e.g., the amplitude of the identified suspected peak 602(1) exceeding the baseline detector signal by any peak detection technique used by peak assembly 314); or an alphanumeric identifier uniquely associated with the hardware or model of the mass spectrometer 302 equipped with the chromatograph.
[0111] In any case, the valid peak set 804 may include any suspected peak whose validity classification label indicates it is valid from among the multiple suspected peaks 602. Conversely, the invalid peak set 806 may include any suspected peak whose validity classification label indicates it is invalid from among the multiple suspected peaks 602.
[0112] In various implementations, execution component 318 may perform any suitable mass spectrometry analysis (e.g., any suitable metabolomics, proteomics, statistical, or computational processing flow or algorithm) electronically on any mass spectrum of the plurality of mass spectra 404 corresponding to the valid peak set 804 (e.g., any time-intensity tuple corresponding to any peak belonging to any peak in the valid peak set 804). However, conversely, execution component 318 may avoid performing any such mass spectrometry analysis electronically on any mass spectrum of the plurality of mass spectra 404 corresponding to the invalid peak set 806 (e.g., any time-intensity tuple corresponding to any peak belonging to any peak in the invalid peak set 806). In some cases, execution component 318 may therefore be considered to ignore, disregard, or discard any mass spectrum of the plurality of mass spectra 404 corresponding to a time-intensity tuple that was incorrectly detected as a peak by peak component 314. Therefore, no time or computational resources are wasted on further or downstream analysis or processing of the mass spectra of such incorrectly detected peaks.
[0113] To reliably perform the post-detection peak validation described herein with the machine learning classifier 802 employing a deep learning internal architecture, the machine learning classifier 802 may first undergo training. Non-limiting examples of such training are described with reference to Figures 10 through 12.
[0114] Figure 10 shows a block diagram of an exemplary non-limiting system comprising a training component and a training dataset that can facilitate post-detection chromatographic peak validation according to one or more embodiments described herein.
[0115] In various implementations, one or more software components 311 may include a training component 1002. In various aspects, the training component 1002 may electronically train a machine learning classifier 802 using a training dataset 1004.
[0116] Figure 11 shows an exemplary non-limiting block diagram of a training dataset 1004 according to one or more embodiments described herein.
[0117] In various implementations, the training dataset 1004 may include a training peak set 1102. In various aspects, the training peak set 1102 may have a total of m One peak m For any suitable positive integer: training peak 1102(1) to training peak 1102( mIn various instances, each of the multiple training peaks 1102 may be a different or corresponding string or sequence of time-intensity tuples that has been (potentially incorrectly) identified as constituting a chromatographic peak. In various cases, any training peak in the set 1102 may be obtained from any suitable chromatographic apparatus (e.g., even one different from the mass spectrometer 302 equipped with a chromatograph) for or for any suitable sample (e.g., even one different from sample 304).
[0118] In various cases, the training dataset 1004 may further include a set of true classification labels 1104. In each respect, the set of true classification labels 1104 may correspond (e.g., in a one-to-one manner) to a training peak set 1102. Therefore, since the training peak set 1102 can have... m Each peak, a real category tag set of 1104, can have m One label: True category label 1104(1) to True category label 1104( m In various instances, each true classification label in the true classification label set 1104 can be considered a correct or accurate validity classification label that is known or considered to correspond to a corresponding training peak in the training peak set 1102. As a non-limiting example, true classification label 1104(1) can correspond to training peak 1102(1). Therefore, true classification label 1104(1) can be any suitable electronic data (e.g., having the same size, format, or dimensions as any of the multiple validity classification labels 902) that correctly or accurately indicates whether training peak 1102(1) is a valid or invalid chromatographic peak. As another non-limiting example, true classification label 1104( m This can correspond to training peak 1102. m Therefore, the true classification label is 1104. m This can be any suitable electronic data that correctly or accurately indicates training peak 1102. m Is it a valid or invalid chromatographic peak?
[0119] Figure 12 shows an exemplary non-limiting block diagram illustrating how a machine learning classifier 802 can be trained according to one or more embodiments described herein.
[0120] In all respects, before training begins, the trainable intrinsic parameters of the machine learning classifier 802 (e.g., convolutional kernel, weight matrix, bias values) can be initialized by the training component 1002 in any suitable manner (e.g., via random initialization).
[0121] In various implementations, the training component 1002 can select any suitable training peak and corresponding true classification label from the training dataset 1004. These can be referred to as training peak 1202 and true classification label 1204, respectively.
[0122] In various aspects, training component 1002 can enable machine learning classifier 802 to execute on training peak 1202, thereby enabling machine learning classifier 802 to produce output 1206. More specifically, in some cases, training peak 1202 can be fed or routed to the input layer of machine learning classifier 802, training peak 1202 can be forward-passed through one or more hidden layers of machine learning classifier 802, and the output layer of machine learning classifier 802 can compute output 1206 based on activation maps or feature maps provided by one or more hidden layers of machine learning classifier 802.
[0123] It should be noted that the format, size, or dimension of output 1206 can be determined by the number, arrangement, size, or other characteristics of neurons, convolutional kernels, attention modules, or other internal parameters of the output layer (or any other layer) of the machine learning classifier 802. Therefore, by adding, deleting, or otherwise adjusting the characteristics of the output layer (or any other layer) of the machine learning classifier 802, output 1206 can be forced to have any desired format, size, or dimension.
[0124] In all respects, output 1206 can be considered a valid classification label synthesized by machine learning classifier 802 based on training peak 1202, either a prediction or an inference. Conversely, the true classification label 1204 can be considered any correct or accurate valid classification label known or regarded as corresponding to training peak 1202. It should be noted that if machine learning classifier 802 has not been trained or has been trained very little to date, output 1206 may be highly inaccurate. In other words, output 1206 may be very different from the true classification label 1204.
[0125] In various aspects, the training component 1002 can compute a loss 1208 (e.g., mean absolute error, mean squared error, cross-entropy error) between the output 1206 and the true classification label 1204. In various instances, the training component 1002 can progressively update the trainable intrinsic parameters of the machine learning classifier 802 based on the loss 1208 via backpropagation (e.g., stochastic gradient descent).
[0126] In various cases, this execution and update procedure can be repeated for any suitable number of training peaks (e.g., for each training peak in the training dataset 1004). This may ultimately lead to iterative optimization of the trainable intrinsic parameters of the machine learning classifier 802 to accurately distinguish or differentiate between valid and invalid chromatographic peaks. In all aspects, any suitable training batch size, any suitable error / loss function, or any suitable training termination criterion can be used during such training.
[0127] While this document primarily describes the supervised training of the machine learning classifier 802, this is merely a non-limiting example for ease of explanation and illustration. In various implementations, any other suitable training paradigm can be used to train the machine learning classifier 802, such as unsupervised training, semi-supervised training, or reinforcement learning, any of which can employ federated or non-federated architectures.
[0128] Figure 13 shows exemplary non-limiting experimental results according to one or more embodiments described herein.
[0129] Figure 1302 shows a time-intensity tuple string or sequence (represented by a continuous solid line) that has been identified as a chromatographic peak (e.g., represented by a shaded bell curve) using curve fitting peak detection techniques. An implementation of machine learning classifier 802 is performed on the time-intensity tuple string or sequence shown in figure 1302, and such execution produces a valid classification label indicating validity. It should be noted that this is reasonable because the continuous solid line matches the shaded bell curve very well.
[0130] Number 1304 shows a time-intensity tuple string or sequence (represented by a continuous solid line) that has been identified as a chromatographic peak (e.g., represented by a shaded bell curve) using curve fitting peak detection techniques. An implementation of machine learning classifier 802 is performed on the time-intensity tuple string or sequence shown in number 1304, and such execution produces a validity classification label indicating invalidity. It should be noted that this is reasonable because the continuous solid line is extremely noisy and does not match the shaded bell curve very well.
[0131] Number 1306 illustrates two separate time-intensity tuple strings or sequences that have been identified as chromatographic peaks using curve-fitting peak detection techniques. An implementation of machine learning classifier 802 is performed on both of these time-intensity tuple strings or sequences, and such execution produces two valid classification labels indicating invalidity.
[0132] These experimental results help demonstrate how to effectively implement the various implementation schemes described in this paper to distinguish between valid and invalid peaks, so as not to waste time and resources on analyzing the mass spectra of invalid peaks.
[0133] Although the various embodiments described herein refer to chromatograms as intensity sequences organized according to retention time, these are merely non-limiting examples for ease of interpretation and illustration. In all aspects, retention times in any embodiment can be readily replaced with unitless retention indices as appropriate.
[0134] In various instances, machine learning algorithms or models can be implemented in any suitable manner to facilitate any suitable aspect described herein. To facilitate some of the aforementioned machine learning aspects of various implementations, consider the following discussion of artificial intelligence (AI). The various implementations described herein may employ artificial intelligence to facilitate the automation of one or more features or functions. These components may employ various AI-based schemes to perform the various implementations / examples disclosed herein. To provide or assist the numerous determinations described herein (e.g., determining, ascertaining, inferring, calculating, predicting, estimating, deducing, predicting, detecting, computing), the components described herein may examine the entirety or a subset of the data authorized to their access and may provide inference about the system or environment or determine its state from a set of observations captured, such as by events or data. For example, determining a probability distribution that can be used to identify a specific context or action, or to generate states. Determination can be probabilistic; that is, calculating the probability distribution of states of interest based on considerations of data and events. Determination can also refer to techniques used to compose higher-level events from a set of events or data.
[0135] Such determinations can lead to the construction of new events or actions from a set of observed events or stored event data, regardless of whether the events are closely related in time or whether the events and data originate from one or more event and data sources. The components disclosed herein can employ various classification schemes or systems (e.g., support vector machines, neural networks, expert systems, Bayesian belief networks, fuzzy logic, data fusion engines, etc.) that can be explicitly trained (e.g., via training data) or implicitly trained (e.g., via observed behavior, preferences, historical information, received external information, etc.) for performing automated or deterministic operations relevant to the subject matter of the claims. Therefore, classification schemes or systems can be used to automatically learn and perform multiple functions, actions, or determinations.
[0136] The classifier can take the input attribute vector z = (z1, z2, z3, z4, ... z n This maps the input to the confidence level of belonging to a certain category, such as based on f(z) = confidence ( class This classification can employ probabilistic analysis or statistical analysis (e.g., considering utility and cost in the analysis) to determine the actions that should be automated. Support Vector Machines (SVMs) are a possible example of a classifier that can be used. SVMs operate by finding a hypersurface in the space of possible inputs, where the hypersurface attempts to separate triggering conditions from non-triggering events. Intuitively, this makes the classification correct for test data that is close to but not exactly the same as the training data. Other directed and non-directed model classification methods include, for example, Naive Bayes, Bayesian networks, decision trees, neural networks, fuzzy logic models, or probabilistic classification models that provide different independent patterns, any of which can be employed. Classification as used in this paper also includes statistical regression for developing prioritization models.
[0137] To provide additional context for the various implementations described herein, Figure 14 and the following discussion are intended to briefly and generally describe suitable computing environments 1400 in which the various implementations described herein can be implemented. Although the implementations have been described above in the general context of computer-executable instructions that can run on one or more computers, those skilled in the art will recognize that these implementations can also be implemented in combination with other program modules or as a combination of hardware and software.
[0138] Typically, program modules include routines, programs, components, data structures, etc., that perform specific tasks or implement specific abstract data types. Furthermore, those skilled in the art will understand that the method of this invention can be implemented in conjunction with other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, Internet of Things (IoT) devices, distributed computing systems, and personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, each of which can be operatively coupled to one or more associated devices.
[0139] The implementation schemes shown in this paper can also be implemented in a distributed computing environment, where certain tasks are performed by remote processing devices linked via a communication network. In a distributed computing environment, program modules can reside on both local and remote memory storage devices.
[0140] Computing devices typically include various media, which may include computer-readable storage media, machine-readable storage media, or communication media. These two terms are used differently from each other herein. A computer-readable storage media or a machine-readable storage media can be any available storage medium accessible to a computer, including volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, a computer-readable storage media or a machine-readable storage media can be implemented in conjunction with any information storage method or technology, such as computer-readable or machine-readable instructions, program modules, structured data, or unstructured data.
[0141] Computer-readable storage media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc storage (CD-ROM), digital versatile optical disc (DVD), Blu-ray disc (BD) or other optical disc storage, cassette tape, magnetic tape, disk storage or other magnetic storage devices, solid-state drives or other solid-state storage devices, or other tangible and / or non-transient media that can be used to store desired information. In this regard, the terms “tangible” or “non-transient” used herein to describe storage devices, memories, or computer-readable media should be understood to exclude only the propagation of transient signals themselves as a modifier, and do not waive the rights of all standard storage devices, memories, or computer-readable media that are not merely about the propagation of transient signals themselves.
[0142] Computer-readable storage media can be accessed by one or more local or remote computing devices, for example, via access requests, queries or other data retrieval protocols, to perform various operations on the information stored in the media.
[0143] Communication media typically embody computer-readable instructions, data structures, program modules, or other structured or unstructured data in data signals (such as modulated data signals, for example, carrier waves or other transmission mechanisms), and include any information transmission or delivery medium. The term "modulated data signal" or signal refers to a signal whose one or more characteristics are set or altered in such a way that information is encoded in one or more signals. By way of example and not limitation, communication media include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, RF, infrared, and other wireless media).
[0144] Referring again to Figure 14, an example environment 1400 for implementing various embodiments of the aspects described herein includes a computer 1402, which includes a processing unit 1404, system memory 1406, and a system bus 1408. The system bus 1408 couples system components (including, but not limited to, system memory 1406) to the processing unit 1404. The processing unit 1404 can be any of a variety of commercially available processors. Dual microprocessors and other multi-processor architectures can also be used as the processing unit 1404.
[0145] System bus 1408 can be any of a variety of bus architectures, which can be further interconnected with memory buses (with or without memory controllers), peripheral buses, and local buses using any of a variety of commercially available bus architectures. System memory 1406 includes ROM 1410 and RAM 1412. The basic input / output system (BIOS) can be stored in non-volatile memory, such as ROM, erasable programmable read-only memory (EPROM), or EEPROM, where the BIOS contains basic routines that facilitate the transfer of information between internal components of computer 1402, such as during startup. RAM 1412 may also include high-speed RAM, such as static RAM for caching data.
[0146] Computer 1402 also includes an internal hard disk drive (HDD) 1414 (e.g., EIDE, SATA), one or more external storage devices 1416 (e.g., floppy disk drive (FDD) 1416, memory stick or flash drive reader, memory card reader, etc.), and a drive 1420, such as a solid-state drive or optical disc drive, which can read from or write to disk 1422 (e.g., CD-ROM, DVD, BD, etc.). Alternatively, in cases involving solid-state drives, disk 1422 is not included unless provided separately. Although the internal HDD 1414 is shown as being located inside computer 1402, the internal HDD 1414 can also be configured for external use in a suitable chassis (not shown). Furthermore, although not shown in environment 1400, a solid-state drive (SSD) can be used to supplement or replace HDD 1414. HDD 1414, external storage device 1416, and drive 1420 can be connected to system bus 1408 via HDD interface 1424, external storage interface 1426, and drive interface 1428, respectively. Interface 1424 for the external drive implementation may include at least one or both of Universal Serial Bus (USB) and Institute of Electrical and Electronics Engineers (IEEE) 1394 interface technologies. Other external drive connection technologies are also within the scope of the embodiments described herein.
[0147] The drive and its associated computer-readable storage medium provide non-volatile storage of data, data structures, computer-executable instructions, etc. For computer 1402, its drive and storage medium can support the storage of any data in a suitable digital format. Although the description of computer-readable storage media above refers to corresponding types of storage devices, those skilled in the art will understand that other types of computer-readable storage media, whether existing or developed in the future, can also be used in the example operating environment, and further, any such storage medium can contain computer-executable instructions for performing the methods described herein.
[0148] Numerous program modules can be stored in the drive and RAM 1412, including the operating system 1430, one or more application programs 1432, other program modules 1434, and program data 1436. All or part of the operating system, application programs, modules, or data may also be cached in RAM 1412. The systems and methods described herein can be implemented using various commercially available operating systems or combinations of operating systems.
[0149] Computer 1402 may optionally include emulation technology. For example, a virtual machine monitor (not shown) or other intermediary may emulate the hardware environment of operating system 1430, and the emulated hardware may optionally be different from the hardware shown in Figure 14. In such embodiments, operating system 1430 may include one of a plurality of virtual machines (VMs) hosted on computer 1402. Furthermore, operating system 1430 may provide a runtime environment for application 1432, such as the Java Runtime Environment or the .NET Framework. A runtime environment is a consistent execution environment that allows application 1432 to run on any operating system that includes that runtime environment. Similarly, operating system 1430 may support containers, and application 1432 may exist as a container, which is a lightweight, standalone, executable software package that includes, for example, application code, runtime, system tools, system libraries, and settings.
[0150] Furthermore, computer 1402 may be equipped with a security module, such as a Trusted Processing Module (TPM). For example, with the help of a TPM, the boot component hashes the next boot component over time and waits for the result to match a security value before loading the next boot component. This process can occur at any level of the computer 1402's code execution stack, for example, at the application execution level or the operating system (OS) kernel level, thus achieving security at any level of code execution.
[0151] Users can input commands and information to computer 1402 through one or more wired / wireless input devices (e.g., keyboard 1438, touchscreen 1440, and positioning devices such as mouse 1442). Other input devices (not shown) may include microphones, infrared (IR) remote controls, radio frequency (RF) remote controls or other remote controls, joysticks, virtual reality controllers or virtual reality headsets, game controllers, styluses, image input devices (e.g., cameras), gesture sensor input devices, visual motion sensor input devices, emotion or face detection devices, biometric input devices (e.g., fingerprint or iris scanners), etc. These and other input devices are typically connected to processing unit 1404 via input device interface 1444, which may be coupled to system bus 1408, but may also be connected via other interfaces such as parallel ports, IEEE 1394 serial ports, game ports, USB ports, infrared interfaces, BLUETOOTH® interfaces, etc.
[0152] Monitor 1446 or other types of display devices may also be connected to system bus 1408 via an interface such as video adapter 1448. In addition to monitor 1446, computers typically include other peripheral output devices (not shown), such as speakers, printers, etc.
[0153] Computer 1402 can operate in a networked environment that uses logical connections to one or more remote computers (such as remote computer 1450) via wired or wireless communication. Remote computer 1450 can be a workstation, server computer, router, personal computer, laptop computer, microprocessor-based entertainment device, peer-to-peer device, or other common network node, and typically includes many or all of the elements described relative to computer 1402; however, for simplicity, only memory / storage device 1452 is shown. The logical connections shown include wired / wireless connections to a local area network (LAN) 1454 or a larger network (e.g., a wide area network (WAN) 1456). Such LAN and WAN networking environments are common in offices and companies and are advantageous for enterprise-wide computer networks (such as intranets), all of which can connect to global communication networks (e.g., the Internet).
[0154] When used in a LAN networking environment, computer 1402 can connect to local network 1454 via a wired or wireless communication network interface or adapter 1458. Adapter 1458 can facilitate wired or wireless communication with LAN 1454, which may also include a wireless access point (AP) configured thereon to communicate with adapter 1458 in wireless mode.
[0155] When used in a WAN networking environment, computer 1402 may include modem 1460, or may otherwise connect to a communication server on WAN 1456 to establish communication via WAN 1456 (such as via the Internet). Modem 1460 may be built-in or external, and may be a wired or wireless device, connected to system bus 1408 via input device interface 1444. In a networking environment, program modules shown relative to computer 1402 or portions thereof may be stored in remote memory / storage device 1452. It should be understood that the network connection shown is an example, and other methods of establishing inter-computer communication links may be used.
[0156] When used in a LAN or WAN networking environment, computer 1402 can access cloud storage systems or other network-based storage systems, such as, but not limited to, network virtual machines that provide one or more aspects of information storage or processing, in addition to or as an alternative to external storage device 1416 as described above. Typically, the connection between computer 1402 and the cloud storage system can be established via LAN 1454 or WAN 1456, for example, via adapter 1458 or modem 1460. When computer 1402 is connected to an associated cloud storage system, external storage interface 1426 can manage the storage provided by the cloud storage system with the help of adapter 1458 or modem 1460, just as it would manage other types of external storage. For example, external storage interface 1426 can be configured to provide access to cloud storage sources as if these sources were physically connected to computer 1402.
[0157] Computer 1402 is operable to communicate with any wireless device or entity operable in a wireless communication manner (e.g., printer, scanner, desktop or portable computer, portable data assistant, communications satellite, any equipment or location associated with a wirelessly detectable tag (e.g., kiosk, newsstand, store shelf, etc.) and telephone). This can include Wi-Fi and BLUETOOTH® wireless technologies. Therefore, communication can be a predefined structure like a traditional network, or simply self-organizing communication between at least two devices.
[0158] Figure 15 is a schematic block diagram of an example computing environment 1500 that can interact with the subject matter of this disclosure. The example computing environment 1500 includes one or more clients 1510. Clients 1510 can be hardware or software (e.g., threads, processes, computing devices). The example computing environment 1500 also includes one or more servers 1530. Servers 1530 can also be hardware or software (e.g., threads, processes, computing devices). For example, server 1530 can accommodate threads for conversion using one or more embodiments as described herein. One possible form of communication between client 1510 and server 1530 may be data packets suitable for transmission between two or more computer processes. The example computing environment 1500 includes a communication framework 1550 that can be used to facilitate communication between client 1510 and server 1530. Client 1510 is operatively connected to one or more client data repositories 1520, which can be used to store information local to client 1510. Similarly, server 1530 is operatively connected to one or more server data repositories 1540, which can be used to store information locally on server 1530.
[0159] Various implementations can be systems, methods, apparatus, or computer program products at any possible level of technical detail integration. A computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon to cause a processor to execute aspects of various implementations. A computer-readable storage medium can be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media may also include: portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only storage media (CD-ROM), digital versatile optical disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punched cards or recessed protrusions on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium must not be interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical cables), or electrical signals transmitted through metallic wires.
[0160] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device or downloaded via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network) to an external computer or external storage device. This network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the corresponding computing / processing device. The computer-readable program instructions used to perform operations in various implementation schemes may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) or procedural programming languages (such as the "C" programming language or similar programming languages). Computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet through an Internet service provider). In some implementations, electronic circuitry, including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can be personalized by utilizing the state information of the computer-readable program instructions to execute the computer-readable program instructions, intended to perform various functions.
[0161] This document describes various aspects with reference to flowchart illustrations or block diagrams of methods, apparatus (systems), and computer program products according to various embodiments. It should be understood that each block in the flowchart illustration or block diagram, and combinations of blocks in the flowchart illustration or block diagram, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create a machine, thereby creating a manner for implementing the functions / actions specified in the flowchart or block diagram blocks. These computer-readable program instructions can also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, or other device to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored constitutes an article of manufacture containing instructions for implementing aspects of the functions / actions specified in the flowchart or block diagram blocks. Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operations to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, thereby implementing the functions / actions specified in the flowchart or block diagram blocks or blocks as executed on the computer, other programmable apparatus or other device.
[0162] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may not occur in the order shown in the figures. For example, two consecutively shown blocks may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented by a system based on dedicated hardware, which performs the specified function or action, or implements a combination of dedicated hardware and computer instructions.
[0163] While the subject matter has been described above in the general context of computer-executable instructions for a computer program product running on one or more computers, those skilled in the art will recognize that this disclosure can also be implemented in conjunction with other program modules. Typically, program modules include routines, programs, components, data structures, etc., that perform specific tasks or implement specific abstract data types. Furthermore, those skilled in the art will recognize that various aspects can be implemented in conjunction with other computer system configurations, including single-processor or multi-processor computer systems, small computing devices, mainframe computers, handheld computing devices (e.g., PDAs, telephones), or microprocessor-based or programmable consumer and / or industrial electronic products. The aspects shown can also be implemented in a distributed computing environment, where tasks are performed by remote processing devices linked via a communication network. However, some, if not all, aspects of this disclosure can be implemented on a standalone computer. In a distributed computing environment, program modules can reside on both local and remote memory storage devices.
[0164] As used herein, the terms “component,” “system,” “platform,” “interface,” etc., may refer to or include computer-related entities or entities associated with an operational machine having one or more specific functions. Entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, or a computer. As an example, an application running on a server and the server itself can both be components. One or more components may reside in a process or execution thread, and components may be located on a single computer or distributed among two or more computers. In another example, a corresponding component may be executed from various computer-readable media on which various data structures are stored. These components may communicate, for example, via local or remote processes based on signals having one or more data packets (e.g., data from one component interacts with another component in a local system, a distributed system, or with other systems via a network such as the Internet). As another example, a component may be a device having specific functions provided by mechanical parts operated by electrical or electronic circuitry, the device being operated by a software or firmware application executed by a processor. In this case, the processor may be located inside or outside the device and may execute at least a portion of the software or firmware application. As yet another example, a component can be a device that provides a specific function through electronic components without mechanical parts, wherein the electronic components may include a processor or other means for executing software or firmware that at least partially endows the electronic components with function. In one aspect, a component can be simulated by, for example, a virtual machine within a cloud computing system.
[0165] Furthermore, the term “or” is intended to mean inclusive “or” rather than exclusive “or”. That is, unless otherwise specified or clearly apparent from the context, “X adopts A or B” means any naturally inclusive permutation. That is, if X adopts A; X adopts B; or X adopts both A and B, then “X adopts A or B” is satisfied in any of the foregoing cases. As used herein, the term “and / or” is intended to have the same meaning as “or”. Furthermore, unless otherwise specified or clearly apparent from the context involving the singular form, the article “a” as used in the subject matter description and accompanying drawings should generally be interpreted as meaning “one or more”. As used herein, the terms “example” or “exemplary” are used to indicate that something is used as an example, instance, or illustration. For the avoidance of doubt, the subject matter disclosed herein is not limited to such examples. Moreover, any aspect or design described herein as an “example” or “exemplary” is not necessarily construed as preferred or superior to other aspects or designs, nor does it imply exclusion of equivalent exemplary structures and techniques known to those skilled in the art.
[0166] The disclosure herein describes non-limiting examples. For ease of description or explanation, the terms “each,” “every,” or “all” are used in discussions of various examples throughout the disclosure. Such use of the terms “each,” “every,” or “all” is not restrictive. In other words, when the disclosure herein provides a description of “each,” “every,” or “all” applicable to a particular object or component, it should be understood that this is a non-limiting example, and it should also be understood that in various other examples, such a description may apply to fewer than “each,” “all,” or “all” of that particular object or component.
[0167] As used herein, the term "processor" can refer to substantially any computing processing unit or device, including but not limited to a single-core processor; a single processor with software multithreading capabilities; a multi-core processor; a multi-core processor with software multithreading capabilities; a multi-core processor with hardware multithreading technology; a parallel platform; or a parallel platform with distributed shared memory. Furthermore, a processor can refer to an integrated circuit, application-specific integrated circuit (ASIC), digital signal processor (DSP), field-programmable gate array (FPGA), programmable logic controller (PLC), complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. Additionally, processors can utilize nanoscale architectures, such as, but not limited to, molecular and quantum dot-based transistors, switches, or gates, to optimize space utilization or enhance the performance of user equipment. Processors can also be implemented as a combination of computing processing units. In this disclosure, terms such as "storage," "memory," "database," and any other information storage component related to the operation and function of a component are used to refer to a "memory component," an entity embodied in "memory," or a component containing memory. It should be understood that the memory or memory component described herein may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. By way of illustration and not limitation, non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). Volatile memory may include RAM, for example, which may act as external cache memory. By way of illustration and not limitation, RAM may be provided in a variety of forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), or Rambus dynamic RAM (RDRAM). Furthermore, the memory components of the systems or computer implementation methods disclosed herein are intended to include, but are not limited to, these or any other suitable types of memory.
[0168] The foregoing description includes only examples of systems and computer-implemented methods. Of course, it is impossible to describe every conceivable combination of components or computer-implemented methods for the purposes of describing this disclosure, but many further combinations and arrangements of the contents of this disclosure are possible. Furthermore, the terms “comprising,” “having,” “possessing,” etc., are used to such an extent in the detailed description, claims, appendices, or drawings that such terms are intended to be inclusive in a manner similar to the term “comprising,” as interpreted when “comprising” is used as a transitional word in the claims.
[0169] Descriptions of various embodiments have been presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to existing technologies in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0170] Various non-limiting aspects are described in the following embodiments.
[0171] Example 1: The system may include a processor capable of executing computer-executable components stored in a non-transitory computer-readable storage memory, wherein the computer-executable components may include: a scanning component capable of enabling a chromatographic device coupled to a mass spectrometer to scan a sample, thereby generating a chromatogram and a mass spectrum; a peak component capable of identifying multiple suspected peaks in the chromatogram via a peak detection algorithm; a model component capable of separating multiple suspected peaks into a valid peak set and an invalid peak set by performing a machine learning classifier on the corresponding suspected peaks among the multiple suspected peaks; and an execution component capable of performing mass spectrometry analysis on a first portion of the mass spectrum corresponding to the valid peak set, but not on a second portion of the mass spectrum corresponding to the invalid peak set.
[0172] Example 2: A system that can be implemented in any of the foregoing embodiments, wherein, for a first suspected peak among a plurality of suspected peaks, the model component can feed the first suspected peak or one or more properties of the first suspected peak as input to a machine learning classifier, and wherein the machine learning classifier can generate a classification label as output, the classification label indicating whether the first suspected peak is a valid peak or an invalid peak.
[0173] Example 3: A system that can be implemented in any of the foregoing embodiments, wherein one or more properties of the first suspected peak may include a first time-intensity tuple representing the vertex of the first suspected peak.
[0174] Example 4: A system that can implement any of the foregoing embodiments, wherein one or more properties of the first suspected peak may further include: a second time-intensity tuple representing the start point of the first suspected peak; or a third time-intensity tuple representing the end point of the first suspected peak.
[0175] Example 5: A system that can be implemented in any of the foregoing embodiments, wherein one or more properties of the first suspected peak may further include the width of the first suspected peak.
[0176] Example 6: A system that can be implemented in any of the foregoing embodiments, wherein one or more properties of the first suspected peak may further include the height of the first suspected peak.
[0177] Example 7: A system that can be implemented in any of the foregoing embodiments, wherein one or more properties of the first suspected peak may further include a hardware identifier associated with the chromatographic device.
[0178] Example 8: A system that can implement any of the foregoing embodiments, wherein the computer-executable component may further include: a training component that can train a machine learning classifier based on a training dataset, wherein the training dataset may include: multiple training peaks; and multiple true classification labels corresponding to the multiple training peaks, each of the multiple true classification labels indicating whether the corresponding training peak is valid or invalid in a binary manner.
[0179] In various implementation schemes, any one or more combinations of Examples 1 to 8 may be implemented.
[0180] Example 9: A computer-implemented method may include: scanning a sample with a chromatographic device coupled to a mass spectrometer via a device operably coupled to a processor, thereby generating a chromatogram and a mass spectrum; identifying multiple suspected peaks in the chromatogram via the device and via a peak detection algorithm; separating the multiple suspected peaks into a valid peak set and an invalid peak set via the device and via performing a machine learning classifier on the corresponding suspected peaks among the multiple suspected peaks; and performing mass spectrometry analysis on a first portion of the mass spectrum corresponding to the valid peak set, but not on a second portion of the mass spectrum corresponding to the invalid peak set, via the device.
[0181] Example 10: A computer-implemented method that can be implemented in any of the foregoing embodiments, wherein, for a first suspected peak among a plurality of suspected peaks, the device can feed the first suspected peak or one or more properties of the first suspected peak as input to a machine learning classifier, and wherein the machine learning classifier can generate a classification label as output, the classification label indicating whether the first suspected peak is a valid peak or an invalid peak.
[0182] Example 11: A computer-implemented method of any of the foregoing embodiments can be implemented, wherein one or more properties of the first suspected peak may include a first time-intensity tuple representing the vertex of the first suspected peak.
[0183] Example 12: A computer-implemented method that can be implemented in any of the foregoing embodiments, wherein one or more properties of the first suspected peak may further include: a second time-intensity tuple representing the start point of the first suspected peak; or a third time-intensity tuple representing the end point of the first suspected peak.
[0184] Example 13: A computer-implemented method of any of the foregoing embodiments can be implemented, wherein one or more properties of the first suspected peak may further include the width of the first suspected peak.
[0185] Example 14: A computer-implemented method of any of the foregoing embodiments can be implemented, wherein one or more properties of the first suspected peak may further include the height of the first suspected peak.
[0186] Example 15: A computer-implemented method of any of the foregoing embodiments can be implemented, wherein one or more properties of the first suspected peak may further include a hardware identifier associated with the chromatographic apparatus.
[0187] Example 16: A computer-implemented method that can implement any of the foregoing embodiments, the computer-implemented method further comprising: training a machine learning classifier based on a training dataset using the device, wherein the training dataset may include: multiple training peaks; and multiple true classification labels corresponding to the multiple training peaks, each of the multiple true classification labels indicating whether the corresponding training peak is valid or invalid in a binary manner.
[0188] In various implementation schemes, any one or more combinations of Examples 9 to 16 can be implemented.
[0189] Example 17: A computer program product for facilitating post-detection chromatographic peak validation may include a non-transitory computer-readable storage device having program instructions embodied thereon. In various aspects, the program instructions may be executed by a processor to cause the processor to: cause a chromatographic device coupled to a mass spectrometer to scan a sample, thereby generating a chromatogram and a mass spectrum; identify multiple suspected peaks in the chromatogram using a peak detection algorithm; separate the multiple suspected peaks into a valid peak set and an invalid peak set by performing a machine learning classifier on the corresponding suspected peaks among the multiple suspected peaks; and perform mass spectrometry analysis on a first portion of the mass spectrum corresponding to the valid peak set, but not on a second portion of the mass spectrum corresponding to the invalid peak set.
[0190] Example 18: A computer program product that can implement any of the foregoing embodiments, wherein, for a first suspected peak among a plurality of suspected peaks, the processor can feed the first suspected peak or one or more properties of the first suspected peak as input to a machine learning classifier, and wherein the machine learning classifier can generate a classification label as output, the classification label indicating whether the first suspected peak is a valid peak or an invalid peak.
[0191] Example 19: A computer program product that can implement any of the foregoing embodiments, wherein one or more properties of the first suspected peak may include a first time-intensity tuple representing the vertex of the first suspected peak.
[0192] Example 20: A computer program product that can implement any of the foregoing embodiments, wherein the program instructions can be further executed to cause the processor to: train a machine learning classifier based on a training dataset, wherein the training dataset may include: multiple training peaks; and multiple true classification labels corresponding to the multiple training peaks, each of the multiple true classification labels indicating whether the corresponding training peak is valid or invalid in a binary manner.
[0193] In various implementation schemes, any one or more combinations of Examples 17 to 20 can be implemented.
[0194] In various implementation schemes, any one or more combinations of Examples 1 to 20 can be implemented.
Claims
1. A system comprising: A processor that executes a computer-executable component stored in a non-transitory computer-readable storage memory, wherein the computer-executable component includes: A scanning component that enables chromatographic equipment coupled to a mass spectrometer to scan a sample, thereby generating chromatograms and mass spectra; A peak component, wherein the peak component identifies multiple suspected peaks in the chromatogram using a peak detection algorithm; A model component, which separates the plurality of suspected peaks into a set of valid peaks and a set of invalid peaks by performing a machine learning classifier on the corresponding suspected peaks among the plurality of suspected peaks; and An execution component performs mass spectrometry analysis on a first portion of the mass spectrum corresponding to the set of valid peaks, but does not perform mass spectrometry analysis on a second portion of the mass spectrum corresponding to the set of invalid peaks.
2. The system according to claim 1, wherein, For the first suspected peak among the plurality of suspected peaks, the model component feeds the first suspected peak or one or more properties of the first suspected peak as input to the machine learning classifier, wherein the machine learning classifier produces a classification label as output, the classification label indicating whether the first suspected peak is a valid peak or an invalid peak.
3. The system of claim 2, wherein the one or more properties of the first suspected peak include a first time-intensity tuple representing the vertex of the first suspected peak.
4. The system of claim 3, wherein the one or more properties of the first suspected peak further include: The second time-intensity tuple represents the starting point of the first suspected peak; or The third time-intensity tuple represents the endpoint of the first suspected peak.
5. The system of claim 3, wherein the one or more properties of the first suspected peak further include the width of the first suspected peak.
6. The system of claim 3, wherein the one or more properties of the first suspected peak further include the height of the first suspected peak.
7. The system of claim 3, wherein the one or more properties of the first suspected peak further include a hardware identifier associated with the chromatographic apparatus.
8. The system of claim 1, wherein the computer-executable component further comprises: Training component, which trains the machine learning classifier based on a training dataset, wherein the training dataset includes: Multiple training peaks; and Each of the multiple training peaks has a corresponding real classification label, and each real classification label indicates whether the corresponding training peak is valid or invalid in a binary manner.
9. A computer-implemented method, the computer-implemented method comprising: A chromatographic device coupled to a mass spectrometer is operatively coupled to a processor to scan a sample, thereby generating chromatograms and mass spectra. The device and a peak detection algorithm are used to identify multiple suspected peaks in the chromatogram. The device separates the plurality of suspected peaks into a set of valid peaks and a set of invalid peaks by performing a machine learning classifier on the corresponding suspected peaks among the plurality of suspected peaks. as well as The device performs mass spectrometry analysis on the first portion of the mass spectrum corresponding to the set of valid peaks, but does not perform mass spectrometry analysis on the second portion of the mass spectrum corresponding to the set of invalid peaks.
10. The computer-implemented method according to claim 9, wherein, For the first suspected peak among the plurality of suspected peaks, the device feeds the first suspected peak or one or more properties of the first suspected peak as input to the machine learning classifier, wherein the machine learning classifier generates a classification label as output, the classification label indicating whether the first suspected peak is a valid peak or an invalid peak.
11. The computer-implemented method of claim 10, wherein the one or more properties of the first suspected peak include a first time-intensity tuple representing the vertex of the first suspected peak.
12. The computer-implemented method of claim 11, wherein the one or more properties of the first suspected peak further include: The second time-intensity tuple represents the starting point of the first suspected peak; or The third time-intensity tuple represents the endpoint of the first suspected peak.
13. The computer-implemented method of claim 11, wherein the one or more properties of the first suspected peak further include the width of the first suspected peak.
14. The computer-implemented method of claim 11, wherein the one or more properties of the first suspected peak further include the height of the first suspected peak.
15. The computer-implemented method of claim 11, wherein the one or more properties of the first suspected peak further include a hardware identifier associated with the chromatographic apparatus.
16. The computer-implemented method according to claim 9, further comprising: The machine learning classifier is trained using the device based on a training dataset, wherein the training dataset includes: Multiple training peaks; and Each of the multiple training peaks has a corresponding real classification label, and each real classification label indicates whether the corresponding training peak is valid or invalid in a binary manner.
17. A computer program product for facilitating post-detection chromatographic peak validation, the computer program product comprising a non-transitory computer-readable storage device having program instructions embodied thereon, the program instructions being executable by a processor to cause the processor to: The chromatographic equipment coupled to the mass spectrometer scans the sample, thereby generating chromatograms and mass spectra; Multiple suspected peaks in the chromatogram were identified using a peak detection algorithm; By performing a machine learning classifier on the corresponding suspected peaks among the plurality of suspected peaks, the plurality of suspected peaks are separated into a set of valid peaks and a set of invalid peaks; and Mass spectrometry analysis is performed on the first portion of the mass spectrum corresponding to the set of valid peaks, but not on the second portion of the mass spectrum corresponding to the set of invalid peaks.
18. The computer program product according to claim 17, wherein, For the first suspected peak among the plurality of suspected peaks, the processor feeds the first suspected peak or one or more properties of the first suspected peak as input to the machine learning classifier, wherein the machine learning classifier generates a classification label as output, the classification label indicating whether the first suspected peak is a valid peak or an invalid peak.
19. The computer program product of claim 18, wherein the one or more properties of the first suspected peak include a first time-intensity tuple representing the vertex of the first suspected peak.
20. The computer program product of claim 17, wherein the program instructions are further executable to cause the processor to: The machine learning classifier is trained based on a training dataset, wherein the training dataset includes: Multiple training peaks; as well as Each of the multiple training peaks has a corresponding real classification label, and each real classification label indicates whether the corresponding training peak is valid or invalid in a binary manner.