A computer-assisted method for computing confidence scores for metabolite identifications in metabolomics

The DecolD2 system addresses the challenge of metabolite identification in metabolomics by computing posterior probabilities using Bayesian modeling, enhancing metabolite detection accuracy and reducing noise, thereby expanding metabolome coverage and biological insights.

WO2026090249A1PCT designated stage Publication Date: 2026-04-30WASHINGTON UNIV IN SAINT LOUIS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
WASHINGTON UNIV IN SAINT LOUIS
Filing Date
2025-10-22
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Current metabolomics software lacks the ability to accurately identify metabolites with high confidence, leading to a significant bottleneck in data processing and limited metabolome coverage, with 90% of detected signals being false positives or noise and over 80% of metabolites remaining unidentified.

Method used

The DecolD2 system uses Bayesian modeling to compute posterior probabilities for metabolite identifications based on MS/MS match, retention time, and natural abundance isotope patterns, integrating a comprehensive metabolite library to provide quantitative confidence scores and automate filtering.

Benefits of technology

DecolD2 significantly increases the number of identified metabolites, reduces false discovery rates, and enables more robust biological interpretation by providing fine-grained confidence assessments, facilitating high-throughput metabolite analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052005_30042026_PF_FP_ABST
    Figure US2025052005_30042026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for computing confidence scores for metabolite identifications in metabolomics is provided. The method includes a) receiving experimental data for a subject to be identified; b) calculating similarity scores based on the experimental data; c) determining at least one probability of achieving the similarity scores based upon a plurality of historical identifications; d) comparing the at least one probability of achieving the similarity scores to the plurality of historical identifications; and e) determining at least one metabolite identification for the subject based upon the comparison.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A COMPUTER-ASSISTED METHOD FOR

[0002] COMPUTING CONFIDENCE SCORES FOR METABOLITE IDENTIFICATIONS IN METABOLOMICS CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to U.S. Provisional Application No.

[0004] 63 / 710,451, filed October 22, 2024, which is hereby incorporated by reference in its entirety.

[0005] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0006] Not applicable.

[0007] MATERIAL INCORPORATED-BY-REFERENCE

[0008] Not applicable.

[0009] FIELD OF THE INVENTION

[0010] The present disclosure generally relates to a method for computing confidence scores for metabolite identifications in metabolomics.

[0011] BACKGROUND OF THE INVENTION

[0012] Reprogramming of cellular metabolism is a hallmark of cancer. The majority of tumor types display high glucose uptake in vivo, as measured by to occur, as up to 80% of advanced cancer patients develop cachexia and 30% of cancer deaths are due to cachexia syndrome. Despite the ubiquity and importance of the tumorigenic metabolic program, it has been challenging to fully investigate due to limitations in metabolomics technologies.

[0013] Metabolomics is the ‘omics-level assessment of small-molecules occurring in organisms and their cells, biofluids, or tissues. The collection of the small molecules present within a biological system is referred to as its metabolome. These small molecules serve as substrates and products of metabolic reactions and can be produced endogenously or come from exogenous sources (food, drugs, environment). Metabolomics can be used to assess tumor heterogeneity, identify novel disease markers, and develop individualized treatment plans. Human metabolism is extremely complex, however, with tens-of-thousands of known endogenous metabolites. When considering small molecules derived from the environment, the chemical space that metabolomics aims to profile is nearly infinite.

[0014] Despite the immense potential of metabolomics, its true impact for human health has yet to be realized. To date, the vast majority of commercially available metabolomics platforms rely on targeted, panel-based approaches of select chemical classes and pathways. Even the broadest workflows only aim to profile a few thousand metabolites, which is far fewer than the number expected to occur in human biospecimens. As systemic metabolism is broadly dysregulated in cancer, metabolomics is well positioned for early disease detection, patient stratification, and novel drug target identification, but to fulfill this potential it requires truly unbiased approaches that provide accurate data on all metabolites present in a biological sample. The term “untargeted metabolomics” is exclusively reserved for this unbiased approach.

[0015] Given its broad coverage and high sensitivity, liquid chromatography / mass spectrometry (LC / MS) is the analytical instrument most often used in metabolomics. For a typical sample (e.g., plasma), LC / MS experiments produce tens of thousands of signals. Untargeted metabolomics is the analysis of all of these signals, which are also commonly referred to as “features.” However, such analysis is challenging due to the size of data generated and the large chemical diversity of metabolites. To reduce the data burden and complexity, metabolite signals are often searched for in a targeted fashion (e.g., targeted metabolomics), where a metabolite library is constructed by running authentic metabolite standards individually and then comparing detected peaks in a biological sample to those in the metabolite library. Targeted methods can be useful for some applications, such as testing specific hypotheses or when absolute quantification is required. However, the analytical potential of these targeted assays is limited to the <10% of human metabolites where authentic standards are commercially available.

[0016] In untargeted metabolomics, metabolite signals are detected in an unbiased fashion and compared to reference databases to make putative identifications. Downstream, compounds found to be of biological significance are confirmed using a standard during validation work ahead of clinical development (which is also required in targeted approaches). While LC / MS methods can always be improved to increase compound coverage, wide-spread adoption of untargeted metabolomics has been limited by insufficient analysis software used to detect, filter, and identify the tens-of-thousands of signals in LC / MS data. Existing software options do not yet make enough accurate metabolite identifications, leading to missed hits of potential biological significance and limited metabolome coverage. Metabolism is highly non-linear and observation of multiple species in a pathway is critical for robust data interpretation and enhances patient stratification and other precision medicine tasks. In addition, there are a multitude of examples of untargeted metabolomics uncovering potential drug targets that are novel compounds that would be missed with a targeted approach. These include, among others, N,N-dimethylsphingosine in chronic pain, 2HG in glioma, and Lac-Phe in obesity. These examples provide strong evidence for the value of unbiased metabolome assessment within metabolomics.

[0017] Untargeted metabolomics offers global profiling of the small molecule abundances within biological systems. For its sensitivity and coverage, liquid chromatography / mass spectrometry (LC / MS) is most often employed for this task. Computational processing of LC / MS data begins with peak detection where peaks in chromatographic retention time (RT) and mass to charge ratio (m / z) are detected in the LC / MS data and represent analytes measured in the experiment. The detected peaks are frequently referred to as “features” in untargeted metabolomics and defined by a unique pair of RT and m / z values. After peak detection and subsequent filtering, it is common to have >1,000 unique biological features detected in an experiment. The levels of these features can be assessed in each biological sample to determine analytes of interest. However, the challenge remains to structurally identify what metabolites these LC / MS features correspond to. The task of metabolite identification is considered the largest bottleneck in metabolomic data processing. Even when using state of the art instrumentation and computational approaches, the number of identified metabolites is generally less than 20% of the total number of presumed unique metabolites detected. The most common method of metabolite identification is based on tandem mass spectrometry (MS / MS) fragmentation data. The workflow for metabolite identification starts with querying the accurate mass and MS / MS spectra for a feature of interest against large metabolomics databases. Afterwards, typically several candidate compounds are returned and scored based on mass error and MS / MS similarity. Next, the compound with the highest MS / MS similarity is taken as the putative identification, and if available, the reference standard for the compound is purchased to confirm the identification with both MS / MS and retention-time data. While effective, this process can be very slow, expensive, and can take multiple iterations to correctly identify the compound.

[0018] Current software is limited by two major factors. Firstly, approximately 90% of the detected signals in an untargeted analysis do not represent unique metabolites but are instead false positives or redundant signals derived from chemical or electronic noise. This figure has been corroborated by multiple laboratories from across the world and was a primary point of agreement in a recent metabolomics workshop. Unfortunately, current software does not reliably remove these noise signals before performing downstream analyses, inflating the data burden, and limiting untargeted studies to small sample sizes. Secondly, more than 80% of detected metabolites cannot be identified with high confidence and are therefore removed from downstream analysis. This is in large part due to current software not providing statistical scores for metabolite identification certainty. These limitations lead to the misidentification of metabolites, creating a lack of trust in untargeted metabolomics data. They also make it difficult to prioritize leads worthy of downstream work, as there are often hundreds of statistically significant metabolites detected that are not identified or annotated as noise. There is a major need for innovative approaches to untargeted analysis software that accurately identify a broad range of metabolites, instill confidence in results, and can be scaled to the sample sizes required for true biomarker discovery. These needs must be addressed before metabolomics can reach its full impact in cancer research.

[0019] This Background section is intended to introduce the reader to various aspects of art that may be related to various aspects of the present disclosure, which are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present disclosure. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art. BRIEF SUMMARY

[0020] Among the various aspects of the present disclosure is the provision of a computer-assisted method for computing confidence scores for metabolite identifications in metabolomics.

[0021] In one aspect, a computer device for computing confidence scores for metabolite identifications in metabolomics is provided. The computer device includes at least one processor in communication with at least one memory device. The at least one processor may be configured to: a) receive experimental data for a subject to be identified; b) calculate similarity scores based on the experimental data; c) determine at least one probability of achieving the similarity scores based upon a plurality of historical identifications; d) compare the at least one probability of achieving the similarity scores to the plurality of historical identifications; and e) determine at least one metabolite identification for the subject based upon the comparison. The computer device may have additional, less, or alternate functionalities, including those discussed elsewhere herein.

[0022] In another aspect, a computer-implemented method for computing confidence scores for metabolite identifications in metabolomics is provided. The method includes a) receiving experimental data for a subject to be identified; b) calculating similarity scores based on the experimental data; c) determining at least one probability of achieving the similarity scores based upon a plurality of historical identifications; d) comparing the at least one probability of achieving the similarity scores to the plurality of historical identifications; and e) determining at least one metabolite identification for the subject based upon the comparison. The method may have additional, less, or alternate functionalities, including those discussed elsewhere herein.

[0023] In a further aspect, a system for computing confidence scores for metabolite identifications in metabolomics is provided. The system and includes a computer device that includes at least one processor in communication with at least one memory device. The at least one processor may be configured to: a) receive experimental data for a subject to be identified; b) calculate similarity scores based on the experimental data; c) determine at least one probability of achieving the similarity scores based upon a plurality of historical identifications; d) compare the at least one probability of achieving the similarity scores to the plurality of historical identifications; and e) determine at least one metabolite identification for the subject based upon the comparison. The system may have additional, less, or alternate functionalities, including those discussed elsewhere herein.

[0024] Advantages will become more apparent to those skilled in the art from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.

[0025] DESCRIPTION OF THE DRAWINGS

[0026] The Figures described below depict various aspects of the systems and methods disclosed. Each Figure depicts an embodiment of a particular aspect of the disclosed systems and methods, and that each of the Figures is intended to accord with a possible embodiment. Further, wherever possible, the following description refers to the reference numerals included in the following Figures, in which features depicted in multiple Figures are designated with consistent reference numerals.

[0027] Figure 1 illustrates an example process 100 for computing confidence scores for metabolite identifications in metabolomics, in accordance with at least one embodiment.

[0028] Figure 2 illustrates an exemplary computer system that may be used with the process shown in Figure 1.

[0029] Figure 3 illustrates a component configuration of the DecolD2 computing device shown in Figure 2.

[0030] Figure 4 illustrates an example configuration of a client system, in accordance with one embodiment of the present disclosure.

[0031] Figure 5 illustrates an example configuration of a server system, in accordance with one embodiment of the present disclosure.

[0032] FIG. 6A illustrates an example graph of the application of DecolD2 to colon cancer samples.

[0033] FIG. 6B illustrates an example graph if ROC curves comparing Pr- based metabolite ID with the MS / MS metric (MS2 sim.) used by DecolD.

[0034] FIG. 60 illustrates an example graph of a comparison of empirical FDR vs expected FDRs that shows that returned ID probabilities are accurate.

[0035] Like reference symbols in the various drawings indicate like elements.

[0036] DETAILED DESCRIPTION

[0037] The present disclosure generally relates to a method for computing confidence scores for metabolite identifications in metabolomics. More specifically, the present disclosure relates to systems and methods for computing confidence scores for metabolite identifications in metabolomics by using a statistical method for integrating liquid chromatography / mass spectrometry (LC / MS) data and computing identification probabilities. More specifically, a DecolD2 system enables automated filtering of metabolite hits with defined confidence. Furthermore, the DecolD2 system stores previously identified metabolites and uses them to preferentially weight metabolite hits in subsequent experiments, serving as a “search history” to accelerate and improve future metabolite identifications.

[0038] The DecolD2 system fundamentally enables a fine-grained assessment of metabolite confidence. This is significantly different from the qualitative scoring that was traditionally used where metabolites were binned into categories in a somewhat arbitrary manner. These confidence values also enable the false discovery rate within a metabolomics dataset to be calculated. With the fined grained confidence values, several new uses of the output become possible. In particular, biomarker selection for downstream validation and clinical development can be done based on the candidate biomarker metabolites that have the highest DecolD2 confidence score. Additionally, benchmarking of different platforms / workflow for performing metabolomics can be achieved by counting the number of identified metabolites at a defined false discovery rate.

[0039] In the example embodiment, the DecolD2 system is powered by a comprehensive metabolite library of >285,000 compounds sourced from public databases with fragmentation data and retention times for >99% of all database compounds. The fragmentation data can be sourced from sources such as, but not limited to, the Human Metabolome Database (HMDB), Lipid Maps, RefMet, the Global Natural Products Social Molecular Network (GNPS) database, and Competitive Fragmentation Modeling for Metabolite Identification (CFM-ID). In some embodiments, the fragmentation data may contain more than 1.5 million tandem mass spectrometry (MS / MS) spectra. Retention-time data is from analysis of >1,000 metabolite standards with the LC / MS methodology, which the DecolD2 system then utilizes to predict retention times for all database compounds.

[0040] The probability metric that underpins the DecolD2 system computes the posterior probability that a metabolite identification, m, is correct given the degree of match to the available experimental data (9) relative to all other potential metabolites in database (M) as well as prior metabolite identification probabilities for all compounds (y). This may be accomplished through Bayesian modeling of metabolite identification probabilities (Eq. 1).

[0041] Pr(m|0,y) — Pr(0|m, y) Pr (m|y) / (2(m' G M)[Pr(0|m', y) Pr(m'|y)] EQ. 1

[0042] Here, 6 corresponds to the degree of MS / MS match (a), degree of retention time match ( ), and degree of natural abundance isotope pattern match ( . As these metrics come from independent sources and are independent of prior identifications, Pr(61 m,y) is decomposed as Pr (a| m) Pr(|31 m) Pr( | m). Pr(a| y), Pr([31 y), and Pr( | y) can be computed by first computing standard similarity metrics for each piece of experimental data (entropy similarity, retention-time deviation, and reverse dot-product). Intuitively, Pr(a| y), Pr([3| y), and Pr(^| y) represent the probabilities of achieving these similarity scores given the metabolite identification is correct. Probability density functions (PDFs) for Pr(a | y), and Pr( | y) are learned on-the-fly through an iterative cross validation procedure to ensure a well suited prior. These PDFs enable conversion of similarity scores into Pr(a | y), Pr(|31 y), and Pr(^| y) values. To infer prior metabolite identification probabilities, Pr(m | y), an iterative expectation-maximization (EM) inference routine is applied to learn the y that maximizes the likelihood of the observed matches to the experimental data. It has been validated that the DecolD2 system probabilities are accurate through the analysis of chemical standards and that the DecolD2 system expands the number of identified metabolites. As shown herein, development of the DecolD2 system is described. The disclosed methods enable the DecolD2 system to perform automated metabolite identifications with quantitative confidence assessments. Accordingly, the DecolD2 system provides for more high throughput metabolite analysis and expanded biological insights from such data, a functionality needed for metabolomics service providers.

[0043] Until now, metabolites were either fully accepted or rejected from downstream statistical analyses based on qualitative scoring of metabolite identification confidence and arbitrary cutoffs. This reduces the coverage of many metabolic pathways. The DecolD2 system not only expands the number of identified metabolites but also enables the use of partial compound identifications in pathway analysis, enabling greater biological interpretation of results.

[0044] In various aspects, the disclosed metabolite confidence score methods may be implemented using a computing system or computing device.

[0045] Figure 1 illustrates an example process 100 for computing confidence scores for metabolite identifications in metabolomics, in accordance with at least one embodiment. More specifically, process 100 describes computing confidence scores for metabolite identifications in metabolomics by using a statistical method for integrating liquid chromatography / mass spectrometry (LC / MS) data and computing identification probabilities. More specifically, the DecolD2 system 200 (shown in Figure 2) enables automated filtering of metabolite hits with defined confidence. Furthermore, the DecolD2 system 200 stores previously identified metabolites and uses them to preferentially weight metabolite hits in subsequent experiments, serving as a “search history” to accelerate and improve future metabolite identifications.

[0046] In the example embodiment, the DecolD2 computer device 210 (shown in Figure 2) of the DecolD2 system 200 performs the steps of process 100.

[0047] In the example embodiment, the DecolD2 computer device 210 is in communication with a database 220 (shown in Figure 2) that includes a plurality of previously successfully identified metabolites. In these embodiments, the DecolD2 computer device 210 is also in communication with a database 220 that includes a comprehensive metabolite library of compounds. In some embodiments, the metabolite library has over 285,000 compounds. In some further embodiments, the metabolite library includes fragmentation data and retention times for the compounds in the database 220. In some embodiments, the library of compounds is source from sources such as, but not limited to, the Human Metabolome Database (HMDB), Lipid Maps, RefMet, the Global Natural Products Social Molecular Network (GNPS) database, and Competitive Fragmentation Modeling for Metabolite Identification (CFM-ID). In some embodiments, the fragmentation data may contain more than 1.5 million MS / MS spectra records. Retention-time data is obtained from analysis of >1,000 metabolite standards with the LC / MS methodology. The DecolD2 system 200 utilizes the analysis of the metabolite standards with the LC / MS methodology to predict retention times for all database compounds.

[0048] The probability metric that underpins the DecolD2 system 200 computes the posterior probability that a metabolite identification, m, is correct given the degree of match to the available experimental data (9) relative to all other potential metabolites in database (M) as well as prior metabolite identification probabilities for all compounds (y). This may be accomplished through Bayesian modeling of metabolite identification probabilities (Eq. 1).

[0049] Pr(m|0,y) — Pr(0|m, y) Pr (m|y) / (2(m' G M)[Pr(0|m', y) Pr(m'|y)] EQ. 1

[0050] Here, 6 corresponds to the degree of MS / MS match (a), degree of retention time match ( ), and degree of natural abundance isotope pattern match ( . As these metrics come from independent sources and are independent of prior identifications, Pr(6| m,Y) is decomposed as Pr (a| m) Pr(p| m) Pr(^| m). The DecolD2 system 200 computes Pr(a| y), Pr((3| y), and Pr(^| y) by first computing standard similarity metrics for each piece of experimental data (entropy similarity, retention-time deviation, and reverse dot-product). Intuitively, Pr(a| y), Pr(|31 y), and Pr(^| y) represent the probabilities of achieving these similarity scores given the metabolite identification is correct. The DecolD2 system 200 learns probability density functions (PDFs) for Pr(a | y), and Pr(^| y) through an iterative cross validation procedure to ensure a well suited prior. These PDFs enable the DecolD2 system to convert similarity scores into Pr(a| y), Pr([3| y), and Pr(^| y) values. To infer prior metabolite identification probabilities, the DecolD2 system 200 applies Pr(m | Y), an iterative expectation-maximization (EM) inference routine, to learn the Y that maximizes the likelihood of the observed matches to the experimental data. In the example embodiment, the DecolD2 computer device 210 receives 105 experimental data for a subject to be identified. In some embodiments, the experimental data includes mass spectrometry data from a sample of the subject to be identified.

[0051] In the example embodiment, the DecolD2 computer device 210 calculates 110 similarity scores based on the experimental data. In some embodiments, the similarity scores include a degree of tandem mass spectrometry match, a degree of retention time match, and a degree of natural abundance isotope pattern match. In other embodiments, the similarity scores include at least one of entropy similarity, retention time deviation, and reverse dot product.

[0052] In the example embodiment, the DecolD2 computer device 210 determines 115 at least one probability of achieving the similarity scores based upon a plurality of historical identifications.

[0053] In the example embodiment, the DecolD2 computer device 210 compares 120 the at least one probability of achieving the similarity scores to the plurality of historical identifications.

[0054] In the example embodiment, the DecolD2 computer device 210 determines at least one metabolite identification for the subject based upon the comparison. In some further embodiments, the DecolD2 computer device 210 calculates a degree of match to the experimental data. The DecolD2 computer device 210 calculates a posterior probability that a metabolite identification is correct based upon the degree of match to the experimental data relative to a plurality of other potential metabolites. The DecolD2 computer device 210 calculates a posterior probability that a metabolite identification is correct based upon the degree of match to the experimental data relative to a plurality of historical metabolite identification probabilities.

[0055] In some further embodiments, the DecolD2 computer device 210 performs Bayesian modeling of metabolite identification probabilities. The DecolD2 computer device 210 calculates a posterior probability that a metabolite identification is correct, wherein the posterior probability is quantified with Bayesian modeling.

[0056] In some further embodiments, the DecolD2 computer device 210 utilizes probability density functions to convert similarity scores into probability values

[0057] Figure 2 illustrates an exemplary computer system 200 that may be used with the process 100 (shown in Figure 1). In the exemplary embodiment, the system 200 for implementing computing confidence scores for metabolite identifications in metabolomics. In some embodiments, the system 200 analyzes the sensor information such as, but not limited to, mass spectrometry data to compute posterior probabilities (confidence scores) that metabolite identifications are correct based on matching to all other potential metabolics in a database.

[0058] As described herein in more detail, a DecolD2 computer device 210 (also known as a DecolD2 server 210) is programmed to analyze sensor information, such as mass spectrometry data to identify metabolites. In addition, the DecolD2 server 210 is programmed to continuously update and “train” one or more models used to computer posterior probabilities. The DecolD2 server 210 is programmed to a) receive 105 experimental data for a subject to be identified; b) calculate 110 similarity scores based on the experimental data; c) determine 115 at least one probability of achieving the similarity scores based upon a plurality of historical identifications; d) compare 120 the at least one probability of achieving the similarity scores to the plurality of historical identifications; and e) determine 125 at least one metabolite identification for the subject based upon the comparison.

[0059] In the exemplary embodiment, user computing devices 205 are computers that include a web browser or a software application, which enables user computer devices 205 to communicate with DecolD2 server 210 and / or mass spectrometry system 230 using the Internet, a local area network (LAN), or a wide area network (WAN). In some embodiments, the user computing devices 205 are communicatively coupled to one or more networks 225, such as the Internet, through many interfaces including, but not limited to, at least one of a network, such as the Internet, a LAN, a WAN, or an integrated services digital network (ISDN), a dial-up- connection, a digital subscriber line (DSL), a cellular phone connection, a satellite connection, and a cable modem. User computing devices 205 can be any device capable of accessing a network, such as the Internet, including, but not limited to, a desktop computer, a laptop computer, a personal digital assistant (PDA), a cellular phone, a smartphone, a tablet, a phablet, wearable electronics, smart watch, virtual headsets or glasses (e.g., AR (augmented reality), VR (virtual reality), or XR (extended reality) headsets or glasses), chat bots, voice bots, ChatGPT bots or ChatGPT-based bots, or other web-based connectable equipment or mobile devices.

[0060] In the exemplary embodiment, DecolD2 computer device 210 (also known as DecolD2 server 210) is a computer that includes a web browser or a software application, which enables DecolD2 server 210 to communicate with user computing devices 205 and mass spectrometry systems 230 using the Internet, a local area network (LAN), and / or a wide area network (WAN). In some embodiments, the DecolD2 server 210 is communicatively coupled to one or more networks 225, such as the Internet, through many interfaces including, but not limited to, at least one of a network, such as the Internet, a LAN, a WAN, or an integrated services digital network (ISDN), a dial-up-connection, a digital subscriber line (DSL), a cellular phone connection, a satellite connection, and a cable modem. DecolD2 server 210 can be any device capable of accessing a network, such as the Internet, including, but not limited to, a desktop computer, a laptop computer, a personal digital assistant (PDA), a cellular phone, a smartphone, a tablet, a phablet, wearable electronics, smart watch, virtual headsets or glasses (e.g., AR (augmented reality), VR (virtual reality), orXR (extended reality) headsets or glasses), chat bots, voice bots, ChatGPT bots or ChatGPT-based bots, or other web-based connectable equipment or mobile devices.

[0061] Mass spectrometry system 230 may be any system that provides mass spectrometry data include the device that captures that data or a device that stores the data after capture. The mass spectrometry system 230 is in communication with the DecolD2 server 210. In the example embodiment, mass spectrometry systems 230 are computers that include a web browser or a software application, which enables mass spectrometry systems 230 to communicate with user computing devices 205 and DecolD2 server 210 using the Internet, a local area network (LAN), or a wide area network (WAN). In some embodiments, the mass spectrometry systems 230 are communicatively coupled to one or more networks 225, such as the Internet, through many interfaces including, but not limited to, at least one of a network, such as the Internet, a LAN, a WAN, or an integrated services digital network (ISDN), a dial-up-connection, a digital subscriber line (DSL), a cellular phone connection, a satellite connection, and a cable modem. Mass spectrometry systems 230 can be any device capable of accessing a network, such as the Internet, including, but not limited to, a desktop computer, a laptop computer, a personal digital assistant (PDA), a cellular phone, a smartphone, a tablet, a phablet, wearable electronics, smart watch, virtual headsets or glasses (e.g., AR (augmented reality), VR (virtual reality), orXR (extended reality) headsets or glasses), chat bots, voice bots, ChatGPT bots or ChatGPT-based bots, or other web-based connectable equipment or mobile devices.

[0062] A database server 215 is communicatively coupled to a database 220 that stores data. In one embodiment, the database 220 is a database that includes prior metabolite identification probabilities. In some embodiments, the database 220 is stored remotely from the DecolD2 server 210. In some embodiments, the database 220 is decentralized. In the example embodiment, a person can access the database 220 via the user computing devices 205 by logging onto DecolD2 server 210.

[0063] Figure 3 illustrates a component configuration 300 of the DecolD2 computing device 210 (shown in Figure 2). The component configuration 300 includes database 220 along with other related computing components. A user 305 may access components of DecolD2 computing device 210.

[0064] In one aspect, database 220 includes LC / MS data 310 and confidence score data 320. LC / MS data 310 may include data used to operate an LC / MS system using the acquisition methods as disclosed herein. Non-limiting examples of LC / MS data 310 include various measurements of LC / MS signals and any parameters used to control the operation of an LC / MS device for the metabolite confidence score methods as disclosed herein. Confidence score data 320 may include any parameters defining equations or other algorithms used to implement the metabolite identification probability methods as disclosed herein. DecolD2 computing device 210 also includes a number of components that perform specific tasks. In the exemplary aspect, DecolD2 computing device 210 includes a data storage device 330, an LC / MS acquisition component 340, an identification probability component 350, and a communication component 360. The LC / MS acquisition component 340 is configured to implement LC / MS methods as described herein. The identification probability component 350 is configured to implement the metabolite confidence score methods as disclosed herein. The data storage device 330 is configured to store data received or generated by DecolD2 computing device 210, such as any of the data stored in database 220 or any outputs of processes implemented by any component of DecolD2 computing device 210.

[0065] The communication component 360 is configured to enable communications between DecolD2 computing device 210 and other devices (e.g., user computing device 205 shown in Figure 2) over a network, such as a network 225 (shown in Figure 2), or a plurality of network connections using predefined network protocols such as TCP / IP (Transmission Control Protocol / lnternet Protocol).

[0066] Figure 4 illustrates an example configuration 400 of a user computer device 402, in accordance with one embodiment of the present disclosure. User computer device 402 is operated by a user 401. User computer device 402 may include, but is not limited to, user computing device 205, DecolD2 server 210, and mass spectrometry system 230 (all shown in Figure 2). User computer device 402 includes a processor 405 for executing instructions. In some embodiments, executable instructions are stored in a memory area 410. Processor 405 may include one or more processing units (e.g., in a multi-core configuration). Memory area 410 is any device allowing information such as executable instructions and / or transaction data to be stored and retrieved. Memory area 410 may include one or more computer-readable media.

[0067] User computer device 402 also includes at least one media output component 415 for presenting information to user 401. Media output component 415 is any component capable of conveying information to user 401. In some embodiments, media output component 415 includes an output adapter (not shown) such as a video adapter and / or an audio adapter. An output adapter is operatively coupled to processor 405 and operatively coupleable to an output device such as a display device (e.g., a cathode ray tube (CRT), liquid crystal display (LCD), light emitting diode (LED) display, or “electronic ink” display) or an audio output device (e.g., a speaker or headphones). In some embodiments, media output component 415 is configured to present a graphical user interface (e.g., a web browser and / or a client application) to user 401. A graphical user interface may include, for example, an interface for viewing compounds. In some embodiments, user computer device 402 includes an input device 420 for receiving input from user 401. User 401 may use input device 420 to, without limitation, select mass spectrometry data to analyze. Input device 420 may include, for example, a keyboard, a pointing device, a mouse, a stylus, a touch sensitive panel (e.g., a touch pad or a touch screen), a gyroscope, an accelerometer, a position detector, a biometric input device, and / or an audio input device. A single component such as a touch screen may function as both an output device of media output component 415 and input device 420.

[0068] User computer device 402 may also include a communication interface 425, communicatively coupled to a remote device such as mass spectrometry system 230 (shown in Figure 2). Communication interface 425 may include, for example, a wired or wireless network adapter and / or a wireless data transceiver for use with a mobile telecommunications network.

[0069] Stored in memory area 410 are, for example, computer-readable instructions for providing a user interface to user 401 via media output component 415 and, optionally, receiving and processing input from input device 420. A user interface may include, among other possibilities, a web browser and / or a client application. Web browsers enable users, such as user 401, to display and interact with media and other information typically embedded on a web page or a website from DecolD2 server 210. A client application allows user 401 to interact with, for example, DecolD2 server 210. For example, instructions may be stored by a cloud service, and the output of the execution of the instructions sent to the media output component 415.

[0070] Processor 405 executes computer-executable instructions for implementing aspects of the disclosure. In some embodiments, the processor 405 is transformed into a special purpose microprocessor by executing computerexecutable instructions or by otherwise being programmed.

[0071] Figure 5 illustrates an example configuration 500 of a server system 502, in accordance with one embodiment of the present disclosure. Server computer device 502 may include, but is not limited to, DecolD2 server 210, database server 315, and mass spectrometry system 230 (all shown in Figure 2). Server computer device 502 also includes a processor 505 for executing instructions. Instructions may be stored in a memory area 510. Processor 505 may include one or more processing units (e.g., in a multi-core configuration).

[0072] Processor 505 is operatively coupled to a communication interface 515 such that server computer device 502 is capable of communicating with a remote device such as another server computer device 502, another DecolD2 server 210, mass spectrometry system 230, or user computing device 205 (shown in Figure 2). For example, communication interface 515 may receive requests from user computing device 205 via the Internet, as illustrated in Figure 2.

[0073] Processor 505 may also be operatively coupled to a storage device 525. Storage device 525 is any computer-operated hardware suitable for storing and / or retrieving data, such as, but not limited to, data associated with database 220 (shown in Figure 2). In some embodiments, storage device 525 is integrated in server computer device 501. For example, server computer device 501 may include one or more hard disk drives as storage device 525. In other embodiments, storage device 525 is external to server computer device 502 and may be accessed by a plurality of server computer devices 502. For example, storage device 525 may include a storage area network (SAN), a network attached storage (NAS) system, and / or multiple storage units such as hard disks and / or solid state disks in a redundant array of inexpensive disks (RAID) configuration.

[0074] In some embodiments, processor 505 is operatively coupled to storage device 534 via a storage interface 520. Storage interface 520 is any component capable of providing processor 505 with access to storage device 525. Storage interface 520 may include, for example, an Advanced Technology Attachment (ATA) adapter, a Serial ATA (SATA) adapter, a Small Computer System Interface (SCSI) adapter, a RAID controller, a SAN adapter, a network adapter, and / or any component providing processor 505 with access to storage device 534.

[0075] Processor 505 executes computer-executable instructions for implementing aspects of the disclosure. In some embodiments, the processor 505 is transformed into a special purpose microprocessor by executing computerexecutable instructions or by otherwise being programmed. For example, the processor 505 is programmed with instructions such as illustrated in Figure 1.

[0076] SCREENING

[0077] The subject methods find use in the screening of a variety of different candidate molecules (e.g., potentially therapeutic candidate molecules). Candidate substances for screening according to the methods described herein include, but are not limited to, fractions of tissues or cells, nucleic acids, polypeptides, siRNAs, antisense molecules, aptamers, ribozymes, triple helix compounds, antibodies, and small (e.g. , less than about 2000 mw, or less than about 1000 mw, or less than about 800 mw) organic molecules or inorganic molecules including but not limited to salts or metals.

[0078] Candidate molecules encompass numerous chemical classes, for example, organic molecules, such as small organic compounds having a molecular weight of more than 50 and less than about 2,500 Daltons. Candidate molecules can comprise functional groups necessary for structural interaction with proteins, particularly hydrogen bonding, and typically include at least an amine, carbonyl, hydroxyl, or carboxyl group, and usually at least two of the functional chemical groups. The candidate molecules can comprise cyclical carbon or heterocyclic structures and / or aromatic or polyaromatic structures substituted with one or more of the above functional groups.

[0079] A candidate molecule can be a compound in a library database of compounds. One of skill in the art will be generally familiar with, for example, numerous databases for commercially available compounds for screening (see e.g., ZINC database, LICSF, with 2.7 million compounds over 12 distinct subsets of molecules; Irwin and Shoichet (2005) J Chem Inf Model 45, 177-182). One of skill in the art will also be familiar with a variety of search engines to identify commercial sources or desirable compounds and classes of compounds for further testing (see e.g., ZINC database; eMolecules.com; and electronic libraries of commercial compounds provided by vendors, for example: ChemBridge, Princeton BioMolecular, Ambinter SARL, Enamine, ASDI, Life Chemicals etc.).

[0080] Candidate molecules for screening according to the methods described herein include both lead-like compounds and drug-like compounds. A lead-like compound is generally understood to have a relatively smaller scaffold-like structure (e.g., molecular weight of about 150 to about 350 kD) with relatively fewer features (e.g., less than about 3 hydrogen donors and / or less than about 6 hydrogen acceptors; hydrophobicity character xlogP of about -2 to about 4) (see e.g., Angewante (1999) Chemie Int. ed. Engl. 24, 3943-3948). In contrast, a drug-like compound is generally understood to have a relatively larger scaffold (e.g., molecular weight of about 150 to about 500 kD) with relatively more numerous features (e.g., less than about 10 hydrogen acceptors and / or less than about 8 rotatable bonds; hydrophobicity character xlogP of less than about 5). Initial screening can be performed with lead-like compounds.

[0081] When designing a lead from spatial orientation data, it can be useful to understand that certain molecular structures are characterized as being “druglike.” Such characterization can be based on a set of empirically recognized qualities derived by comparing similarities across the breadth of known drugs within the pharmacopoeia. While it is not required for drugs to meet all, or even any, of these characterizations, it is far more likely for a drug candidate to meet with clinical successful if it is drug-like.

[0082] Several of these “drug-like” characteristics have been summarized into the four rules of Lipinski (generally known as the “rules of fives” because of the prevalence of the number 5 among them). While these rules generally relate to oral absorption and are used to predict bioavailability of compound during lead optimization, they can serve as effective guidelines for constructing a lead molecule during rational drug design efforts such as may be accomplished by using the methods of the present disclosure.

[0083] The four “rules of five” state that a candidate drug-like compound should have at least three of the following characteristics: (i) a weight less than 500 Daltons; (ii) a log of P less than 5; (iii) no more than 5 hydrogen bond donors (expressed as the sum of OH and NH groups); and (iv) no more than 10 hydrogen bond acceptors (the sum of N and O atoms). Also, drug-like molecules typically have a span (breadth) of between about 8A to about 15A.

[0084] A control sample or a reference sample as described herein can be a sample from a healthy subject. A reference value can be used in place of a control or reference sample, which was previously obtained from a healthy subject or a group of healthy subjects. A control sample or a reference sample can also be a sample with a known amount of a detectable compound or a spiked sample.

[0085] Compositions and methods described herein utilizing molecular biology protocols can be according to a variety of standard techniques known to the art.

[0086] Examples

[0087] The following non-limiting examples are provided to further illustrate the present disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the examples that follow represent approaches the inventors have found function well in the practice of the present disclosure, and thus can be considered to constitute examples of modes for its practice. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments that are disclosed and still obtain a like or similar result without departing from the spirit and scope of the present disclosure.

[0088] Example 1 - DecolD2

[0089] Beyond removing noise, metabolite identification is widely regarded as the largest bottleneck in downstream processing of metabolomics data. To identify a metabolite, experimental data (such as fragmentation data, accurate masses, and retention time for detected metabolite signals) are used to query potential metabolite identifications from reference databases. MS / MS fragmentation data is often the most salient piece of experimental data to inform metabolite identification as it can be acquired rapidly and large reference libraries of MS / MS spectra for metabolites have been created. MS / MS data acquired on research samples can be compared to that of pure reference standards in metabolomics databases to inform metabolite identification. Previous software improved such MS / MS matching workflows by >30% without increasing false discovery through deconvolution of MS / MS data. However, an outstanding problem is the lack of a metabolite scoring scheme that quantitates the probability that a metabolite identification is correct. Previous methods make weak assumptions that limit the number of identified metabolites and waste resources on downstream work when misidentifications occur. Others have proposed multiple qualitative scoring schemes, such as the Metabolomics Standards Initiative (MSI) levels. MSI level 1 and 2 metabolites are highly confident structural identifications, MSI level 3 metabolites are only accurate to the molecular formula level, and MSI level 4 metabolites are unknowns. Only MSI level 1 and 2 metabolites are used to guide biological interpretation of metabolomics data, which excludes many metabolites with biological significance.

[0090] In this example, a metabolite scoring system with statistical significance is developed. The method computes posterior probabilities of metabolite identifications based on Bayesian integration of all available experimental information (fragmentation data, retention time, accurate mass, etc.) as well as leveraging metabolites identified in separate experiments of the same sample type to weight matches and form a “search history” of identified metabolites. Such an approach has previously been applied to construct confidence scores in other fields (e.g., transcription factor motif inference) but has never been applied in metabolomics. This scoring method is implemented by the DecolD2 system 200 described herein. The DecolD2 system 200 not only performs automated data deconvolution and database matching, but returns the identification probabilities for each signal. The DecolD2 system 200 can be powered by a comprehensive metabolite library of >285,000 compounds with fragmentation data and retention times for >99% of all database compounds. The fragmentation data is sourced from HMDB, Lipid Maps, RefMet, GNPS, and CFM-ID 4.0, and contains more than 1.5 million MS / MS spectra, making it the most comprehensive MS / MS database. Retention-time data is from analysis of >1 ,000 metabolite standards with our LC / MS methodology, which is then utilized to predict retention times for all database compounds with a method similar to that of Retip. When combined with the other modules of the existing analysis pipelines, the number of identified metabolites can be increased, positive hits can be reduced, and data interpretation can be improved through downstream incorporation of partial metabolite identifications.

[0091] When applying the probability scoring metric embedded into the DecolD2 system 200 to a colon cancer dataset, the identification confidence for all 1,509 metabolites was quantified, enabling the usage of all metabolites in downstream pathway enrichment analyses.

[0092] Figure 6A illustrates an example graph of improved identification counts that are seen across all score thresholds, including 343 identifications at an FDR < 5%. Thirty of these metabolites reached statistical significance but did not meet the standards required by MSI level 1 or 2, including dehydroascorbate and other compounds critical to observing vitamin C dysregulation in colon cancer. Figure 6B illustrates a graph of a scoring metric implemented within the DecolD2 system 200 (shown in Figure 2). The scoring metric was validated with authentic standards where improved identification performance was seen when compared to solely basing metabolite identifications on the quality of MS / MS matches. Figure 6C illustrates a graph of posterior probabilities returned by the DecolD2 system 200. The posterior probabilities reflect the actual identification likelihood, enabling false discovery rate (FDR)-controlled filtering of metabolite hits.

[0093] This example was designed to enable inclusion of all identified metabolites found during an initial assessment of colon cancer samples into downstream analyses, such as pathway analysis. Due to the lack of statistically derived quantitative confidence levels, stringent and somewhat arbitrary cutoffs restrict usage to only a small fraction of identified metabolites, removing informative but ambiguous identifications. The probability metric is at the core of the DecolD2 system 200. Further, validation in complex biological matrices (e.g., human plasma, tissue, etc.) through spike-in studies with authentic standards can be performed.

[0094] In some aspects, the present disclosure describes the integration of Bayesian scoring of metabolite matches to creating the DecolD2 system 200. The probability metric included in the DecolD2 system 200 computes the posterior probability that a metabolite identification, m, is correct given the degree of match to the available experimental data (0) relative to all other potential metabolites in the database (M) as well as prior metabolite identification probabilities for all compounds (Y). This will be accomplished through Bayesian modeling of metabolite identification probabilities (EQ. 1).

[0095] Pr(m|0,y) — Pr(0|m, y) Pr (m|y) / (2(m' G M)[Pr(0|m', y) Pr(m'|y)]

[0096] EQ. 1

[0097] Here, 0 corresponds to the degree of MS / MS match (a), degree of retention time match ( ), and degree of natural abundance isotope pattern match ( . As these metrics come from independent sources and are independent of prior identifications, Pr(0|m,Y) is decomposed as Pr(a|m) Pr(|3|m) Pr( |m). Pr(a|y), PT([3|Y), and Pr(^|y) can be computed by first computing standard similarity metrics for each piece of experimental data (entropy similarity, retention-time deviation, and reverse dot-product). Intuitively, Pr(a|y), Pr(P|y), and Pr(^|y) represent the probabilities of achieving these similarity scores given the metabolite identification is correct. Probability density functions (PDFs) for Pr(a|y), and Pr( | y) are learned on-the-fly through an iterative cross validation procedure to ensure a well suited prior. These PDFs enable conversion of similarity scores into Pr(a|v), PT(|3|Y), and Pr(^lv) values.

[0098] To infer prior metabolite identification probabilities, Pr(m|v), an iterative expectation-maximization (EM) inference routine can be applied to learn the Y that maximizes the likelihood of the observed matches to the experimental data. The DecolD2 system 200 can then enable LC / MS / MS data to be automatically processed into metabolite identifications with associated probabilities of the compound identity being correct.

[0099] All publications, patents, patent applications, and other references cited in this application are incorporated herein by reference in their entirety for all purposes to the same extent as if each individual publication, patent, patent application or other reference was specifically and individually indicated to be incorporated by reference in its entirety for all purposes. Citation of a reference herein shall not be construed as an admission that such is prior art to the present disclosure. ADDITIONAL CONSIDERATIONS

[0100] As will be appreciated based upon the foregoing specification, the above-described embodiments of the disclosure may be implemented using computer programming or engineering techniques including computer software, firmware, hardware or any combination or subset thereof. Any such resulting program, having computer-readable code means, may be embodied or provided within one or more computer-readable media, thereby making a computer program product, i.e. , an article of manufacture, according to the discussed embodiments of the disclosure. The computer-readable media may be, for example, but is not limited to, a fixed (hard) drive, diskette, optical disk, magnetic tape, semiconductor memory such as read-only memory (ROM), and / or any transmitting / receiving medium such as the Internet or other communication network or link. The article of manufacture containing the computer code may be made and / or used by executing the code directly from one medium, by copying the code from one medium to another medium, or by transmitting the code over a network.

[0101] These computer programs (also known as programs, software, software applications, “apps,” or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” “computer-readable medium” refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The “machine-readable medium” and “computer-readable medium,” however, do not include transitory signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0102] As used herein, the terms “processor” and “computer” and related terms, e.g., “processing device”, “computing device”, and “controller” are not limited to just those integrated circuits referred to in the art as a computer, but broadly refers to a microcontroller, a microcomputer, a programmable logic controller (PLC), a reduced instruction set circuit (RISC), an application specific integrated circuit (ASIC), logic circuits, and any other circuit or processor capable of executing the functions described herein. The above examples are example only and are thus not intended to limit in any way the definition and / or meaning of the term “processor.”

[0103] As used herein, the terms “software” and “firmware” are interchangeable, and include any computer program stored in memory for execution by a processor, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are example only, and are thus not limiting as to the types of memory usable for storage of a computer program.

[0104] As used herein, the term “database” can refer to either a body of data, a relational database management system (RDBMS), or to both. As used herein, a database can include any collection of data including hierarchical databases, relational databases, flatfile databases, object-relational databases, object-oriented databases, and any other structured collection of records or data that is stored in a computer system. The above examples are example only, and thus are not intended to limit in any way the definition and / or meaning of the term database. Examples of RDBMS’ include, but are not limited to including, Oracle® Database, MySQL, IBM® DB2, Microsoft® SQL Server, Sybase®, and PostgreSQL. However, any database can be used that enables the systems and methods described herein. (Oracle is a registered trademark of Oracle Corporation, Redwood Shores, California; IBM is a registered trademark of International Business Machines Corporation, Armonk, New York; Microsoft is a registered trademark of Microsoft Corporation, Redmond, Washington; and Sybase is a registered trademark of Sybase, Dublin, California.)

[0105] In another example, a computer program is provided, and the program is embodied on a computer-readable medium. In an example, the system is executed on a single computer system, without requiring a connection to a server computer. In a further example, the system is being run in a Windows® environment (Windows is a registered trademark of Microsoft Corporation, Redmond, Washington). In yet another example, the system is run on a mainframe environment and a UNIX® server environment (UNIX is a registered trademark of X / Open Company Limited located in Reading, Berkshire, United Kingdom). In a further example, the system is run on an iOS® environment (iOS is a registered trademark of Cisco Systems, Inc. located in San Jose, CA). In yet a further example, the system is run on a Mac OS® environment (Mac OS is a registered trademark of Apple Inc. located in Cupertino, CA). In still yet a further example, the system is run on Android® OS (Android is a registered trademark of Google, Inc. of Mountain View, CA). In another example, the system is run on Linux® OS (Linux is a registered trademark of Linus Torvalds of Boston, MA). The application is flexible and designed to run in various different environments without compromising any major functionality.

[0106] As used herein, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural elements or steps, unless such exclusion is explicitly recited. Furthermore, references to “example” or “one example” of the present disclosure are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. Further, to the extent that terms “includes,” “including,” “has,” “contains,” and variants thereof are used herein, such terms are intended to be inclusive in a manner similar to the term “comprises” as an open transition word without precluding any additional or other elements.

[0107] In some embodiments, numbers expressing quantities of ingredients, properties such as molecular weight, reaction conditions, and so forth, used to describe and claim certain embodiments of the present disclosure are to be understood as being modified in some instances by the term “about.” In some embodiments, the term “about” is used to indicate that a value includes the standard deviation of the mean for the device or method being employed to determine the value. In some embodiments, the numerical parameters set forth in the written description and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameters should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of some embodiments of the present disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable. The numerical values presented in some embodiments of the present disclosure may contain certain errors necessarily resulting from the standard deviation found in their respective testing measurements. The recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. The recitation of discrete values is understood to include ranges between each value.

[0108] Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where the event occurs and instances where it does not.

[0109] The terms “comprise,” “have” and “include” are open-ended linking verbs. Any forms or tenses of one or more of these verbs, such as “comprises,” “comprising,” “has,” “having,” “includes” and “including,” are also open-ended. For example, any method that “comprises,” “has” or “includes” one or more steps is not limited to possessing only those one or more steps and can also cover other unlisted steps. Similarly, any composition or device that “comprises,” “has” or “includes” one or more features is not limited to possessing only those one or more features and can cover other unlisted features.

[0110] Furthermore, as used herein, the term “real-time” refers to at least one of the time of occurrence of the associated events, the time of measurement and collection of predetermined data, the time to process the data, and the time of a system response to the events and the environment. In the examples described herein, these activities and events occur substantially instantaneously.

[0111] All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided with respect to certain embodiments herein is intended merely to better illuminate the present disclosure and does not pose a limitation on the scope of the present disclosure otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the present disclosure. Groupings of alternative elements or embodiments of the present disclosure disclosed herein are not to be construed as limitations. Each group member can be referred to and claimed individually or in any combination with other members of the group or other elements found herein. One or more members of a group can be included in, or deleted from, a group for reasons of convenience or patentability. When any such inclusion or deletion occurs, the specification is herein deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.

[0112] In some embodiments, the system includes multiple components distributed among a plurality of computer devices. One or more components may be in the form of computer-executable instructions embodied in a computer-readable medium. The systems and processes are not limited to the specific embodiments described herein. In addition, components of each system and each process can be practiced independent and separate from other components and processes described herein. Each component and process can also be used in combination with other assembly packages and processes. The present embodiments may enhance the functionality and functioning of computers and / or computer systems.

[0113] The computer-implemented methods discussed herein can include additional, less, or alternate actions, including those discussed elsewhere herein. The methods can be implemented via one or more local or remote processors, transceivers, servers, and / or sensors (such as processors, transceivers, servers, and / or sensors mounted on vehicles or mobile devices, or associated with smart infrastructure or remote servers), and / or via computer-executable instructions stored on non-transitory computer-readable media or medium. Additionally, the computer systems discussed herein can include additional, less, or alternate functionality, including that discussed elsewhere herein. The computer systems discussed herein can include or be implemented via computer-executable instructions stored on non-transitory computer-readable media or medium.

[0114] As used herein, the term “non-transitory computer-readable media” is intended to be representative of any tangible computer-based device implemented in any method or technology for short-term and long-term storage of information, such as, computer-readable instructions, data structures, program modules and sub-modules, or other data in any device. Therefore, the methods described herein can be encoded as executable instructions embodied in a tangible, non-transitory, computer readable medium, including, without limitation, a storage device and / or a memory device. Such instructions, when executed by a processor, cause the processor to perform at least a portion of the methods described herein. Moreover, as used herein, the term “non-transitory computer-readable media” includes all tangible, computer-readable media, including, without limitation, non-transitory computer storage devices, including, without limitation, volatile and nonvolatile media, and removable and non-removable media such as a firmware, physical and virtual storage, CD-ROMs, DVDs, and any other digital source such as a network or the Internet, as well as yet to be developed digital means, with the sole exception being a transitory, propagating signal.

[0115] Additionally or alternatively, the machine learning programs may be trained by inputting sample data sets or certain data into the programs, such as images, object statistics and information, traffic timing, previous trips, and / or actual timing. The machine learning programs may utilize deep learning algorithms that may be primarily focused on pattern recognition, and may be trained after processing multiple examples. The machine learning programs may include Bayesian Program Learning (BPL), voice recognition and synthesis, image or object recognition, signal processing, optical character recognition, and / or natural language processing - either individually or in combination. The machine learning programs may also include natural language processing, semantic analysis, automatic reasoning, and / or machine learning.

[0116] Supervised and unsupervised machine learning techniques may be used. In supervised machine learning, a processing element may be provided with example inputs and their associated outputs, and may seek to discover a general rule that maps inputs to outputs, so that when subsequent novel inputs are provided the processing element may, based upon the discovered rule, accurately predict the correct output. In unsupervised machine learning, the processing element may be required to find its own structure in unlabeled example inputs. In one embodiment, machine learning techniques may be used to determine brain responses to stimuli such as VNS settings. Based upon these analyses, the processing element may learn how to identify characteristics and patterns that may then be applied to analyzing image data, model data, and / or other data. For example, the processing element may learn, to identify glucoses responses to stimuli and the VNS settings for different patients to provide optimal glucose levels. The processing element may also learn how to identify trends that may not be readily apparent based upon collected traffic data, such as trends that simplify identification of groups of metabolites.

[0117] Definitions and methods described herein are provided to better define the present disclosure and to guide those of ordinary skill in the art in the practice of the present disclosure. Unless otherwise noted, terms are to be understood according to conventional usage by those of ordinary skill in the relevant art.

[0118] Although specific features of various embodiments may be shown in some drawings and not in others, this is for convenience only. In accordance with the principles of the systems and methods described herein, any feature of a drawing may be referenced or claimed in combination with any feature of any other drawing.

[0119] This written description uses examples to disclose the invention, including the best mode, and also to enable any person skilled in the art to practice the invention, including making and using any devices or systems and performing any incorporated methods. The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal language of the claims.

Claims

IN THE CLAIMS1. A computer device to compute confidence scores for metabolite identification, wherein the computer device comprises at least one processor in communication with at least one memory device, wherein the at least one processor is programmed to:receive experimental data for a subject to be identified;calculate similarity scores based on the experimental data;determine at least one probability of achieving the similarity scores based upon a plurality of historical identifications;compare the at least one probability of achieving the similarity scores to the plurality of historical identifications; anddetermine at least one metabolite identification for the subject based upon the comparison.

2. The computer device of Claim 1 , wherein the experimental data include mass spectrometry data from a sample of the subject to be identified.

3. The computer device of Claim 1 , wherein the at least one processor is further programmed to calculate a degree of match to the experimental data.

4. The computer device of Claim 3, wherein the at least one processor is further programmed to calculate a posterior probability that a metabolite identification is correct based upon the degree of match to the experimental data relative to a plurality of other potential metabolites.

5. The computer device of Claim 4, wherein the at least one processor is further programmed to calculate a posterior probability that a metabolite identification is correct based upon the degree of match to the experimental data relative to a plurality of historical metabolite identification probabilities.

6. The computer device of Claim 1 , wherein the at least one processor is further programmed to perform Bayesian modeling of metabolite identification probabilities.

7. The computer device of Claim 1 , wherein the at least one processor is further programmed to calculate a posterior probability that a metabolite identification is correct, wherein the posterior probability is quantified with Bayesian modeling.

8. The computer device of Claim 1 , wherein the similarity scores include a degree of tandem mass spectrometry match, a degree of retention time match, and a degree of natural abundance isotope pattern match.

9. The computer device of Claim 1 , wherein the similarity scores include at least one of entropy similarity, retention time deviation, and reverse dot product.

10. The computer device of Claim 1 , wherein at least one processor is further programmed to utilize probability density functions to convert similarity scores into probability values.

11. The computer device of Claim 1 , wherein the at least one processor is further programmed to perform biomarker selection for downstream validation based on the at least one metabolite identification for the subject.

12. The computer device of Claim 1 , wherein the at least one metabolite identification is based upon a highest confidence score.

13. The computer device of Claim 1 , wherein the at least one processor is further programmed to benchmark different platforms or workflows for performing metabolomics by counting a number of identified metabolites at a defined false discovery rate.

14. A computer implemented method to compute confidence scores for metabolite identification, wherein the method is implemented on a computer device comprising at least one processor in communication with at least one memory device, wherein the at least one processor is programmed to:receiving experimental data for a subject to be identified;calculating similarity scores based on the experimental data;determining at least one probability of achieving the similarity scores based upon a plurality of historical identifications;comparing the at least one probability of achieving the similarity scores to the plurality of historical identifications; anddetermining at least one metabolite identification for the subject based upon the comparison.

15. The computer implemented method of Claim 14, wherein the experimental data include mass spectrometry data from a sample of the subject to be identified.

16. The computer implemented method of Claim 14 further comprising calculating a degree of match to the experimental data.

17. The computer implemented method of Claim 16 further comprising calculating a posterior probability that a metabolite identification is correct based upon the degree of match to the experimental data relative to a plurality of other potential metabolites.

18. The computer implemented method of Claim 17 further comprising calculating a posterior probability that a metabolite identification is correct based upon the degree of match to the experimental data relative to a plurality of historical metabolite identification probabilities.

19. The computer implemented method of Claim 14 further comprising performing Bayesian modeling of metabolite identification probabilities.

20. The computer implemented method of Claim 14 further comprising calculating a posterior probability that a metabolite identification is correct, wherein the posterior probability is quantified with Bayesian modeling.

21. The computer implemented method of Claim 14, wherein the similarity scores include a degree of tandem mass spectrometry match, a degree of retention time match, and a degree of natural abundance isotope pattern match.

22. The computer implemented method of Claim 14, wherein the similarity scores include at least one of entropy similarity, retention time deviation, and reverse dot product.

23. The computer implemented method of Claim 14 further comprising utilizing probability density functions to convert similarity scores into probability values.

24. The computer implemented method of Claim 14 further comprising performing biomarker selection for downstream validation based on the at least one metabolite identification for the subject.

25. The computer implemented method of Claim 14, wherein the at least one metabolite identification is based upon a highest confidence score.

26. The computer implemented method of Claim 14 further comprising benchmarking different platforms or workflows for performing metabolomics by counting a number of identified metabolites at a defined false discovery rate.