Methods and systems of mass spectrometry analysis for precision medicine

The two-pass machine learning algorithm directly processes original mass spectrometry data to address the limitations of peak identification methods, preserving comprehensive biomarker information and improving classification accuracy through non-linear analysis.

WO2026120517A1PCT designated stage Publication Date: 2026-06-11PREVIEW HEALTH PTY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PREVIEW HEALTH PTY LTD
Filing Date
2025-12-03
Publication Date
2026-06-11

Smart Images

  • Figure IB2025062396_11062026_PF_FP_ABST
    Figure IB2025062396_11062026_PF_FP_ABST
Patent Text Reader

Abstract

Described herein is a method of processing and analyzing mass spectrometry data in its original state with minimal to no data reduction for analysis directly by a machine learning classifier. This may be for the purposes of diagnosing, detecting, classifying, predicting, stratifying, screening or monitoring one or more medical conditions or disease states, and discovering biomarker signatures.
Need to check novelty before this filing date? Find Prior Art

Description

WSGR Docket No. 70030-701.601METHODS AND SYSTEMS OF MASS SPECTROMETRY ANALYSIS FOR PRECISION MEDICINECROSS-REFERENCE

[0001] This application claims the benefit of US Provisional Application Serial Number 63 / 727,907 filed on December 4, 2024, US Provisional Application Serial Number63 / 800,906 filed on May 6, 2025, and US Provisional Application Serial Number 63 / 880,830 filed on September 12, 2025, each of which is incorporated by reference herein in its entirety.BACKGROUND

[0002] The identification of molecular biomarkers may be important for understanding and diagnosing medical conditions and disease states. Molecular biomarkers can include small molecules such as metabolites, lipids and larger biomolecules including peptides, proteins, and nucleic acids. These biomarkers can be detected in various biological fluids and matrices, including blood, urine, tissue, saliva, or breath, to provide valuable insights into pathological processes.SUMMARY

[0003] In current methods, biomarkers may be identified by a mixture of heuristics and manual analysis but these approaches can come with limitations. First, there can be a critical dependency on a highly targeted biomarker identification approach as the basis for analysis. However, this default reliance on manually identified biomarkers can oversimplify the thousands of complex biological processes that occur simultaneously in the human body. In current methods, biomarkers with the most obvious presentation may be confidently detected and used; however, over 99% of molecular data may be discarded.

[0004] Second, in current methods (for example, regression-based models) model performance can be overly reliant on the linear separability of data derived from prior biomarker analysis. A major limitation to this approach is that these methods can fail to capture complex biological effects that may be non-linear in nature and far more complex on an individual biological level.WSGR Docket No. 70030-701.601

[0005] Therefore, a computational method is needed that can both preserve up to 100% of molecular data and consider non-linear biological effects to identify more accurate biomarkers.

[0006] Mass spectrometry (MS) is an analytical technique that may be capable of detecting hundreds to thousands of molecules in a single biological sample. However, the resulting data may be complex due to the vast number of detected molecules. This includes many molecules that are detected in low abundances that can contribute to ‘noisy’ data and may not provide meaningful information regarding medical conditions, disease states or in providing insights into biological processes.

[0007] Current approaches for analyzing mass spectrometry data may rely on peak identification methods. Such methods may focus on detecting significant peaks in a mass spectrum, corresponding to ion abundances at specific mass-to-charge ratios (m / z) and specific times of detection, which may correspond to chromatography elution times. Peaks can be identified using various methods including comparing ion fragmentation data to databases of fragmentation spectra and data from authentic standards. Identified peaks may be curated into abridged lists of molecules and their respective abundances, which may be used to differentiate “control” (negative) groups from “disease” (positive) groups. Linear statistical methods such as p-value and fold change may be used to identify biomarkers associated with negative or positive groups. This approach may allow the generation of biomarker profiles associated with specific disease states, such as in a discovery phase of the development of a disease biomarker, or panel of biomarkers. Potential biomarkers can progress to validation in which targeted methods are developed to quantify the biomarkers in the biological sample, where such information can then be used to support clinical decision making.

[0008] Despite their utility, peak identification methods may encounter significant limitations. They may rely on approximate heuristics designed to reduce the complexity of the data. While this simplification makes the data more manageable, information can be lost particularly for low abundance ions and ions that cannot be readily identified without further experiments. Thus, the loss of information can result in important potential biomarkers being missed using standard approaches, which can lead to reduced classification accuracy.

[0009] Current methods of analyzing mass spectrometry data face challenges in using up to 100% of original mass spectrometry data. For example, original mass spectrometry data can contain artefacts and / or low abundance background ions that can overshadow low abundance ions that may be relevant as disease biomarkers. To overcome these challenges,WSGR Docket No. 70030-701.601 the present disclosure provides a method of processing and analyzing original mass spectrometry data using a ‘two-pass’ machine learning algorithm.

[0010] Consequently, there remains an unfulfilled need for analytical approaches that can process complex mass spectrometry data with high fidelity, while leveraging advanced computational methods such as machine learning algorithms to enable robust disease prediction and biomarker discovery.

[0011] Recognizing the above-mentioned needs, the present disclosure provides a method of processing and analyzing mass spectrometry data in its original state with minimal to no data reduction for analysis directly by a machine learning classifier. The methods may comprise diagnosing, detecting, classifying, predicting, stratifying, screening, and / or monitoring one or more medical conditions and / or disease states, and discovering biomarker signatures.

[0012] This approach differs from current approaches in that a prior peak identification method may not be needed to process the mass spectrometry data for use with a machine learning classifier.

[0013] The methods and systems of the present disclosure provide various benefits. For example, by minimizing pre-processing or other heuristic operations needed to process mass spectrometry data, greater molecular information can be preserved in its original data form prior to using a machine learning classifier. Therefore, these methods can: inform new biomarker discovery by capturing potentially important biomarkers that are directly targeted to applications which current methods may miss; identify biomarkers that may reveal new biological mechanisms of action or biological interactions to help stratify disease phenotypes; identify more accurate and targeted predictive biomarkers by leveraging untargeted mass spectrometry data; achieve classification accuracy that is comparable or higher than current methods involving prior peak identification methods; and increase scalability as this method is self-reliant and does not require external heuristics to operate.

[0014] In some embodiments, the first pass of the two-pass machine learning algorithm comprises applying up to 100% of original mass spectrometry data with a machine learning algorithm, wherein the full m / z range is analyzed.

[0015] In some embodiments, the second pass of the two-pass machine learning algorithm comprises ‘zooming in’ on the original mass spectrometry data by applying one orWSGR Docket No. 70030-701.601 more targeted m / z ranges. A machine learning algorithm is then applied to the original mass spectrometry data within the one or more targeted m / z ranges.

[0016] The present disclosure provides a computer-implemented system whereby direct injection or hyphenated mass spectrometry data in its original state can be processed directly from a mass spectrometer and used as input data for a machine learning classifier.

[0017] In an aspect, the present disclosure provides a method of processing the original mass spectrometry data file directly from a mass spectrometer as input data for a machine learning classifier.

[0018] As examples, original data files may include .raw, ,wiff2, .D, .LCD, .CDF, .mzXML and mzML.

[0019] In another aspect, the present disclosure provides a method for developing a machine learning classifier based on the processing and analyzing of original direct injection or hyphenated mass spectrometry data from one or more biological sample / s (e.g., as directly obtained from a mass spectrometer either independently or in conjunction with non-chemical data).

[0020] In some embodiments, the method comprises developing a machine learning classifier based on direct injection or hyphenated mass spectrometry data containing metabolomics, proteomics, and / or lipidomics data with minimal to no pre-processing, for example, data reduction. In some embodiments, the method comprises applying the machine learning classifier to an unknown biological sample, correlating the biomarker signature with one or more medical conditions, and generating a prediction or probability score.

[0021] In some embodiments, the medical conditions comprise oncological, immunological, neurological, respiratory, cardiovascular, gastrointestinal, and mental health disease states and / or disorders, or viral and bacterial infections, or their respective clinical phenotypes.

[0022] In some embodiments, the biological samples comprise whole blood, blood serum, blood plasma, saliva, breath, urine, feces, sebum, tissue, cerebrospinal fluid, seminal fluid, vaginal secretion, amniotic fluid, nasal fluid, otic fluid, interstitial fluid, breast milk, tears, sputum, sweat, or hair (e.g., as belonging to human or animal species).

[0023] In some embodiments, the direct injection MS methods comprise ESI-MS, ESI- MS / MS, MALDI-MS, MALDI-MS / MS, EI-MS, EI-MS / MS, CI-MS, CI-MS / MS, APCI-MS, APCI-MS / MS, DART-MS, DART-MS / MS, DESI-MS, DESI-MS / MS, ICP-MS, ICP- MS / MS, FAB-MS, FAB-MS / MS, SIMS, or SIMS / MS.WSGR Docket No. 70030-701.601

[0024] In some embodiments, the hyphenated MS methods comprise GC-MS, GC- MS / MS, LC-MS, LC-MS / MS, UHPLC-MS, UHPLC-MS / MS, HPLC-MS, HPLC-MS / MS, HILIC-MS, HILIC-MS / MS, CE-MS, or CE-MS / MS.

[0025] In some embodiments, the classes of biomarkers comprise metabolites, nucleic acids, proteins, peptides, or lipids, or fragments thereof.

[0026] In some embodiments, the non-chemical data comprises biological sex, age, ethnicity, height, weight, diet, symptomatic information, family history, medication use, biometric data, genetic history, or electronic health records.

[0027] In some embodiments, the machine learning classifiers comprise linear regression, linear discriminant analysis, logistic regression, random forest, support vector machines, extreme gradient boosting, convolutional neural networks, or other neural network methods.

[0028] In another aspect, the present disclosure provides a method of visualizing mass spectrometry data in a one-dimensional array as input data for a machine learning classifier. In some embodiments, the ion abundances of each mass-to-charge ratio (m / z) is aggregated over time, resulting in a single summed abundance value for each m / z value.

[0029] In some embodiments, the ion abundances of each m / z that have been summed over time are further aggregated based on specific m / z ranges (buckets).

[0030] In another aspect, the present disclosure provides a method of visualizing mass spectrometry data as a two-dimensional matrix (e.g., a heatmap) as input data for a machine learning classifier.

[0031] In some embodiments, the x-axis is represented as m / z, the y-axis is represented as time (t), and where the value for each (m / z, t) pair is the ion abundance value corresponding to the m / z and time value. In some embodiments, ion abundances of each m / z are further aggregated based on specific m / z ranges (buckets). In other embodiments, the ion abundances of each m / z bucket are further aggregated based on specific time buckets.

[0032] In some embodiments, the x-axis is represented as m / z, the y-axis is represented as m / z, and where the value for each (m / z, m / z) pair is the ion abundance value corresponding to m / z (from unfragmented ion data) and m / z (from fragmented ion data) value. In some embodiments, ion abundances of each m / z is further aggregated based on specific m / z ranges (buckets).

[0033] In another aspect, the present disclosure provides a method for using the machine learning classifier to identify a biomarker signature comprising one or more biomarkers that are correlated with one or more medical conditions by directly analyzing mass spectrometry data in its original state.WSGR Docket No. 70030-701.601

[0034] In some embodiments, the biomarker signature is determined from Shapley Additive ExPlanation (SHAP), statistical significance, or feature importance methods, where biomarker discovery is optimized using a machine learning classifier.

[0035] In some embodiments, the biomarker signature comprises absolute abundance or relative ratio of one or more biomarkers.

[0036] In another aspect, the present disclosure provides a method for identifying biological interactions between multiple m / z features based on the interpretation of biomarker signatures.

[0037] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.

[0038] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.

[0039] Provided herein in some embodiments is a method of analyzing original mass spectrometry data, comprising (a) performing a first pass comprising analyzing the original mass spectrometry data, and (b) performing a second pass comprising analyzing a targeted mass-to-charge ratio (m / z) range of the original spectrometry data.

[0040] In some embodiments, performing the second pass comprising analyzing the targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data is based at least in part on an output of the first pass.

[0041] In some embodiments, a method provided herein further comprises using a machine learning algorithm.

[0042] In some embodiments, the original mass spectrometry data comprises a full m / z range. In some embodiments, the original mass spectrometry data comprises a onedimensional array or a two-dimensional matrix. In some embodiments, the first pass comprises applying a machine learning classifier to the original mass spectrometry data. In some embodiments, the second pass comprises analyzing one or more targeted m / z range of the original mass spectrometry data. In some embodiments, the second pass comprises applying a machine learning classifier to the one or more targeted m / z range of the original mass spectrometry data.

[0043] In some embodiments, a method provided herein further comprises performing a third pass comprising analyzing a second targeted mass-to-charge ratio (m / z) range of theWSGR Docket No. 70030-701.601 original mass spectrometry data. In some embodiments, a method provided herein further comprises performing a fourth pass comprising analyzing a third targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data.

[0044] In some embodiments, the original mass spectrometry data is obtained using a direct injection mass spectrometry or a hyphenated mass spectrometry. In some embodiments, the direct injection mass spectrometry comprises one or more of electrospray ionization-mass spectrometry (ESI-MS), matrix-assisted laser desorption ionization-mass spectrometry (MALDI-MS), electron ionization-mass spectrometry (EI-MS), chemical ionization-mass spectrometry (CI-MS), atmospheric pressure chemical ionization-mass spectrometry (APCI-MS), direct analysis in real time-mass spectrometry (DART -MS), desorption electrospray ionization-mass spectrometry (DESI-MS), inductively coupled plasma-mass spectrometry (ICP-MS), fast atom bombardment-mass spectrometry (FAB-MS), and secondary ion mass spectrometry (SIMS). In some embodiments, the direct injection mass spectrometry comprises a tandem mass spectrometry (MS / MS), the MS / MS comprising one or more of ESI-MS / MS, MALDI-MS / MS, EI-MS / MS, CI-MS / MS, APCI-MS / MS, DART -MS / MS, DESI-MS / MS, ICP-MS / MS, FAB-MS / MS, and SIMS / MS. In some embodiments, the hyphenated mass spectrometry comprises one or more of gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS), high-performance liquid chromatography -mass spectrometry (HPLC-MS), hydrophilic interaction chromatography-mass spectrometry (HILIC-MS), and capillary electrophoresismass spectrometry (CE-MS). In some embodiments, the hyphenated mass spectrometry comprises a tandem mass spectrometry (MS / MS) method, the MS / MS method comprising one or more of GC-MS / MS, LC-MS / MS, UHPLC-MS / MS, HPLC-MS / MS, HILIC-MS / MS, and CE-MS / MS.

[0045] In some embodiments, the machine learning algorithm comprises one or more of linear regression, linear discriminant analysis, logistic regression, random forest, support vector machines, extreme gradient boosting, recurrent neural networks, convolutional neural networks, and transformers. In some embodiments, the machine learning algorithm comprises a neural network. In some embodiments, a method provided herein further comprises using the machine learning algorithm to detect, classify, predict, stratify, screen, or monitor a disease, disorder, or condition. In some embodiments, the machine learning algorithm is trained using training data. In some specific embodiments, the training data is obtained from a subject having or suspected of having the disease, disorder, or condition.WSGR Docket No. 70030-701.601

[0046] In some embodiments, a method provided herein comprises determining a probability score.

[0047] In some embodiments, the probability score is based at least in part on a detection, classification, or prediction of the machine learning algorithm of a biological sample. In some embodiments, the biological sample is a sample of unknown origin or an anonymized sample obtained or derived from a subject of a plurality of subjects. In some embodiments, the biological sample comprises one or more of whole blood, blood serum, blood plasma, saliva, breath, urine, feces, sebum, tissue, cerebrospinal fluid, seminal fluid, vaginal secretion, amniotic fluid, nasal fluid, otic fluid, interstitial fluid, breast milk, tears, sputum, sweat, and hair.

[0048] In some embodiments, a method provided herein further comprises using the machine learning algorithm to generate a performance metric. In some embodiments, the performance metric comprises one or more of accuracy, sensitivity, specificity, precision, positive predictive value (PPV), negative predictive value (NPV), Mathew’s correlation coefficient (MCC), area under the receiver operating characteristic curve (AUROC), and area under the precision recall curve. In some specific embodiments, the area under the receiver operating characteristic curve (AUROC) is 0.50 or higher.

[0049] In some embodiments, a method provided herein comprises: (a) obtaining a first biological sample of a subject; (b) obtaining a first biological sample of a subject; (c) analyzing the original mass spectrometry data, thereby obtaining processed mass spectrometry data; (d) constructing a machine learning classifier based at least in part on the processed mass spectrometry data; (e) using the machine learning classifier to identify a biomarker signature comprising one or more biomarker that is correlated with a disease, disorder, or condition; and (f) applying the machine learning classifier to a second biological sample, thereby correlating the biomarker signature with the disease, disorder, or condition. In some embodiments, analyzing the original mass spectrometry data comprises: (a) performing a first pass comprising analyzing the original mass spectrometry data, and (b) performing a second pass comprising analyzing a targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data. In some embodiments, (d) constructing a machine learning classifier based at least in part on the processed mass spectrometry data further comprises constructing the machine learning classifier based at least in part on subject data.

[0050] In some embodiments, the subject data comprises one or more of biological sex, age, ethnicity, height, weight, diet, symptomatic information, family history, medication use, genetic history, and electronic health records.WSGR Docket No. 70030-701.601

[0051] In some embodiments, the first or the second biological sample is a sample of unknown origin or an anonymized sample. In some embodiments, the second biological sample is obtained or derived from the first subject or from a second subject. In some embodiments, the first subject or the second subject is a mammal. In some specific embodiments, the first subject or the second subject is a human.

[0052] In some embodiments, the disease, disorder, or condition comprises one or more of an oncological, immunological, neurological, respiratory, cardiovascular, gastrointestinal, or mental health disease, disorder, or condition. In some embodiments, the disease, disorder, or condition comprises a phenotype. In some embodiments, the disease, disorder, or condition comprises a viral infection or a bacterial infection.

[0053] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.

[0054] the first or the second biological sample comprises one or more of whole blood, blood serum, blood plasma, saliva, breath, urine, feces, sebum, tissue, cerebrospinal fluid, seminal fluid, vaginal secretion, amniotic fluid, nasal fluid, otic fluid, interstitial fluid, breast milk, tears, sputum, sweat, and hair.

[0055] In some embodiments, (b) performing a second pass comprising analyzing a targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data comprises using a direct injection mass spectrometry or a hyphenated mass spectrometry. In some embodiments, the direct injection mass spectrometry comprises one or more of electrospray ionization-mass spectrometry (ESI-MS), matrix-assisted laser desorption ionization-mass spectrometry (MALDI-MS), electron ionization-mass spectrometry (EI-MS), chemical ionization-mass spectrometry (CI-MS), atmospheric pressure chemical ionization-mass spectrometry (APCI-MS), direct analysis in real time-mass spectrometry (DART -MS), desorption electrospray ionization-mass spectrometry (DESI-MS), inductively coupled plasma-mass spectrometry (ICP-MS), fast atom bombardment-mass spectrometry (FAB-MS), and secondary ion mass spectrometry (SIMS). In some embodiments, the direct injection mass spectrometry comprises a tandem mass spectrometry (MS / MS), the MS / MS comprising one or more of ESI-MS / MS, MALDI-MS / MS, EI-MS / MS, CI-MS / MS, APCI-MS / MS,WSGR Docket No. 70030-701.601DART -MS / MS, DESI-MS / MS, ICP-MS / MS, FAB-MS / MS, and SIMS / MS. In some embodiments, the hyphenated mass spectrometry comprises one or more of gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS), high-performance liquid chromatography -mass spectrometry (HPLC-MS), hydrophilic interaction chromatography-mass spectrometry (HILIC-MS), and capillary electrophoresismass spectrometry (CE-MS). In some embodiments, the hyphenated mass spectrometry comprises a tandem mass spectrometry (MS / MS) method, the MS / MS method comprising one or more of GC-MS / MS, LC-MS / MS, UHPLC-MS / MS, HPLC-MS / MS, HILIC-MS / MS, and CE-MS / MS.

[0056] In some embodiments, a biomarker provided herein comprises one or more of a metabolite, nucleic acid, protein, peptide, and lipid. In some embodiments, the biomarker comprises a feature. In some embodiments, the biomarker comprises a chemical or a nonchemical feature. In some embodiments, the biomarker comprises an ion feature. In some specific embodiments, the ion feature is detected in a positive or negative ionization mode of the mass spectrometry. In some embodiments, the biomarker is obtained using untargeted mass spectrometry or targeted mass spectrometry.

[0057] In some embodiments, a biomarker signature is determined using one or more of Shapley Additive ExPlanation (SHAP), statistical significance, and feature importance procedure. In some embodiments, the biomarker signature is optimized using the machine learning classifier. In some embodiments, the biomarker signature comprises an absolute abundance or a relative ratio of the biomarker. In some embodiments, the biomarker signature comprises a panel of chemical or non-chemical features. In some embodiments, wherein the biomarker signature comprises an ion heatmap based on m / z values with respect to time. In some embodiments, the biomarker signature comprises an ion heatmap based on m / z values from mass spectrometry data of a first mass spectrometry stage (MSI; mi / zi) with respect to m / z values from mass spectrometry data of a second mass spectrometry stage (MS2; m2 / z2).

[0058] In some embodiments, a method provided herein further comprises using the biomarker signature to monitor disease, disorder, or condition progression.

[0059] In some embodiments, a method provided herein further comprises using the biomarker signature to monitor the response of a subject to a disease-modifying therapy or intervention.WSGR Docket No. 70030-701.601

[0060] In some embodiments, a method provided herein further comprises using the biomarker signature to identify predisposition of the subject to the disease, disorder, or condition.

[0061] In some embodiments, a method provided herein further comprises using the biomarker signature to identify different phenotypes of the disease, disorder, or condition.

[0062] In some embodiments, a method provided herein further comprises identifying a disease mechanism or a druggable target based at least in part on analyzing the original mass spectrometry data.

[0063] In some embodiments, a machine learning classifier provided herein comprises one or more of a linear regression, linear discriminant analysis, logistic regression, random forest, support vector machines, extreme gradient boosting, recurrent neural network, convolutional neural network, and a transformer. In some embodiments, the machine learning classifier comprises a neural network. In some embodiments, the machine learning classifier is trained using training data. In some embodiments, the training data is obtained or derived from a subject having or suspected of having the disease, disorder, or condition. In some embodiments, a method provided herein further comprises using the machine learning classifier to detect, classify, predict, stratify, screen, or monitor the disease, disorder, or condition based on the training of data obtained or derived from the subject having or suspected of having the disease, disorder, or condition.

[0064] In some embodiments, a method provided herein further comprises determining a probability score. In some embodiments, the probability score is based at least in part on a detection, classification, or prediction of the machine learning classifier of the second biological sample. In some embodiments, the machine learning classifier is used to generate a classification performance metric comprising one or more of accuracy, sensitivity, specificity, precision (or positive predictive value, PPV), negative predictive value (NPV), Mathew’s correlation coefficient (MCC), area under the receiver operating characteristic curve, and area under the precision recall curve. In some specific embodiments, the area under the receiver operating characteristic curve (AUROC) is 0.50 or higher.

[0065] In some embodiments, the original mass spectrometry data is an input data for a machine learning classifier. In some embodiments, the original mass spectrometry data comprises one or more of data in a .raw, ,wiff2, .D, .LCD, .CDF, .mzXML, or .mzML format.

[0066] In some embodiments, a method provided herein further comprises processing the original mass spectrometry data using a machine learning classifier, wherein the original mass spectrometry data comprise a one-dimensional data array. In some embodiments, an ionWSGR Docket No. 70030-701.601 abundance of the original mass spectrometry data for an m / z value is summed over time, providing a single summed abundance value for the m / z value. In some embodiments, the summed abundance value is associated with an output indicative of the presence of a disease, disorder, or condition. In some embodiments, the ion abundance of the m / z value is summed over an m / z range, providing a summed abundance value for the m / z range.

[0067] In some embodiments, provided herein is a method of processing original mass spectrometry data, the method comprising, processing the original mass spectrometry data using a machine learning classifier, wherein the original mass spectrometry data comprise a two-dimensional matrix.

[0068] In some embodiments, the x-axis represents m / z and the y-axis represents time (t), and wherein the value for each (m / z, t) pair is the ion abundance value corresponding to the (m / z, t) pair.

[0069] In some embodiments, the x-axis represents an m / z comprising first stage mass spectrometry data (MSI; mi / zi) and the y-axis represents a second m / z comprising second stage mass spectrometry data (MS2; m2 / z2), and wherein the value for each (mi / zi, m2 / z2) pair is the ion abundance value corresponding to the (mi / zi, m2 / z2 ) pair.

[0070] In some embodiments, the x-axis represents an m / z comprising second stage mass spectrometry data (MS2; m2 / z2) and the y-axis represents a second m / z comprising first stage mass spectrometry data (MSI; mi / zi), and wherein the value for each (m2 / z2, mi / zi) pair is the ion abundance value corresponding to the (m2 / z2, mi / zi) pair.

[0071] In some embodiments, an ion abundance of an m / z value of the original mass spectrometry data is summed over an m / z range, providing a summed abundance value for the m / z range. In some embodiments, the ion abundance of the m / z range of the original mass spectrometry data is summed over a second m / z range, providing a summed abundance value for the second m / z range. In some embodiments, the ion abundance of the m / z value of the original mass spectrometry data is summed over a time range, providing a summed abundance value for the time range with respect to the m / z value. In some embodiments, the ion abundance of the m / z value is further summed over an m / z range, providing a summed abundance value for the m / z range with respect to the time range. In some embodiments, the ion abundance of the m / z range is further summed over a second m / z range, providing a summed abundance value for the second m / z range with respect to the time range. In some embodiments, the ion abundance of the time range is further summed over a second time range, providing a summed abundance value for the second time range with respect to the m / z range.WSGR Docket No. 70030-701.601BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0073] FIG. 1 provides an example of a method of the present disclosure.

[0074] FIG. 2 illustrates a non-limiting example of how a mass spectrum can be depicted as a 1-dimensional array.

[0075] FIG. 3 illustrates a non-limiting example of a 1 -dimensional tabular array of a mass spectrum.

[0076] FIG. 4 illustrates a non-limiting example of how a mass spectrum can be depicted as a 2-dimensional heatmap.

[0077] FIG. 5 illustrates a non-limiting example of how mass spectrometry data can be processed and analyzed using a two-pass machine learning algorithm.

[0078] FIG. 6 shows a confusion matrix obtained from the multi-classification model using 1 -dimensional mass spectrometry data; malaria (KM), visceral Leishmaniasis (VL), and Zika virus (ZIK).

[0079] FIG. 7 shows a confusion matrix obtained from the multi-classification model using 2-dimensional mass spectrometry data; malaria (KM), visceral Leishmaniasis (VL), and Zika virus (ZIK).

[0080] FIG. 8 shows an Area Under the Receiver Operating Characteristic (AUROC) curve obtained from applying a machine learning classifier to 1 -dimensional data.

[0081] FIG. 9 shows a distribution plot of m / z features with respect to Shapley Additive ExPlanation (SHAP) value.

[0082] FIG. 10 shows an AUROC curve obtained from applying a machine learning classifier to 2-dimensional data.

[0083] FIG. 11 shows a non-limiting example of a biomarker signature in the form of an ion heatmap.

[0084] FIG. 12 shows a SHAP plot based on combinatorial analysis of top m / z features.WSGR Docket No. 70030-701.601

[0085] FIG. 13 shows an AUROC curve obtained from applying a machine learning classifier to 1 -dimensional data.

[0086] FIG. 14 shows an AUROC curve obtained from applying a machine learning classifier to 1 -dimensional data.

[0087] FIG. 15 shows an AUROC curve obtained from applying a machine learning classifier to 2-dimensional data.

[0088] FIG. 16 shows a confusion matrix obtained from the multi-classification model using 1 -dimensional data; 0 (healthy control), 1 (endometrial polyps), 2 (endometrial cancer), and 3 (endometrial hyperplasia).

[0089] FIG. 17 shows a confusion matrix obtained from the multi-classification model using 2-dimensional data.

[0090] FIG. 18 shows an AUROC curve obtained from applying a machine learning classifier to 2-dimensional data.

[0091] FIG. 19 shows a confusion matrix obtained from the multi-classification model using 2-dimensional data.

[0092] FIG. 20 shows an AUROC curve obtained from applying a machine learning classifier to 1 -dimensional data.

[0093] FIG. 21 shows a distribution plot of m / z features with respect to SHAP value.

[0094] FIG. 22 shows a distribution plot of m / z features with respect to SHAP value.

[0095] FIG. 23 shows an AUROC curve obtained from applying a machine learning classifier to 2-dimensional data.

[0096] FIG. 24 illustrates a non-limiting example of a fragmentation-based biomarker signature.

[0097] FIG. 25A and FIG. 25B illustrate a non-limiting examples of a time-dependent biomarker signature based on before (FIG. 25A) and after (FIG. 25B) intervention.DETAILED DESCRIPTION

[0098] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.WSGR Docket No. 70030-701.601

[0099] As used in the specification and claims, the singular form “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise.

[0100] Various terms used throughout the present description may be read and understood as follows, unless the context indicates otherwise: “or” as used throughout is inclusive, as though written “and / or”; singular articles and pronouns as used throughout include their plural forms, and vice versa; similarly, gendered pronouns include their counterpart pronouns so that pronouns should not be understood as limiting anything described herein to use, implementation, performance, etc. by a single gender; “exemplary” should be understood as “illustrative” or “exemplifying” and not necessarily as “preferred” over other embodiments. Further definitions for terms may be set out herein; these may apply to prior and subsequent instances of those terms, as are understood from a reading of the present description.Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0101] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0102] The term “subject,” as used herein, generally refers to a human such as a patient. In some embodiments, the subject is a mammal. The subject may be a person (e.g., a patient) with a disease, disorder, or condition, or a person that has been treated for a disease, disorder, or condition, or a person that is being monitored for a disease, disorder, or condition, or a person that is suspected of having the disease, disorder, or condition, or a person that does not have or is not suspected of having the disease, disorder, or condition. The disease, disorder, or condition may be an infectious disease, an immune disorder or disease, an injury, or a rare disease. In some embodiments, the disease, disorder, or condition comprises febrile illnesses, chronic pain, silicosis, endometrial disease, bacterial vaginosis, Parkinson’s disease, and synucleinopathies.

[0103] The present disclosure provides a method of processing and analyzing original mass spectrometry data (e.g., mass spectrometry data in its original state with minimal to no data reduction for analysis directly by a machine learning classifier; raw mass spectrometryWSGR Docket No. 70030-701.601 data; whole mass spectrometry data) for analysis directly by a machine learning classifier to diagnose, classify, predict, stratify, screen and / or monitor medical conditions, and to discover biomarker signatures.Diseases, Disorders, and Conditions

[0104] For biologically relevant questions, it is imperative to be able to distinguish populations with particular diseases, disorders, and conditions from each other and to monitor progression as accurately as possible. Examples of biologically relevant questions include whether a subject has biological indicators of a single or multiple diseases, disorders, or conditions and / or disease states, and whether a subject is responsive to a disease-modifying therapy or intervention. Non-limiting examples of diseases, disorders, and conditions to which methods and systems disclosed herein can be applied include, but are not limited to febrile illnesses, chronic pain, silicosis, endometrial disease, bacterial vaginosis, Parkinson’s disease, and synucleinopathies. The disease, disorder, or condition can comprise one or more of an oncological, immunological, neurological, respiratory, cardiovascular, gastrointestinal, or mental health disease, disorder, or condition. The infectious disease may be caused by bacteria (e.g., gram negative or gram positive bacteria) or viruses (e.g., DNA or RNA viruses). For example, the disease or disorder may comprise a viral or bacterial infection. In some embodiments, the disease, disorder, or condition comprises a phenotype.A. Febrile Illnesses

[0105] Acute fever-related illnesses (acute febrile illnesses) resulting from infection is a leading cause of death and morbidity worldwide, particularly among children and those in low- to middle-income nations. For example, in 2010, 64% of deaths in children less than 5 years old were caused by infections such as malaria (Liu et al., 2012, Lancet, 2151-61; incorporated by reference herein in its entirety). The high death rate is partially due to misdiagnosis and delayed treatment where mistreatment can result in antimicrobial resistance. Therefore, it is imperative to accurately diagnose different infectious causes of acute febrile illnesses.B. Chronic Pain

[0106] Chronic pain is a public health epidemic that affects 20% of the global population (Goldberg et al., 2011, BMC Public Health, 770-774; incorporated by reference herein in its entirety). A major challenge with chronic pain treatment is that individual responses toWSGR Docket No. 70030-701.601 analgesic drugs can vary dramatically owing to differences in etiologies and variations in phenotypic presentations. (Robert et al., 2016, PAIN, 1851-1871; incorporated by reference herein in its entirety). Thus, there is a major clinical need to accurately detect chronic pain phenotypes so that treatments can be administered appropriately to particular patient subgroups.C. Silicosis

[0107] Silicosis is an emerging public health epidemic whose prevalence is correlated with increased activity and interest in mining and artificial stone industries. Current methods for detecting silicosis, for example, imaging tools, have limited sensitivity where lung abnormalities are difficult to detect or are indistinguishable from other lung diseases. With global prevalence levels and disability-adjusted life years increasing by over 90% and 20%, respectively, in the last 25 years (Liu et al., 2023, BMC Public Health, 1366; incorporated by reference herein in its entirety), detecting silicosis early and accurately can reduce disease burden.D. Endometrial disease

[0108] Endometrial disease can be caused by abnormal thickening of the uterine lining and is one of the most common gynaecological conditions affecting women worldwide (Liu et al., 2023, JNCI Cancer Spectrum, pkadOOl; incorporated by reference herein in its entirety). Given similar clinical presentations, a major challenge is differentiating between benign endometrial diseases (such as endometrial polyps) and endometrial cancer, as well as monitoring high risk cases such as endometrial hyperplasia that may become cancerous. Thus, there is a need to accurately differentiate between different endometrial diseases so that timely interventions can be administered.E. Bacterial vaginosis

[0109] Bacterial vaginosis is the most common vaginal infection affecting more than 1 in 5 reproductive age women worldwide (Peebles et al., 2019, Sex Transm Dis, 304-311; incorporated by reference herein in its entirety). Commonly associated with a vaginal microbiome that is rich in anaerobic bacteria and low in Lactobacillus, bacterial vaginosis can result in adverse health outcomes such as an increased risk of pre-term birth. Thus, early and accurate detection of bacterial vaginosis and differentiation of vaginal microbiome profiles is needed for early intervention and improved health outcomes.WSGR Docket No. 70030-701.601F. Parkinson’s disease

[0110] Parkinson’s disease is the fastest growing neurological disease worldwide with disability-adjusted life years increasing by over 80% in the last 25 years (Dorsey et al., 2018, J Parkinsons Dis, S3-S8; Xu et al., 2024, Lancet Reg Health West Pac, 101078; incorporated by reference herein in its entirety). In the absence of a specific biological test, Parkinson’s disease is clinically diagnosed by the observation of motor symptoms. However, with a misdiagnosis rate of up to 25% (Rizzo et al., 2016, Neurology, 566-576; incorporated by reference herein in its entirety) and Parkinson’s often being diagnosed too late for interventions to be effective, there is an urgent clinical need for accurate and earlier diagnosis to improve disease management and reduce disease burden.G. Synucleinopathies

[0111] By 2050, the cost of managing neurodegenerative diseases like Parkinson’s disease (PD), Alzheimer’s disease (AD), and Dementia with Lewy Bodies (DLB) is expected to exceed hundreds of billions of dollars. Despite this immense impact, there is currently no reliable method to differentiate between these diseases, posing a major challenge to effective patient management and treatment.

[0112] Using methods and systems of the present disclosure, an Al technology is developed and validated to differentiate between synucleinopathies, such as Parkinson’s disease, Alzheimer’s disease, and Lewy Bodies Dementia. Using Al algorithms and mass spectrometry, the technology analyzes tens to thousands of high-resolution biomarker signatures, achieving over 90% accuracy, at the cost of a single biomarker analysis.

[0113] Three of the most commonly misdiagnosed pathologies include AD, PD, and DLB. Diagnosing neurological conditions like PD, AD, and DLB remains challenging because they share the same abnormal protein, alpha-synuclein (a-syn). Current methods can detect a-syn to rule-in / out these diseases but, due to technical and analytical limitations, they cannot accurately distinguish between different synucleinopathies. For example, a skin biopsy test and a cerebrospinal fluid test both use the protein a-syn to detect synucleinopathy. However, this single-protein-biomarker approach is unable to distinguish between different types of synucleinopathy, and its distinct diseases (e.g., AD, PD, DLB), thus limiting effectiveness of patient management and treatment.

[0114] Accurate differential diagnosis may be crucial, as each disease has distinct pathologies that require targeted interventions for optimal clinical outcomes. Without such aWSGR Docket No. 70030-701.601 classification tool, precision treatment may remain unattainable, forcing clinicians / researchers / pharma to rely on lengthy trial-and-error methods, which impair patient management, treatment, and new therapy discovery. As such, the methods and systems of the present disclosure can enable differential classification capabilities that provide significant value to pharmaceutical companies, contract research organizations, and treatment / diagnostic providers. This can prevent patients’ exposure to ineffective medications and harmful side effects, and enable more targeted therapies.Processing and analyzing mass spectrometry data

[0115] The present disclosure provides a streamlined approach that uses mass spectrometry data in its original state prior to the use of a machine learning classifier may improve classification performance and reveal new biomarker signatures for a particular medical condition or disease state that have not been previously reported.

[0116] FIG. 1 illustrates a non-limiting method for diagnosing, detecting, classifying, predicting, stratifying, screening, and / or monitoring medical conditions and / or disease states based on mass spectrometry data that has been analyzed in its original state either independently or in conjunction with non-chemical data.A. Population cohort

[0117] In some embodiments, a subject related to human or animal species is identified, and a population cohort pertaining to a medical condition or disease state is recruited.Examples of medical conditions include, but are not limited to: oncological, immunological, neurological, respiratory, cardiovascular, gastrointestinal, and mental health disease states and / or disorders, viral and bacterial infections, and their respective clinical phenotypes.B. Non-chemical data

[0118] In some embodiments, non-chemical data pertaining to human or animal characteristics may be collected and used. Non-chemical data provide supplemental information that can help contextualize chemical information obtained from biological samples. Non-limiting examples of non-chemical data include, but are not limited to: biological sex, age, ethnicity, height, weight, diet, symptomatic information, family history, medication use, and electronic health records. Non-chemical data can comprise subject data.C. Biological sampleWSGR Docket No. 70030-701.601

[0119] In some embodiments, biological samples pertaining to human or animal species are collected. Examples of biological samples include, but are not limited to: whole blood, blood serum, blood plasma, saliva, breath, urine, feces, sebum, tissue, cerebrospinal fluid, seminal fluid, vaginal secretion, amniotic fluid, nasal fluid, otic fluid, interstitial fluid, breast milk, tears, sputum, sweat, and hair. These samples may be stored at room temperature or in temperatures lower than 10 degrees Celsius prior to preparing the sample for use with a mass spectrometer. As a non-limiting example, a known concentration of a known analyte or a combination of analytes may be combined with a biological sample, and used as an internal or reference standard. In some embodiments, the biological sample is a biofluid.D. Mass spectrometry

[0120] In some embodiments, prepared biological samples are analyzed using a single or a combination of targeted and / or untargeted MS-based methods. Examples of MS-based methods include direct injection MS methods. These include, but is not limited to: ESI-MS, ESI-MS / MS, MALDI-MS, MALDI-MS / MS, EI-MS, EI-MS / MS, CI-MS, CI-MS / MS, APCI- MS, APCI-MS / MS, DART-MS, DART-MS / MS, DESI-MS, DESI-MS / MS, ICP-MS, ICP- MS / MS, FAB-MS, FAB-MS / MS, SIMS, and SIMS / MS.

[0121] In some embodiments, MS-based methods comprise hyphenated MS methods. These may include, but are not limited to: GC-MS, GC-MS / MS, LC-MS, LC-MS / MS, UHPLC-MS, UHPLC-MS / MS, HPLC-MS, HPLC-MS / MS, HILIC-MS, HILIC-MS / MS, CE- MS, and CE-MS / MS.

[0122] In some embodiments, the method comprises detection of chemical species from a mass spectrometer. These analytes may include but are not limited to metabolites, lipids, peptides, or proteins, or fragments thereof. These ions may be a) a positive or negative ion; b) comprise of a mass-to-charge ratio (m / z) value of at least 50 m / z or higher; c) comprise of an ion charge value of 1 or higher. The analyte molecule or fragment thereof may also comprise a m / z value of at least 50 m / z or higher. Overall, an analyte may be considered a biomarker or as part of a biomarker signature in combination with one or more other analytes.

[0123] In some embodiments, an MS-based method comprises Tandem Mass Spectrometry (MS / MS). In some embodiments, an MS-based method comprises a first mass spectrometry stage (MSI) and a second mass spectrometry stage (MS2). In some embodiments, an MS-based method comprises a stage in which ions are generated from a sample. In some embodiments, an MS-based method comprises a stage in which ion fragments are generated. In some embodiments, an MS-based method comprises a first stageWSGR Docket No. 70030-701.601 in which ions are generated from a sample, and a second stage in which ion fragments are generated from the ions. The ions or ion fragments may be analyzed, for example, by collecting ion abundances corresponding to m / z ratios.E. Machine learning classifier

[0124] In some embodiments, original mass spectrometry data is used as input data for analysis by a machine learning classifier with minimal to no prior data reduction steps. Nonlimiting examples of machine learning classifiers include: linear regression, linear discriminant analysis, logistic regression, random forest, support vector machines, extreme gradient boosting, recurrent neural networks, convolutional neural networks, transformers, and other neural network methods. In some embodiments, the machine learning classifier comprises a neural network. A machine learning algorithm provided herein can comprise a machine learning classifier.

[0125] The machine learning classifier can be trained using training data. In some embodiments, the training data is obtained or derived from a subject (e.g., a subject provided herein) having or suspected of having a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.)

[0126] The machine learning classifier can be constructed based at least in part on subject data. Non-limiting examples of subject data include biological sex, age, ethnicity, height, weight, diet, symptomatic information, family history, medication use, genetic history, and electronic health records. In some embodiments, non-chemical data provided herein comprise subject data.

[0127] A machine learning classifier provided herein can be used in a method provided herein to detect, classify, predict, stratify, screen, or monitor a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.) In some embodiments of a method provided herein, a machine learning classifier is used to detect, classify, predict, stratify, screen, or monitor a disease, disorder, or condition based on training data obtained or derived from a subject having or suspected of having a disease, disorder, or condition. The machine learning classifier can be trained on data obtained or derived from a subject having or suspected of having the same or a different disease, disorder, or condition, as the disease, disorder, or condition that is detected, classified, predicted, stratified, screened, or monitored by the machine learning classifier. In some embodiments, the machine learning classifier can be trained on data obtained or derived from a subject having or suspected of having the same disease, disorder, or condition, as the disease, disorder, or condition that is detected,WSGR Docket No. 70030-701.601 classified, predicted, stratified, screened, or monitored by the machine learning classifier. In some embodiments, the machine learning classifier can be trained on data obtained or derived from a subject having or suspected of having a different disease, disorder, or condition, as the disease, disorder, or condition that is detected, classified, predicted, stratified, screened, or monitored by the machine learning classifier.

[0128] In some embodiments of a method provided herein, a machine learning classifier is used to generate a classification performance metric. The classification performance metric can comprise accuracy, sensitivity, specificity, precision (or positive predictive value, PPV), negative predictive value (NPV), Mathew’s correlation coefficient (MCC), area under the receiver operating characteristic curve, and area under the precision recall curve.F. Biomarker and Biomarker signature

[0129] Provided herein in some embodiments is a method comprising a machine learning classifier to identify a biomarker signature. The biomarker signature can comprise one or more biomarker. In some embodiments, a biomarker is correlated with a disease, disorder, or condition. A method provided herein can comprise correlating a biomarker with a disease, disorder, or condition. A biomarker can comprise a metabolite, a nucleic acid, a protein, a peptide, a lipid, or a combination thereof.

[0130] In some embodiments, a biomarker (e.g., a biomarker provided herein comprises a feature), for example a chemical or a non-chemical feature. A non-chemical feature can comprise subject data. Non-limiting examples of subject data are biological sex, age, ethnicity, height, weight, diet, symptomatic information, family history, medication use, genetic history, and electronic health records. In some embodiments, the biomarker comprises an ion feature. The ion feature can be detected in a positive or a negative ionization mode of a mass spectrometry (e.g., a mass spectrometry provided herein.) Ion features can be obtained using untargeted mass spectrometry or targeted mass spectrometry.

[0131] In some embodiments, a biomarker signature is identified following optimization of a machine learning classifier, a biomarker signature provided herein can be optimized using a machine learning classifier. The biomarker signature can be considered as a selection of chemical and / or non-chemical features that can be used to distinguish different pathological indications. A biomarker signature containing chemical features may comprise (i) an ion heatmap, (ii) the absolute abundance of one or more biomarkers, or (iii) the relative ratio of one or more biomarkers. Examples of procedures used to identify a biomarkerWSGR Docket No. 70030-701.601 signature include, but are not limited to: Shapley Additive ExPlanation (SHAP), statistical significance, and feature importance.

[0132] The biomarker signature can be used in a method provided herein to monitor the progression of a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.) In some embodiments, the biomarker signature is used in a method provided herein to monitor the response of a subject to a therapy or intervention, for example, a disease-modifying therapy or intervention. The biomarker signature can be used in a method provided herein to identify the predisposition of a subject to a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.) In some embodiments, the biomarker signature is used in a method provided herein to identify different phenotypes of a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.)

[0133] Provided herein in some embodiments is a method comprising identifying a disease mechanism or a druggable target based at least in part of analyzing original mass spectrometry data.

[0134] A biomarker signature provided herein can comprise an ion heatmap based on mass-to-charge (m / z) values with respect to time. For example, the biomarker signature can comprise an ion heatmap based on m / z values from mass spectrometry data of a first mass spectrometry stage (MSI; mi / zi) with respect to m / z values from mass spectrometry data of a second mass spectrometry stage (MS2; m2 / z2).G. Classification performance

[0135] In some embodiments, the classification performance of a machine learning classifier is determined based on a series of quantitative metrics for binary or multi classification. These metrics can be used for both self-improvement of a given classifier as well as assessing the utility of the classifier in diagnosing, classifying, predicting, stratifying, screening, and / or monitoring known and unknown medical conditions. Metrics may be computed on the training set, validation set, test set, or final predicted data. Examples of classification performance metrics include, but are not limited to: accuracy, sensitivity, specificity, precision (or positive predictive value, PPV), negative predictive value (NPV) Mathew’s correlation coefficient (MCC), area under the receiver operating characteristic curve, and area under the precision recall curve.

[0136] In some embodiments, the area under the received operating characteristic curve (AUROC) is 0.50, 0.55, 0.60, 0.65, 0.70, 0.75, 0.80, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, 0.999, or higher. In some embodiments, theWSGR Docket No. 70030-701.601AUROC is 0.50 or higher. In some embodiments, the AUROC is 0.55 or higher. In some embodiments, the AUROC is 0.60 or higher. In some embodiments, the AUROC is 0.65 or higher. In some embodiments, the AUROC is 0.70 or higher. In some embodiments, the AUROC is 0.75 or higher. In some embodiments, the AUROC is 0.80 or higher. In some embodiments, the AUROC is 0.85 or higher. In some embodiments, the AUROC is 0.86 or higher. In some embodiments, the AUROC is 0.87 or higher. In some embodiments, the AUROC is 0.88 or higher. In some embodiments, the AUROC is 0.89 or higher. In some embodiments, the AUROC is 0.90 or higher. In some embodiments, the AUROC is 0.91 or higher. In some embodiments, the AUROC is 0.92 or higher. In some embodiments, the AUROC is 0.93 or higher. In some embodiments, the AUROC is 0.94 or higher. In some embodiments, the AUROC is 0.95 or higher. In some embodiments, the AUROC is 0.96 or higher. In some embodiments, the AUROC is 0.97 or higher. In some embodiments, the AUROC is 0.98 or higher. In some embodiments, the AUROC is 0.99 or higher. In some embodiments, the AUROC is 0.999 or higher.H. User output

[0137] In some embodiments, a selected machine learning classifier and its associated parameters is applied to data of unknown status whereby a qualitative or quantitative diagnosis, classification, or prediction of a medical condition is then made. In other embodiments, the diagnosis, classification, or prediction made by a machine learning classifier may be accompanied by a probability score. The ability to report a computergenerated prediction and a probability score is an important parameter for clinical decisionmaking and prognosis.Representation of mass spectrometry data in one or two dimensions

[0138] Mass spectrometry data may be visualized as the concentration of molecules as a function of mass-to-charge (m / z) ratio recorded at different time points. With current methods, the high dimensionality of mass spectrometry data may necessitate the need for data reduction methods at the expense of data resolution. Yet, high data resolution is imperative to the performance of machine learning classifiers.

[0139] FIG. 2, FIG. 3, and FIG. 4 illustrate non-limiting methods for representing mass spectrometry data in one or two dimensions (one dimensional: FIG. 2 and FIG. 3; two dimensional: FIG. 4).WSGR Docket No. 70030-701.601A. One-dimensional representation

[0140] In some embodiments, the mass spectrometry data is coupled with an associated label (e.g., 0 for “control” or “negative”, 1 for “disease X” or “positive”, 2 for “disease Y”, 3 for “disease Z”). The concentration of each analyte as detected from a mass spectrometer is then aggregated over the time dimension, resulting in a single summed abundance value for each m / z value. A tabular representation e.g., an array) of the data is then used as input data for analysis by a machine learning classifier. A one-dimensional array can be, for example, a two-dimensional matrix with one-dimensional spatial coordinates.

[0141] In other embodiments, the m / z values may be aggregated based on a specified m / z range (buckets) and the sum total of abundance for a given m / z bucket is then applied.

[0142] In some embodiments, the data is split into multiple subsets representing a training dataset, a validation dataset, and / or a test dataset. In some embodiments, a cross-fold validation technique is used to assess the machine learning classifier and an average prediction of results obtained from a set of validation datasets is obtained for a given model classifier.B. Two-dimensional representation

[0143] In some embodiments, the mass spectrometry data is coupled with an associated label (e.g., 0 for “control” or “negative”, 1 for “disease X” or “positive”, 2 for “disease Y”, 3 for “disease Z”). A two-dimensional matrix can be depicted with two-dimensional spatial coordinates (e.g., a heatmap). For example, a two-dimensional heatmap can be constructed based on an x-axis representing m / z values and a y-axis representing time (t). A color gradient is then applied to the concentrations of each analyte at each (m / z, t) pair.

[0144] In some embodiments, the 2-dimensional matrix represented as a heatmap can also be constructed based on an x-axis representing m / z values and a y-axis representing m / z values whereby m / z values reflect unfragmented and fragmented ion data, respectively, or vice versa. A color gradient is also applied to the concentrations of each analyte at each (m / z, m / z) pair.

[0145] In some embodiments, the data is split into multiple subsets representing a training dataset, a validation dataset, and / or a test dataset. In some embodiments, a cross-fold validation technique is used to train the machine learning classifier and an average of results obtained from a set of validation datasets is obtained for a given model classifier.WSGR Docket No. 70030-701.601

[0146] A two-dimensional representation of the mass spectrometry data broadly resembles digital images where each (m / z, t) or (mi / zi, 1112 / Z2) or (m2 / z2, mi / zi) pair corresponds to brightness values that can be aggregated to form an image. This approach introduces spatial hierarchy which is a data format that may be better suited for particular machine learning algorithms. As a result, the ability to represent mass spectrometry data in a visual format may improve the performance of particular machine learning classifiers.Methods

[0147] Mass spectrometry data may be processed and analyzed with a machine learning classifier, for example, via a two-step process. With current methods, peaks are pre-selected prior to the use of a machine learning classifier to exclude low abundance or background ions that may not contribute to disease classification. However, exclusion of these ions may unintentionally fail to include low abundance ions that may be relevant to disease classification.

[0148] FIG. 5 illustrates a non-limiting method for processing and analyzing mass spectrometry data using a two-pass machine learning algorithm.A. First pass

[0149] In some embodiments, the first pass of the two-pass machine learning algorithm comprises applying the full m / z range from a distinct set of the mass spectrometry data into a machine learning classifier. In some embodiments, the mass spectrometry data used in the first pass can be represented as either a one-dimensional array or a two-dimensional matrix (e.g. visualized as a heatmap).B. Second pass

[0150] In some embodiments, the second pass of the two-pass machine learning algorithm comprises applying one or more targeted m / z ranges (zoom-in range) from a distinct set of the mass spectrometry data into a machine learning classifier. In some embodiments, the mass spectrometry data used in the second pass can be represented as either a one-dimensional array or a two-dimensional matrix (e.g., visualized as a heatmap).C. Multi-pass method

[0151] In an aspect, provided herein is a method of analyzing original mass spectrometry data. In some embodiments, the method comprises performing a first pass e.g., a first passWSGR Docket No. 70030-701.601 provided herein). The first pass can comprise analyzing the original mass spectrometry data. In some embodiments, the method comprises performing a second pass. The second pass can comprise analyzing a targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data.

[0152] Provided herein, in some embodiments, is a method of analyzing original mass spectrometry data, comprising (a) performing a first pass comprising analyzing the original mass spectrometry data, and (b) performing a second pass comprising analyzing a targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data. In some embodiments, performing the second pass analysis of the targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data is based at least in part on an output of the first pass.

[0153] In some embodiments, original mass spectrometry data comprises a full m / z range. The original mass spectrometry data can comprise a one-dimensional array. A onedimensional array can be, for example, a two-dimensional matrix with one-dimensional spatial coordinates. The original mass spectrometry data can comprise a two-dimensional matrix represented with two dimensional spatial coordinates (e.g., a heatmap). The original mass spectrometry data can comprise one or more of data in a .raw, ,wiff2, .D, .LCD, .CDF, mzXML, or mzML format.

[0154] In some embodiments of a method of analyzing original mass spectrometry data provided herein, the method comprises using a machine learning algorithm. For example, a machine learning algorithm can be a machine learning classifier (e.g., a machine learning classifier provided herein.) The machine learning algorithm can comprise one or more of linear regression, linear discriminant analysis, logistic regression, random forest, support vector machines, extreme gradient boosting, recurrent neural networks, convolutional neural networks, and transformers. In some embodiments, the machine learning algorithm comprises a neural network. The machine learning algorithm can be used (e.g., in a method provided herein) to detect, classify, stratify, screen, or monitor a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.)

[0155] In some embodiments, provided herein is a machine learning algorithm (e.g., a machine learning algorithm provided herein) that is trained using training data. The training data can be obtained from a subject (e.g., a subject provided herein). In some embodiments, the subject has or is suspected of having a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.)WSGR Docket No. 70030-701.601

[0156] In some embodiments, a machine learning algorithm (e.g., a machine learning algorithm provided herein) is applied to original mass spectrometry data (e.g., original mass spectrometry data provided herein.) In some embodiments, the first pass comprises applying the machine learning algorithm to the original mass spectrometry data, e.g., a onedimensional array or a two-dimensional heatmap comprising a full m / z range.

[0157] In some embodiments, the second pass comprises analyzing one or more targeted m / z range of original mass spectrometry data, e.g., a one-dimensional array or a two- dimensional matrix (e.g., a heatmap) comprising a full m / z range. The second pass can comprise applying a machine learning algorithm (e.g., a machine learning algorithm provided herein) to one or more targeted m / z range of original mass spectrometry data, e.g., a onedimensional array or a two-dimensional heatmap comprising a full m / z range.

[0158] In certain embodiments, provided herein is a method of analyzing original mass spectrometry data comprising performing a third pass. In some embodiments, the third pass comprises analyzing a second targeted m / z range of the original mass spectrometry data. The second targeted m / z range can be non-overlapping, partially overlapping, or comprised within the first targeted m / z range. The second m / z range can comprise one or more targeted m / z range.

[0159] In certain embodiments, provided herein is a method of analyzing original mass spectrometry data comprising performing a fourth pass. In some embodiments, the fourth pass comprises analyzing a third targeted m / z range of the original mass spectrometry data. The third targeted m / z range can be non-overlapping, partially overlapping, or comprised within the first or the second targeted m / z range. The third m / z range can comprise one or more targeted m / z range.

[0160] In some embodiments, m / z or time values can be grouped (e.g., summed) in bins (or buckets) of different sizes in different passes. For example, a bin size can be larger in the first pass than in a second (or successive) pass; e.g., in the first pass m / z values can be grouped in 0.1 m / z bins and in the second pass m / z values can be grouped in 0.01 m / z bins.

[0161] Provided herein in some embodiments is a method of analyzing original mass spectrometry data obtained using a direct injection mass spectrometry or a hyphenated mass spectrometry. In some embodiments, the original mass spectrometry data is obtained using tandem mass spectrometry (MS / MS), e.g., direct injection tandem mass spectrometry or hyphenated tandem mass spectrometry.

[0162] The direct injection mass spectrometry can comprise one or more of electrospray ionization-mass spectrometry (ESI-MS), matrix-assisted laser desorption ionization-massWSGR Docket No. 70030-701.601 spectrometry (MALDI-MS), electron ionization-mass spectrometry (EI-MS), chemical ionization-mass spectrometry (CI-MS), atmospheric pressure chemical ionization-mass spectrometry (APCI-MS), direct analysis in real time-mass spectrometry (DART -MS), desorption electrospray ionization-mass spectrometry (DESI-MS), inductively coupled plasma-mass spectrometry (ICP-MS), fast atom bombardment-mass spectrometry (FAB-MS), and secondary ion mass spectrometry (SIMS). The direct injection tandem mass spectrometry (MS / MS) can comprise one or more of ESI-MS / MS, MALDI-MS / MS, EI-MS / MS, CI- MS / MS, APCI-MS / MS, DART-MS / MS, DESI-MS / MS, ICP-MS / MS, FAB-MS / MS, and SIMS / MS.

[0163] In some embodiments, the hyphenated mass spectrometry comprises one or more of gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS), high-performance liquid chromatography-mass spectrometry (HPLC-MS), hydrophilic interaction chromatography-mass spectrometry (HILIC-MS), and capillary electrophoresis-mass spectrometry (CE-MS). In some embodiments, the hyphenated mass spectrometry (MS / MS) comprises a tandem mass spectrometry (MS / MS) method, the MS / MS method comprising one or more of GC -MS / MS, LC-MS / MS, UHPLC-MS / MS, HPLC-MS / MS, HILIC -MS / MS, and CE-MS / MS.

[0164] In some embodiments, provided herein is a method comprising using a machine learning algorithm (e.g., a machine learning algorithm provided herein) to generate a performance metric. The performance metric can comprise one or more of accuracy, sensitivity, specificity, precision, positive predictive value (PPV), negative predictive value (NPV), Mathew’s correlation coefficient (MCC), area under the receiver operating characteristic curve (AUROC), and area under the precision recall curve.

[0165] In some embodiments, the area under the received operating characteristic curve (AUROC) is 0.50, 0.55, 0.60, 0.65, 0.70, 0.75, 0.80, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, 0.999, or higher. In some embodiments, the AUROC is 0.50 or higher. In some embodiments, the AUROC is 0.55 or higher. In some embodiments, the AUROC is 0.60 or higher. In some embodiments, the AUROC is 0.65 or higher. In some embodiments, the AUROC is 0.70 or higher. In some embodiments, the AUROC is 0.75 or higher. In some embodiments, the AUROC is 0.80 or higher. In some embodiments, the AUROC is 0.85 or higher. In some embodiments, the AUROC is 0.86 or higher. In some embodiments, the AUROC is 0.87 or higher. In some embodiments, the AUROC is 0.88 or higher. In some embodiments, the AUROC is 0.89 or higher. In someWSGR Docket No. 70030-701.601 embodiments, the AUROC is 0.90 or higher. In some embodiments, the AUROC is 0.91 or higher. In some embodiments, the AUROC is 0.92 or higher. In some embodiments, the AUROC is 0.93 or higher. In some embodiments, the AUROC is 0.94 or higher. In some embodiments, the AUROC is 0.95 or higher. In some embodiments, the AUROC is 0.96 or higher. In some embodiments, the AUROC is 0.97 or higher. In some embodiments, the AUROC is 0.98 or higher. In some embodiments, the AUROC is 0.99 or higher. In some embodiments, the AUROC is 0.999 or higher.

[0166] In some embodiments a method of analyzing original mass spectrometry data provided herein comprises determining a probability score. The probability score can be based at least in part on a detection, classification, or prediction of a machine learning algorithm (e.g., a machine learning algorithm provided herein) relating to a biological sample. In some embodiments, the probability score is based at least in part on a detection, classification, or prediction of the machine learning algorithm of a biological sample.

[0167] The biological sample can be an anonymized sample (e.g., an anonymized sample obtained or derived from a subject of a plurality of subjects.) In some embodiments, the biological sample is a sample of unknown origin. In some embodiments, the biological sample comprises one or more of whole blood, blood serum, blood plasma, saliva, breath, urine, feces, sebum, tissue, cerebrospinal fluid, seminal fluid, vaginal secretion, amniotic fluid, nasal fluid, otic fluid, interstitial fluid, breast milk, tears, sputum, sweat, and hair. In some embodiments, a method provided herein comprises obtaining a first biological sample of a subject. In some embodiments, the method comprises processing the biological sample (e.g., a biological sample provided herein) using mass spectrometry (e.g., a mass spectrometry provided herein), thereby obtain original mass spectrometry data (e.g., original mass spectrometry data provided herein.) In some embodiments, the method comprises analyzing the original spectrometry, thereby obtaining processed spectrometry data. In some embodiments, the method comprises constructing a machine learning classifier (e.g., a machine learning classifier provided herein) based at least in part on the process spectrometry data. In some embodiments, the method comprises using the machine learning classifier to identify a biomarker signature (e.g., a biomarker signature provided herein) comprising one or more biomarker (e.g., one or more biomarker provided herein) that is correlated with a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.) In some embodiments, the method comprises applying the machine learning classifier to a second biological sample (e.g., a biological sample provided herein), thereby correlating the biomarker signature with the disease, disorder, or condition.WSGR Docket No. 70030-701.601

[0168] In an aspect, provided herein is a method comprising: (a) obtaining a first biological sample of a subject; (b) processing the biological sample using mass spectrometry, thereby obtaining original mass spectrometry data; (c) analyzing the original mass spectrometry data, thereby obtaining processed spectrometry data; (d) constructing a machine learning classifier based at least in part on the processed spectrometry data; (e) using the machine learning classifier to identify a biomarker signature comprising one or more biomarker that is correlated with a disease, disorder, or condition; and (f) applying the machine learning classifier to a second biological sample, thereby correlating the biomarker signature with the disease, disorder, or condition. In some embodiments, analyzing the original mass spectrometry data comprises: (a) performing a first pass (e.g., a first pass provided herein) comprising analyzing the original mass spectrometry data, and (b) performing a second pass (e.g., a second pass provided herein) comprising analyzing a targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data.

[0169] In some embodiments, provided herein is a method comprising determining a probability score. The probability score can be based at least in part on a detection, classification, or prediction of a machine learning classifier (e.g., a machine learning classifier provided herein) of a biological sample. In some embodiments, the probability score is based at least in part on a detection, classification, or prediction of a machine learning classifier (e.g., a machine learning classifier provided herein) of a first biological sample. In some embodiments, the probability score is based at least in part on a detection, classification, or prediction of a machine learning classifier (e.g., a machine learning classifier provided herein) of a second biological sample.

[0170] In some embodiments, the first or the second biological sample is an anonymized sample (e.g., a sample of unknown origin.) In some embodiments, the first biological sample is an anonymized sample (e.g., an anonymized sample obtained or derived from a subject of a plurality of subjects.) In some embodiments, the second biological sample is an anonymized sample (e.g., a sample obtained or derived from a subject of a plurality of subjects.) In some embodiments, the second biological sample is obtained or derived from the first or the second subject. In some embodiments, the second biological sample is obtained or derived from the first subject. In some embodiments, the second biological sample is obtained or derived from the second subject. The second biological sample can be obtained or derived from a subject after the first biological sample was obtained or derived from the subject (e.g., to track progression of a disease, disorder, or condition or the effect of an intervention.) The first or the second biological sample can be a sample provided herein.WSGR Docket No. 70030-701.601

[0171] In an aspect, provided herein is a method of processing original mass spectrometry data (e.g., original mass spectrometry data provided herein.) In some embodiments, a method of processing original mass spectrometry data provided herein comprises using a machine learning classifier (e.g., a machine learning classifier provided herein.)

[0172] In some embodiments, provided herein is a method of processing original mass spectrometry data, the original mass spectrometry data comprising a one-dimensional array, comprising using a machine learning classifier. In some embodiments, when the original mass spectrometry data is a one-dimensional array, an ion abundance of the original mass spectrometry data for an m / z value is summed over time, resulting in a single summed abundance value for the m / z value. In some embodiments, when the original mass spectrometry data is a one-dimensional array, an ion abundance of the original mass spectrometry data for an m / z value is summed over an m / z range, providing a single summed abundance value for the m / z range. A summed abundance value, e.g., for an m / z value or an m / z range, can be associated with an output indicative of the presence of a disease, disorder, or condition (e.g., a disease, disorder, or condition provided herein.)

[0173] In some embodiments, provided herein is a method of processing original mass spectrometry data, the original mass spectrometry data comprising a two-dimensional matrix represented as a two-dimensional heatmap, comprising using a machine learning classifier. In some embodiments, when the original mass spectrometry data comprises a two-dimensional matrix represented as a two-dimensional heatmap, a dimension or an axis (e.g., an x-axis) of the data represents m / z and a dimension or an axis (e.g., a y-axis) of the data represents time (t), and the value of each (m / z, t) pair is an ion abundance value corresponding to the (m / z, t) pair. In some embodiments, when the original mass spectrometry data comprises a two- dimensional matrix represented as a two-dimensional heatmap, a first dimension or axis (e.g., an x-axis) of the data represents first stage mass spectrometry data (MSI; mi / zi) and a second dimension or axis (e.g., a y-axis) of the data represents second stage mass spectrometry data (MS2; m2 / z2), and the value of each (mi / zi, m2 / z2) pair is an ion abundance value corresponding to the (mi / zi, 1112 / Z2) pair. In some embodiments, when the original mass spectrometry data comprises a two-dimensional matrix represented as a two-dimensional heatmap, a first dimension or axis (e.g., an x-axis) of the data represents second stage mass spectrometry data (MS2; m2 / z2) and a second dimension or axis (e.g., a y-axis) of the data represents first stage mass spectrometry data (MSI; mi / zi), and the value of each (m2 / z2, mi / zi) pair is an ion abundance value corresponding to the (m2 / z2, mi / zi) pair.WSGR Docket No. 70030-701.601

[0174] In some embodiments, when the original mass spectrometry data comprises a two- dimensional matrix represented as a two-dimensional heatmap, an ion abundance of an m / z value of the original mass spectrometry data can be summed over an m / z range, providing a summed abundance value for the m / z range. In some embodiments, when the original mass spectrometry data comprises a two-dimensional matrix represented as a two-dimensional heatmap, an ion abundance of an m / z range of the original mass spectrometry data is summed over a second m / z range, providing a summed abundance value for the second m / z range with respect to time.

[0175] In certain embodiments, when the original mass spectrometry data comprises a two-dimensional matrix represented as a two-dimensional heatmap, an ion abundance of an m / z range of the original mass spectrometry data is summed over a time range, providing an abundance value for the time range with respect to the m / z value. In some embodiments, the ion abundance of the m / z value is further summed over an m / z range, providing a summed abundance value for the m / z range with respect to the time range. In some specific embodiments, the ion abundance of the m / z range is further summed over a second m / z range, providing a summed abundance value for the second m / z range with respect to the time range. In some specific embodiments, the ion abundance of the time range is further summed over a second time range, providing a summed abundance value for the second time range with respect to the m / z range.Identification of non-linear biological effects

[0176] Combinatorial analysis can be conducted on lead biomarker features to identify the presence of non-linear biological or network effects. In some embodiments, the lead biomarker features are identified by methods including but not limited to SHAP. In other embodiments, a combinatorial analysis can be conducted on the lead features whereby interaction terms between all combinations can be represented mathematically (e.g., a / (a + b), a * b, a / b, where a and b are the abundance values of features A and B, respectively).WSGR Docket No. 70030-701.601EXAMPLESExample 1 - Machine learning algorithm for multi-class classification of viral infections

[0177] The efficacy of systems and methods disclosed herein in the classification of acute febrile diseases was analyzed.Method:

[0178] Briefly, a Thermo Orbitrap QExactive (Thermo Fisher Scientific) mass spectrometer coupled with a Dionex UltiMate 3000 RSLC system (Thermo Fisher Scientific, Hemel Hempstead, UK) with a ZIC-pHILIC column (150 mm x 4.6 mm, 5 pm column, Merck SeQuant) was used (Nastase et al., 2023, PLoS Negl Trop Dis, eOOl 1133; incorporated by reference herein in its entirety). Untargeted LC-MS metabolomic data from serum was collected in both positive and negative ion mode. Mass spectrometry data belonging to 37 healthy controls, and 7 malaria, 18 visceral Leishmaniasis, and 10 Zika virus samples were analyzed.

[0179] A multi-classification machine learning model was developed using the LC-MS data in its original form with no prior peak identification steps. A bucket of 0.5 m / z was applied to create a structured array. Two different machine learning classifiers were developed with one representing mass spectrometry data as a one-dimensional array and another as a two-dimensional heatmap. Classification performance was assessed for each classifier based on 5-fold cross-validation using a) combination of positive and negative ions, b) positive ions only, and c) negative ions only.Results and Discussion:

[0180] By visualizing mass spectrometry data as one-dimensional, the application of an extreme gradient boosting classifier resulted in overall average classification accuracies of 83.4%, 83.3%, and 78.3% for positive and negative ions, positive ions only, and negative ions only, respectively (Table 1). For different ionization modes, the accuracy in classifying healthy controls, malaria, visceral Leishmaniasis, and Zika virus ranged from 75.1% to 97.2%.

[0181] By visualizing mass spectrometry data as two-dimensional, the application of a convolutional neural network classifier resulted in overall average classification accuracies of 86.2%, 84.7%, and 86.1% for positive and negative ions, positive ions only, and negativeWSGR Docket No. 70030-701.601 ions only, respectively (Table 1). For different ionization modes, the accuracy in classifying healthy controls, malaria, visceral Leishmaniasis, and Zika virus ranged from 71.4% to 100%.Table 1. Summary of classification accuracy for each classification and ionization mode using one-dimensional and two-dimensional data

[0182] In comparison, in the method as described by Nastase et a incorporated by reference herein in its entirety, biomarkers were pre-selected based on whether the concentration of biomarkers found in disease versus healthy control had a p-value < 0.05. However, for the majority of these biomarkers, the analyte levels of each single biomarker often overlapped with the analyte levels of the same biomarker for healthy controls for each of the infections.

[0183] In contrast, our results demonstrated that when mass spectrometry data was analyzed with no data reduction, the resulting biomarker signatures accurately differentiatedWSGR Docket No. 70030-701.601 healthy controls from each infection. An example of a confusion matrix obtained from a multi-classification model of 1 -dimensional or 2-dimensional mass spectrometry data using only positive ions are shown in FIG. 6 and FIG. 7, respectively. Overall, these results demonstrated that by using the systems and methods disclosed herein, different infectious causes of acute febrile illnesses were classified with up to 100% accuracy.Example 2 - Machine learning algorithm for identifying metabolic biomarker signatures to differentiate chronic pain phenotypes

[0184] The efficacy of systems and methods disclosed herein in the classification of chronic pain phenotypes was analyzed.Method:

[0185] Briefly, a Bruker timsTOF Pro 2 (Bruker) mass spectrometer coupled with a proprietary LC column was used. Untargeted LC-MS metabolomic data from urine was collected in positive ion mode. Mass spectrometry data belonging to 98 chronic pain patients with a pain score greater than 75 and 102 chronic pain patients with a pain score of 0 were analyzed.

[0186] A binary classification machine learning model was developed using the LC-MS data in its original form with no prior peak identification steps. To reduce data dimensionality and to ensure consistency in the number of data scans between different samples, m / z bucketing was applied. A two-pass machine learning algorithm using mass spectrometry data represented as a one-dimensional array or two-dimensional heatmap was used to further boost classification performance. Classification performance was calculated based on 10-fold cross validation.Results and Discussion:

[0187] In the first pass of the two-pass machine learning algorithm, mass spectrometry data was visualized as one-dimensional. Application of an extreme gradient boosting classifier with 0.2 m / z buckets resulted in an average accuracy of 87.5% and an Area Under the Receiver Operating Characteristic (AUROC) of 0.936 (FIG. 8). Detailed classification performance metrics are provided in Table 2.

[0188] Shapley Additive ExPlanation (SHAP) can reveal m / z features that contribute most to model predictions. SHAP analysis was used to interpret model results from the firstWSGR Docket No. 70030-701.601 pass. Using SHAP, a distribution plot showed that the vast majority of important m / z features to differentiate different chronic pain phenotypes were localized in the 105-165 m / z range (with full m / z range being 45-1705) (FIG. 9).

[0189] In the second pass of the two-pass machine learning algorithm, mass spectrometry data was visualized as two-dimensional. The mass spectrometry data was ‘zoomed-in’ on the 105-165 m / z range and further refined with 0.01 m / z bucket - a 20 times increase in resolution compared to the first pass. Using a convolutional neural network classifier, this ‘zoomed-in’ method resulted in accuracy of 96.4% and AUROC of 0.982 (FIG. 10). In comparison, applying the same classifier on the full 45-1705 m / z range with 0.2 bucket resulted in accuracy of 91.5% and AUROC of 0.937 (Table 2). These results demonstrated that classification performance can significantly increase when the second pass of the machine learning algorithm was used for more localized and higher-resolution analysis.Table 2. Summary of classification performance for the two-pass machine learning algorithm.WSGR Docket No. 70030-701.601

[0190] SHAP analysis was then applied to the convolutional neural network classifier to identify biomarker signatures resulting in high classification accuracy. The resulting SHAP values were visualized as a two-dimensional ion heatmap, where each pixel represents the contribution of a specific (m / z, time) pair to the classification outcome. This visualization method resulted in a biomarker signature comprising spatial patterns with respect to m / z and time. Given that each individual sample produced unique biomarker signatures, identifying biomarker regions (or clusters) from these ion heatmaps may be further used for stratification and biomarker discovery (FIG. 11).

[0191] Results obtained from the method disclosed were compared with a current approach (i.e., CRANK-MS, Zhang et al., 2023, ACS Cent. Sci, 9, 1035-1045; incorporated by reference herein in its entirety). Using current methods of biomarker discovery, a machine learning classifier was applied to a pre-specified list of 24 m / z biomarkers that had a log2 fold change greater than 2 or a p-value less than IO'20. Using a linear support vector machine learning classifier, the accuracy was 86.6% and AUROC of 0.946 (Table 2). These results demonstrated that classification performance can be significantly improved by using original mass spectrometry data with no prior peak identification steps.

[0192] In addition, the top biomarker hits identified from the current approach differed from the hits identified from the method described herein. For example, in comparing the top 10 m / z features obtained from the current method based on p-value with the top 10 features obtained from the method described herein, 50% of the top features were found in common between the two approaches. The remaining 50% of features were different for each method. This difference can be explained by the fact that pre-selected biomarkers may be based on the assumption that biomarkers behave linearly and independently of other biomarkers. However, metabolites may be interdependent and the relationship between different biomarkers may be non-linear. Given that classification accuracy can significantly improve by analyzing whole mass spectrometry data compared to pre-selected biomarkers, this result indicates that this method can identify more accurate biomarkers through considering non-linear effects.

[0193] To demonstrate that this method can identify non-linear or network effects, combinatorial SHAP analysis was applied to the top 20 m / z features identified from the first pass. Here, feature interactions were represented as A / (A + B) and calculated as, a / (a + b), where a and b are the abundance values of the features A and B, respectively. This resulted in top m / z features belonging to combinations of features as opposed to individual m / z features (FIG. 12). These results indicated that the method described can identify new biomarkers that were influenced by non-linear effects and that such results were an improvement over resultsWSGR Docket No. 70030-701.601 obtained using current methods. Overall, these results demonstrated that by using the systems and methods disclosed herein, new biomarkers can be identified which can lead to significant improvements in classification performance when compared to current methods.Example 3 - Machine learning algorithm for binary classification of silicosis

[0194] The efficacy of systems and methods disclosed herein in the classification of silicosis was analyzed.Method:

[0195] Briefly, a custom-made APCI source connected to a LTQ-XL mass spectrometer (Thermo Fisher Scientific) was used (Baker et al., 2025, J Breath Res, 026011; incorporated by reference herein in its entirety). Untargeted APCI-MS data from breath was collected in positive ion mode. Mass spectrometry data belonging to 31 subjects diagnosed with having silicosis and 60 healthy control samples were analyzed.

[0196] A binary classification machine learning model was developed using the APCI- MS data in its original form with no prior peak identification steps. A bucket of 0.1 m / z was applied to create a structured array. A machine learning classifier was developed with mass spectrometry data represented as a one-dimensional array. Classification performance was assessed based on 10-fold cross-validation.Results and Discussion:

[0197] By visualizing mass spectrometry data as one-dimensional, the application of an extreme gradient boosting classifier resulted in an average accuracy of 88.1% and an Area Under the Receiving Operating Characteristic (AUROC) of 0.925 (FIG. 13). Detailed classification performance metrics are provided in Table 3.Table 3. Summary of classification performance for a binary machine learning classifier.WSGR Docket No. 70030-701.601

[0198] Results obtained from the method disclosed were compared with the method, CRANK -MS (Baker et al., 2025, J Breath Res, 026011; incorporated by reference herein in its entirety). Using an extreme gradient boosting classifier, classification performance from the 1 -dimensional model was significantly higher across all metrics compared to the CRANK-MS method that used either all 550 features obtained from a peak identification software or 26 top features as identified by Shapley Additive ExPlanation. For example, sensitivity increased by 26.8% and 17.5% when the original mass spectrometry data was used compared to the CRANK -MS method where all 550 and 26 features were pre-selected, respectively.

[0199] Crucially, these results demonstrated that despite both methods using ‘whole’ mass spectrometry data, there was improved performance from utilizing original mass spectrometry data versus all features as obtained from peak identification software. That is, these results indicated that the use of original mass spectrometry data with no prior peak identification steps can preserve data quality and reduce information loss, and can significantly improve classification performance compared to current methods.WSGR Docket No. 70030-701.601Example 4 - Machine learning algorithm for identifying metabolic biomarker signatures to differentiate different endometrial diseases

[0200] The efficacy of systems and methods disclosed herein in the classification of different endometrial diseases was analyzed.Method:

[0201] Briefly, a quadrupole time-of-flight mass spectrometer (Waters, USA) coupled with a Waters ACQUITY UPLC system (Waters Company) with a ACQUITY UPLC BEH C18 column (2.1 x 100 mm, 1.7 pm; Waters Company, Milford, USA) was used (Yan et al., 2022, Int J Cancer, 1549-1559; incorporated by reference herein in its entirety). Untargeted LC-MS lipidomic data from serum was collected in positive ion mode. Mass spectrometry data belonging to 201 endometrial polyps, 73 endometrial cancer, 52 endometrial hyperplasia, and 225 health control samples were analyzed.

[0202] A binary and multi-class classification machine learning model was developed using the LC-MS data in its original form with no prior peak identification steps. To reduce data dimensionality and to ensure consistency in the number of data scans between different samples, m / z bucketing was applied. Two different machine learning approaches were used, namely one that represented mass spectrometry data as a one-dimensional array and another as a two-dimensional heatmap. Classification performance for binary and multi-class classification using each machine learning approach was evaluated based on 5-fold cross- validation.Results and Discussion:

[0203] By visualizing mass spectrometry data as one-dimensional, the application of a logistic regression classifier resulted in average accuracy of 98.7% and 93.3% for binary classification of endometrial polyps versus cancer and endometrial polyps versus hyperplasia, respectively. Similarly, by visualizing mass spectrometry data as two-dimensional, the application of a convolutional neural network classifier resulted in average accuracy of 99.6% and 91.6% for binary classification of endometrial polyps versus cancer and endometrial polyps versus hyperplasia, respectively. Detailed classification performance metrics are provided in Table 4, and example Area Under the Receiver Operating Characteristic (AUROC) curves for binary classification of endometrial polyp versus cancerWSGR Docket No. 70030-701.601 using the one-dimensional and two-dimensional approach are shown in FIG. 14 and FIG. 15, respectively.Table 4. Summary of binary classification accuracy using one-dimensional and two- dimensional data.

[0204] Results obtained from the method disclosed were compared with the method described by Yan et al., 2022, Int J Cancer, 1549-1559; incorporated by reference herein in its entirety. Classification performance obtained from either the 1 -dimensional or 2- dimensional method was significantly higher across all metrics when compared to the method of Yan et al. In particular, classification performance obtained from the 1 -dimensional method which utilized all lipid features from the original mass spectrometry data was significantly higher compared to a four lipid biomarker panel that were pre-selected, despite a logistic regression classifier being used in both methods. For example, the sensitivity score in classifying endometrial polyps versus cancer increased by 28.3% and endometrial polyps versus hyperplasia increased by 30.8% when the 1 -dimensional method was used comparedWSGR Docket No. 70030-701.601 to the method of Yan et al. These results demonstrated that the use of mass spectrometry data with no prior peak identification steps can significantly improve classification performance.

[0205] Further, by using either the one-dimensional or two-dimensional method, multiclass classification of different endometrial diseases was obtained with comparable classification performance to binary classification. For example, the average multi-class classification accuracy of healthy controls, endometrial polyps, cancer, and hyperplasia using the one-dimensional and two-dimensional method was 91.3% and 94.1%, respectively. A confusion matrix obtained from the multi-class classification model using 1 -dimensional and 2-dimensional mass spectrometry data is shown in FIG. 16 and FIG. 17, respectively. Overall, these results demonstrated that by using the systems and methods disclosed herein, different endometrial diseases can be classified with accuracy as high as 99%.Example 5 - Machine learning algorithm for identifying metabolic biomarker signatures to differentiate different vaginal microbial profiles

[0206] The efficacy of systems and methods disclosed herein in the classification of different vaginal microbial profiles was analyzed.Method:

[0207] Briefly, a LTQ-Orbitrap Discovery mass spectrometer (Thermo Scientific, Bremen, Germany) coupled with a DESI-MS source was used (Pruski et al., 2021, Nat Commun, 5967; incorporated by reference herein in its entirety). Untargeted DESI-MS metabolomic data from cervicovaginal fluid was collected in positive and negative ion mode. Mass spectrometry data from 1,028 cervicovaginal swab samples belonging to 365 pregnant women across two independent cohorts (VMET and VMET2) were analyzed. Specifically, 404 ZrzctoZ>rzcz7 / z-dominant and 51 ZrzctoZ>rzcz7 / z-depleted samples belonging to 160 women from VMET, and 451 ZrzctoZ>rzcz7 / z-dominant and 1227 zctoZ>rzcz7 / z-depleted samples belonging to 205 women from VMET2 were used.

[0208] A binary and multi-class classification machine learning model was developed using the DESI-MS data in its original form with no prior peak identification steps. A bucket of 0.1 m / z was applied to create a structured array. A machine learning classifier was developed with mass spectrometry data represented as a two-dimensional array. Classification performance was assessed based on 7-fold cross-validation.WSGR Docket No. 70030-701.601Results and Discussion:

[0209] By visualizing mass spectrometry data as two-dimensional, the application of a convolutional neural network classifier resulted in average accuracies of 99.6%, 99.3%, 99.1%, and 99.1% for binary classification of ZrzctoZ>rzcz7 / z-dominant versus Lactobacilli- depleted vaginal microbial profiles from VMET (negative ion mode), VMET2 (negative ion mode), VMET (positive ion mode), and VMET2 (positive ion mode) cohorts, respectively. For different ionization modes and patient cohorts, the Area Under the Receiver Operating Characteristic (AUROC) ranged from 0.981 to 0.999. An example of an AUROC curve for VMET2 (negative ion mode) is depicted in FIG. 18. Detailed classification performance metrics are provided in Table 5.Table 5. Summary of classification performance for a binary machine learning classifier.WSGR Docket No. 70030-701.601

[0210] Results obtained from the method disclosed were compared with the current literature method (z.e., Pruski et al., 2021, Nat Commun, 5967; incorporated by reference herein in its entirety). Classification performance obtained from the 2-dimensional method was significantly higher across all metrics when compared to the method of Pruski et al., regardless of the ionization mode or patient cohort. In particular, in contrast to the Pruski etWSGR Docket No. 70030-701.601 al. method where high specificity was obtained at the expense of sensitivity, results obtained using the 2-dimensional method resulted in both high sensitivity and specificity scores. For example, when classifying ZrzctoZ>rzcz7 / z-dominant versus ZrzctoZ>rzcz7 / z-depleted vaginal microbial profiles from VMET2 (positive ion mode), sensitivity increased from 44.1% to 97.5% and specificity increased from 97.0% to 99.6% when the 2-dimensional method was used compared to the method of Pruski et al. These results demonstrated that the use of original mass spectrometry data can improve overall classification performance by identifying spatially localized patterns and correlated features. This is in contrast to the Pruski et al. method that utilizes a data reductionist approach where 113 biomarkers were preselected.

[0211] Further, by using the two-dimensional method, multi-class classification of different vaginal microbial profiles was obtained. Specifically, three Lactobacilli related community state types (CST) were studied, namely Lactobacilli cris[)atiis- om na.iQ (CST I), Lactobacilli zz / cz'.s-dominated (CST III) and mixed anaerobic communities (CST IV). The average multi-class classification accuracy in differentiating between the three vaginal microbial profiles ranged from 74.3% to 91.8%. An example of a confusion matrix obtained from VMET (negative ion mode) is shown in FIG. 19. Detailed classification performance metrics are provided in Table 6. Overall, these results demonstrated that by using the systems and methods disclosed herein, different vaginal microbial profiles can be classified with accuracy as high as 99%.Table 6. Summary of multi-class classification accuracy for each patient cohort and ionization modeWSGR Docket No. 70030-701.601Example 6 - Machine learning algorithm for binary classification of Parkinson’s disease

[0212] The efficacy of systems and methods disclosed herein in the classification of Parkinson’s disease was analyzed.Method:WSGR Docket No. 70030-701.601

[0213] Briefly, a Thermo Q Exactive HF-X Hybrid Quadrupole-Orbitrap mass spectrometer coupled with a Vanquish Horizon UHPLC system (Thermo Fisher Scientific, Massachusetts, USA) and either an ACQUITY Premier BEH Amide (HILIC) column with VanGuard FIT Cartridge (1.7 pm x 2.1 mm x 100 mm, Waters Corporation, USA) or ACQUITY Premier CSH C18 (C18) column with VanGuard FIT Cartridge (1.7 pm x 2.1 mm x 100 mm, Waters Corporation, USA) was used. Untargeted LC-MS metabolomic data from plasma were collected in positive ion mode. Mass spectrometry data belonging to 51 healthy controls, and 40 Parkinson's disease samples were analyzed.

[0214] A binary classification machine learning model was developed using the LC-MS data in its original form with no prior peak identification steps. To reduce data dimensionality and to ensure consistency in the number of data scans between different samples, m / z bucketing was applied. A two-pass machine learning algorithm using mass spectrometry data represented as a one-dimensional array or two-dimensional heatmap was used to further boost classification performance. Classification performance was evaluated based on 10-fold cross- validation.Results and Discussion:

[0215] In the first pass of the two-pass machine learning algorithm, mass spectrometry data was visualized as one-dimensional. Application of an extreme gradient boosting classifier with 0.1 m / z buckets resulted in an average accuracy of 92.2% and an Area Under the Receiver Operating Characteristic (AUROC) of 0.990 for the HILIC method (FIG. 20). Comparable results were obtained for the Cl 8 method. Detailed classification performance metrics are provided in Table 7.

[0216] Shapley Additive ExPlanation (SHAP) analysis was used to interpret model results from the first pass. Using SHAP, a distribution plot showed that the vast majority of important m / z features to classify Parkinson’s disease versus healthy control were below 500 m / z for both HILIC (FIG. 21) and Cl 8 method (FIG. 22), respectively (with full m / z range being 100 - 1000).

[0217] In the second pass of the two-pass machine learning algorithm, mass spectrometry data was visualized as two-dimensional. The mass spectrometry data was ‘zoomed-in’ on the 150 - 400 m / z range and further refined with 0.05 m / z bucket - a 2 times increase in resolution compared to the first pass. Using a convolutional neural network classifier, this ‘zoomed-in’ method resulted in accuracy of up to 100.0% and AUROC of 1.0 (FIG. 23). In comparison, applying the same classifier on the full 100 - 1000 m / z range with 0.1 bucketWSGR Docket No. 70030-701.601 resulted in accuracy of 91.2% and AUROC of 0.916 (Table 7). Comparable results were obtained for the C18 method. These results demonstrated that classification performance significantly increased when the second pass of the machine learning algorithm was used for more localized and higher-resolution analysis.Table 7. Summary of classification performance for a binary machine learning classifierWSGR Docket No. 70030-701.601

[0218] Results obtained from the method disclosed were compared to the approach for classification of Parkinson’s disease using metabolomic data described in Zhang et al., 2023, ACS Cent. Sci, 9, 1035-1045; incorporated by reference herein in its entirety. Using CRANK-MS on all metabolite features that were pre-identified in LC-MS positive ion mode, Zhang et al. provided the highest accuracy, sensitivity, specificity, precision, and AUROC values of 94.6%, 95.1%, 94.4%, 94.5% and 0.983 using a neural network classifier, respectively. In contrast, using original mass spectrometry data and the two-pass machine learning algorithm, classification performance increased to 100% across all metrics. Overall, these results demonstrated that by using the systems and methods disclosed herein, Parkinson’s disease can be classified with accuracy as high as 100%.Example 7 - Use of fragmentation-based biomarker signatures for disease assessment using machine learning algorithms

[0219] Using methods and systems of the present disclosure, a machine learning-based technology for disease assessment is developed and validated. The machine learning-based technology uses fragmentation-based biomarker signatures corresponding to one or more medical condition over time, based on the chemical analysis of a biological sample (e.g., a biofluid) including blood, urine, saliva, breath, sebum, tissue, or fecal samples. Using Al algorithms (e.g., a convolutional neural network or a deep neural network) and mass spectrometry methods, this approach utilizes mass spectrometry data from precursor ions obtained from first stage mass spectrometry data (MSI) and fragment ions obtained from second stage mass spectrometry data (MS2). This approach results in fragmentation-basedWSGR Docket No. 70030-701.601 biomarker signatures that can be used to detect, classify, predict, stratify, screen, and / or monitor one or more medical conditions. In some cases, the fragmentation-based biomarker signature can be represented as a two-dimensional heatmap. A non-limiting example of a fragmentation-based biomarker signature is shown in FIG. 24.

[0220] The technology comprises use of fragmentation-based biomarker signatures wherein fragment ions are obtained using data dependent acquisition (DDA) or data independent acquisition (DIA) mode from the mass spectrometer. In some cases, the use of fragmentation-based biomarker signatures obtained from DIA mode results in comparable or higher diagnostic performance than biomarker signatures obtained from DDA modeExample 8 - Multimodal machine learning algorithm for disease assessment

[0221] Using methods and systems of the present disclosure, a multimodal machine learning-based technology for disease assessment is developed and validated. The multimodal machine learning-based technology uses biomarker signatures obtained from non-chemical data in addition to chemical data (e.g., mass spectrometry). These biomarker signatures correspond to one or more medical condition over time, based on the chemical analysis of biological samples (e.g., biofluids) including blood, urine, saliva, breath, sebum, tissue, or fecal samples. Using Al algorithms (e.g., a convolutional neural network or a deep neural network) and mass spectrometry methods, this approach utilizes various non-chemical data modalities including demographic information, clinical information, historical clinical data, third-party health records, electronic health records, self-reported symptoms, selfreported journal entries, food diary entries, imaging data, and data from wearable technologies. When using the machine learning algorithm to detect, classify, predict, stratify, screen, and / or monitor one or more medical conditions, this approach results in comparable or improved diagnostic performance compared to machine learning algorithms utilizing mass spectrometry data alone.Example 9 - Time-based monitoring of biomarker signatures using machine learning algorithms

[0222] Using methods and systems of the present disclosure, a machine learning-based technology for monitoring biomarker signatures corresponding to one or more medical condition over time is developed and validated. The machine learning-based technology for monitoring biomarker signatures is based on the chemical analysis of biological samples (e.g., biofluids) including blood, urine, saliva, breath, sebum, tissue, or fecal samples. UsingWSGR Docket No. 70030-701.601Al algorithms (e.g., a convolutional neural network or a deep neural network) and mass spectrometry methods, this approach can detect time-dependent changes to biomarker signatures as a result of biological changes, for example changes resulting from natural or disease-driven processes, disease progression, or response to drug and non-drug interventions. These time-dependent changes may be reflected in the original mass spectrometry data, such as the addition of new m / z features, or changes in the ion abundance or relative ratio of specific m / z features, biomarkers or biochemical pathways. In some cases, the biomarker signature can be represented as a two-dimensional heatmap. A non-limiting example of a time-dependent biomarker signature obtained before and after the use of a drugbased intervention is shown in FIG. 25. FIG. 25A provides the biomarker signature before intervention, and FIG. 25B provides the biomarker signature after intervention.

[0223] The technology comprises use of machine learning algorithms to analyze mass spectrometry data for monitoring transition from healthy to disease state, early onset of disease, disease progression, and response to interventions. This technology can demonstrate capability for real-time health monitoring including use in analyzing pharmacodynamics and pharmacokinetics of drug-based interventions.

[0224] The platform technology can be integrated into any laboratory setting where a mass spectrometer is used e.g., industrial or academic laboratories, contract / clinical research organizations, test providers) or portable mass spectrometers. For example, a computer- interfaced software may be coupled with an experimental protocol for mass spectrometry analysis.Example 10 - Machine learning algorithm for identifying metabolic biomarker signatures to differentiate synucleinopathies

[0225] Using methods and systems of the present disclosure, a machine learning-based technology is developed and validated to differentiate between synucleinopathies, such as Parkinson’s disease, Alzheimer’s disease, and Lewy Bodies Dementia. Using machine learning algorithms and mass spectrometry, the technology analyzes tens to thousands of high-resolution biomarkers (e.g., metabolites or lipids), achieving over 90% accuracy, at the cost of a single biomarker analysis.

[0226] Using methods and systems of the present disclosure, a machine learning-based technology will be developed that has the potential to stratify synucleinopathies with clinically viable accuracy based on the chemical analysis of tissue samples. Machine learningWSGR Docket No. 70030-701.601 algorithms and mass spectrometry methods are leveraged to reveal unique biomarker signatures of each disease. This approach may achieve accuracies over 90%.

[0227] Using skin biopsy testing for synucleinopathy, methods and systems of the present disclosure are validated for clinical use in differentiating a-syn pathologies. The technology is able to generate signatures containing tens to thousands of biomarkers simultaneously, at the same cost as a single biomarker, whilst simultaneously producing comparable or higher classification accuracy than current approaches.

[0228] The technology comprises use of machine learning algorithms to analyze mass spectrometry data for disease diagnosis and stratification. This technology can outperform proof-of-concept machine learning models in classification accuracy, and demonstrate capability for differential diagnosis.

[0229] The platform technology can be seamlessly integrated into any laboratory setting where a mass spectrometer is used (e.g., pharma, CROs, test providers). For example, a computer-interfaced software may be coupled with an experimental protocol for mass spectrometry analysis.

[0230] Clinical studies are performed including model training for clinical validation on skin tissue. The machine learning analysis is validated on a new modality, using skin tissue samples, to establish platform classification capabilities.

[0231] Clinical studies are performed including clinical validation of biological stratification. Stratification capabilities are validated between Parkinson’s disease, Alzheimer’s disease, and Lewy Bodies Dementia, and also between various phenotypes of Parkinson’s disease. This provides evidence of not only differential classification capabilities, but also of platform classification capabilities.

[0232] Using methods and systems of the present disclosure, various AD, PD, and DLB biomarker signatures are obtained from analysis of skin tissue samples.

[0233] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are notWSGR Docket No. 70030-701.601 limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

[0234] All publications, patents and patent applications are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety. Where a term in the present description is found to be defined differently in a document incorporated herein by reference, the definition provided herein is to serve as the definition for the term.

Claims

WSGR Docket No. 70030-701.601CLAIMSWhat is claimed is:

1. A method of analyzing original mass spectrometry data, comprising a. performing a first pass comprising analyzing the original mass spectrometry data, and b. performing a second pass comprising analyzing a targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data.

2. The method of claim 1, wherein performing the second pass comprising analyzing the targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data is based at least in part on an output of the first pass.

3. The method of claim 1 or claim 2, further comprising using a machine learning algorithm.

4. The method of any one of claims 1-3, wherein the original mass spectrometry data comprises a full m / z range.

5. The method of any one of claims 1-4, wherein the original mass spectrometry data comprises a one-dimensional array or a two-dimensional matrix.

6. The method of claim 5, wherein the first pass comprises applying a machine learning classifier to the original mass spectrometry data.

7. The method of claim 6, wherein the original mass spectrometry data comprises the one-dimensional array or the two-dimensional matrix.

8. The method of claim 5, wherein the second pass comprises analyzing one or more targeted m / z range of the original mass spectrometry data.

9. The method of claim 8, wherein the second pass comprises applying a machine learning classifier to the one or more targeted m / z range of the original mass spectrometry data.WSGR Docket No. 70030-701.60110. The method of claim 8 or claim 9, wherein the original mass spectrometry data comprises the one-dimensional array or the two-dimensional matrix.

11. The method of any one of claims 1-10, further comprising performing a third pass comprising analyzing a second targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data.

12. The method of any one of claims 1-11, further comprising performing a fourth pass comprising analyzing a third targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data.

13. The method of any one of claims 1-12, wherein the original mass spectrometry data is obtained using a direct injection mass spectrometry or a hyphenated mass spectrometry.

14. The method of claim 13, wherein the direct injection mass spectrometry comprises one or more of electrospray ionization-mass spectrometry (ESI-MS), matrix-assisted laser desorption ionization-mass spectrometry (MALDI-MS), electron ionization-mass spectrometry (EI-MS), chemical ionization-mass spectrometry (CI-MS), atmospheric pressure chemical ionization-mass spectrometry (APCI-MS), direct analysis in real time-mass spectrometry (DART-MS), desorption electrospray ionization-mass spectrometry (DESIMS), inductively coupled plasma-mass spectrometry (ICP-MS), fast atom bombardmentmass spectrometry (FAB-MS), and secondary ion mass spectrometry (SIMS).

15. The method of claim 13, wherein the direct injection mass spectrometry comprises a tandem mass spectrometry (MS / MS), the MS / MS comprising one or more of ESI-MS / MS, MALDI-MS / MS, EI-MS / MS, CI-MS / MS, APCI-MS / MS, DART -MS / MS, DESI-MS / MS, ICP -MS / MS, FAB-MS / MS, and SIMS / MS.

16. The method of claim 13, wherein the hyphenated mass spectrometry comprises one or more of gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS), high-performance liquid chromatography-mass spectrometry (HPLC-MS), hydrophilic interaction chromatography-mass spectrometry (HILIC-MS), and capillary electrophoresis-mass spectrometry (CE-MS).WSGR Docket No. 70030-701.60117. The method of claim 13, wherein the hyphenated mass spectrometry comprises a tandem mass spectrometry (MS / MS) method, the MS / MS method comprising one or more of GC -MS / MS, LC-MS / MS, UHPLC-MS / MS, HPLC-MS / MS, HILIC-MS / MS, and CE- MS / MS.

18. The method of any one of claims 3-17, wherein the machine learning algorithm comprises one or more of linear regression, linear discriminant analysis, logistic regression, random forest, support vector machines, extreme gradient boosting, recurrent neural networks, convolutional neural networks, and transformers.

19. The method of any one of claims 3-17, wherein the machine learning algorithm comprises a neural network.

20. The method of any one of claims 3-19, further comprising using the machine learning algorithm to detect, classify, predict, stratify, screen, or monitor a disease, disorder, or condition.

21. The method of any one of claims 3-20, wherein the machine learning algorithm is trained using training data.

22. The method of claim 21, wherein the training data is obtained from a subject having or suspected of having the disease, disorder, or condition.

23. The method of any one of claims 3-22, further comprising determining a probability score.

24. The method of claim 23, wherein the probability score is based at least in part on a detection, classification, or prediction of the machine learning algorithm of a biological sample.

25. The method of claim 24, wherein the biological sample is a sample of unknown origin or an anonymized sample obtained or derived from a subject of a plurality of subjects.

26. The method of claim 24 or claim 25, wherein the biological sample comprises one or more of whole blood, blood serum, blood plasma, saliva, breath, urine, feces, sebum, tissue,WSGR Docket No. 70030-701.601 cerebrospinal fluid, seminal fluid, vaginal secretion, amniotic fluid, nasal fluid, otic fluid, interstitial fluid, breast milk, tears, sputum, sweat, and hair.

27. The method of any one of claims 3-26, further comprising using the machine learning algorithm to generate a performance metric.

28. The method of claim 27, wherein the performance metric comprises one or more of accuracy, sensitivity, specificity, precision, positive predictive value (PPV), negative predictive value (NPV), Mathew’s correlation coefficient (MCC), area under the receiver operating characteristic curve (AUROC), and area under the precision recall curve.

29. The method of claim 28, wherein the area under the receiver operating characteristic curve (AUROC) is 0.50 or higher.

30. A method comprising: a. obtaining a first biological sample of a subject; b. processing the biological sample using mass spectrometry, thereby obtaining original mass spectrometry data; c. analyzing the original mass spectrometry data, thereby obtaining processed mass spectrometry data; d. constructing a machine learning classifier based at least in part on the processed mass spectrometry data; e. using the machine learning classifier to identify a biomarker signature comprising one or more biomarker that is correlated with a disease, disorder, or condition; and f. applying the machine learning classifier to a second biological sample, thereby correlating the biomarker signature with the disease, disorder, or condition.

31. The method of claim 30, wherein analyzing the original mass spectrometry data comprises: a. performing a first pass comprising analyzing the original mass spectrometry data, and b. performing a second pass comprising analyzing a targeted mass-to-charge ratio (m / z) range of the original mass spectrometry data.WSGR Docket No. 70030-701.60132. The method of claim 30 or claim 31, wherein (d) further comprises constructing the machine learning classifier based at least in part on subject data.

33. The method of claim 32, wherein the subject data comprises one or more of biological sex, age, ethnicity, height, weight, diet, symptomatic information, family history, medication use, genetic history, and electronic health records.

34. The method of any one of claims 30-33, wherein the first or the second biological sample is a sample of unknown origin or an anonymized sample.

35. The method of any one of claims 30-34, wherein the second biological sample is obtained or derived from the first subject or from a second subject.

36. The method of claim 35, wherein the first subject or the second subject is a mammal.

37. The method of claim 36, wherein the first subject or the second subject is a human.

38. The method of any one of claims 30-37, wherein the disease, disorder, or condition comprises one or more of an oncological, immunological, neurological, respiratory, cardiovascular, gastrointestinal, or mental health disease, disorder, or condition.

39. The method of any one of claims 30-38, wherein the disease, disorder, or condition comprises a phenotype.

40. The method of any one of claims 30-39, wherein the disease, disorder, or condition comprises a viral infection or a bacterial infection.

41. The method of any one of claims 30-40, wherein the first or the second biological sample comprises one or more of whole blood, blood serum, blood plasma, saliva, breath, urine, feces, sebum, tissue, cerebrospinal fluid, seminal fluid, vaginal secretion, amniotic fluid, nasal fluid, otic fluid, interstitial fluid, breast milk, tears, sputum, sweat, and hair.

42. The method of any one of claims 30-41, wherein (b) comprises using a direct injection mass spectrometry or a hyphenated mass spectrometry.WSGR Docket No. 70030-701.60143. The method of claim 42, wherein the direct injection mass spectrometry comprises one or more of electrospray ionization-mass spectrometry (ESI-MS), matrix-assisted laser desorption ionization-mass spectrometry (MALDI-MS), electron ionization-mass spectrometry (EI-MS), chemical ionization-mass spectrometry (CI-MS), atmospheric pressure chemical ionization-mass spectrometry (APCI-MS), direct analysis in real time-mass spectrometry (DART-MS), desorption electrospray ionization-mass spectrometry (DESIMS), inductively coupled plasma-mass spectrometry (ICP-MS), fast atom bombardmentmass spectrometry (FAB-MS), and secondary ion mass spectrometry (SIMS).

44. The method of claim 42, wherein the direct injection mass spectrometry comprises a tandem mass spectrometry (MS / MS), the MS / MS comprising one or more of ESI-MS / MS, MALDI-MS / MS, EI-MS / MS, CI-MS / MS, APCI-MS / MS, DART -MS / MS, DESI-MS / MS, ICP -MS / MS, FAB-MS / MS, and SIMS / MS.

45. The method of claim 42, wherein the hyphenated mass spectrometry comprises one or more of gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS), high-performance liquid chromatography-mass spectrometry (HPLC-MS), hydrophilic interaction chromatography-mass spectrometry (HILIC-MS), and capillary electrophoresis-mass spectrometry (CE-MS).

46. The method of claim 42, wherein the hyphenated mass spectrometry comprises a tandem mass spectrometry (MS / MS) method, the MS / MS method comprising one or more of GC -MS / MS, LC-MS / MS, UHPLC-MS / MS, HPLC-MS / MS, HILIC-MS / MS, and CE- MS / MS.

47. The method of any one of claims 30-46, wherein the biomarker comprises one or more of a metabolite, nucleic acid, protein, peptide, and lipid.

48. The method of any one of claims 30-47, wherein the biomarker comprises a feature.

49. The method of claim 48, wherein the biomarker comprises a chemical or a nonchemical feature.WSGR Docket No. 70030-701.60150. The method of claim 49, wherein the biomarker comprises an ion feature.

51. The method of claim 50, wherein the ion feature is detected in a positive or negative ionization mode of the mass spectrometry.

52. The method of any one of claims 30-51, wherein the biomarker is obtained using untargeted mass spectrometry or targeted mass spectrometry.

53. The method of any one of claims 30-52, wherein the biomarker signature is determined using one or more of Shapley Additive ExPlanation (SHAP), statistical significance, and feature importance procedure.

54. The method of any one of claims 30-53, wherein the biomarker signature is optimized using the machine learning classifier.

55. The method of any one of claims 30-54, wherein the biomarker signature comprises an absolute abundance or a relative ratio of the biomarker.

56. The method of claim 55, wherein the biomarker signature comprises a panel of chemical or non-chemical features.

57. The method of claim 55, wherein the biomarker signature comprises an ion heatmap based on m / z values with respect to time.

58. The method of claim 55, wherein the biomarker signature comprises an ion heatmap based on m / z values from mass spectrometry data of a first mass spectrometry stage (MSI; mi / zi) with respect to m / z values from mass spectrometry data of a second mass spectrometry stage (MS2; m2 / z2).

59. The method of any one of claims 30-58, further comprising using the biomarker signature to monitor disease, disorder, or condition progression.

60. The method of any one of claims 30-59, further comprising using the biomarker signature to monitor the response of a subject to a disease-modifying therapy or intervention.WSGR Docket No. 70030-701.60161. The method of any one of claims 30-60, further comprising using the biomarker signature to identify predisposition of the subject to the disease, disorder, or condition.

62. The method of any one of claims 30-61, further comprising using the biomarker signature to identify different phenotypes of the disease, disorder, or condition.

63. The method of any one of claims 30-62, further comprising identifying a disease mechanism or a druggable target based at least in part on analyzing the original mass spectrometry data.

64. The method of any one of claims 30-63, wherein the machine learning classifier comprises one or more of a linear regression, linear discriminant analysis, logistic regression, random forest, support vector machines, extreme gradient boosting, recurrent neural network, convolutional neural network, and a transformer.

65. The method of any one of claims 30-63, wherein the machine learning classifier comprises a neural network.

66. The method of any one of claims 30-65, wherein the machine learning classifier is trained using training data.

67. The method of claim 66, wherein the training data is obtained or derived from a subject having or suspected of having the disease, disorder, or condition.

68. The method of claim 66 or claim 67, further comprising using the machine learning classifier to detect, classify, predict, stratify, screen, or monitor the disease, disorder, or condition based on the training of data obtained or derived from the subject having or suspected of having the disease, disorder, or condition.

69. The method of any one of claims 30-68, further comprising determining a probability score.

70. The method of claim 69, wherein the probability score is based at least in part on a detection, classification, or prediction of the machine learning classifier of the second biological sample.WSGR Docket No. 70030-701.60171. The method of any one of claims 30-70, wherein the machine learning classifier is used to generate a classification performance metric comprising one or more of accuracy, sensitivity, specificity, precision (or positive predictive value, PPV), negative predictive value (NPV), Mathew’s correlation coefficient (MCC), area under the receiver operating characteristic curve, and area under the precision recall curve.

72. The method of claim 71, wherein the area under the receiver operating characteristic curve (AUROC) is 0.50 or higher.

73. A method of processing original mass spectrometry data, wherein the original mass spectrometry data is an input data for a machine learning classifier.

74. The method of claim 73, wherein the original mass spectrometry data comprises one or more of data in a .raw, ,wiff2, .D, .LCD, .CDF, .mzXML, or .mzML format.

75. A method of processing original mass spectrometry data, comprising processing the original mass spectrometry data using a machine learning classifier, wherein the original mass spectrometry data comprise a one-dimensional data array.

76. The method of claim 75, wherein an ion abundance of the original mass spectrometry data for an m / z value is summed over time, providing a single summed abundance value for the m / z value.

77. The method of claim 76, wherein the summed abundance value is associated with an output indicative of the presence of a disease, disorder, or condition.

78. The method of claim 76 or claim 77, wherein the ion abundance of the m / z value is summed over an m / z range, providing a summed abundance value for the m / z range.

79. A method of processing original mass spectrometry data, comprising processing the original mass spectrometry data using a machine learning classifier, wherein the original mass spectrometry data comprise a two-dimensional matrix.WSGR Docket No. 70030-701.60180. The method of claim 79, wherein the x-axis represents m / z and the y-axis represents time (t), and wherein the value for each (m / z, t) pair is the ion abundance value corresponding to the (m / z, t) pair.

81. The method of claim 79, wherein the x-axis represents an m / z comprising first stage mass spectrometry data (MSI; mi / zi) and the y-axis represents a second m / z comprising second stage mass spectrometry data (MS2; m2 / z2), and wherein the value for each (mi / zi, m2 / z2) pair is the ion abundance value corresponding to the (mi / zi, 1112 / Z2 ) pair.

82. The method of claim 79, wherein the x-axis represents an m / z comprising second stage mass spectrometry data (MS2; m2 / z2) and the y-axis represents a second m / z comprising first stage mass spectrometry data (MSI; mi / zi), and wherein the value for each (m2 / z2, mi / zi) pair is the ion abundance value corresponding to the (m2 / z2, mi / zi) pair.

83. The method of any one of claims 80-82, wherein an ion abundance of an m / z value of the original mass spectrometry data is summed over an m / z range, providing a summed abundance value for the m / z range.

84. The method of claim 83, wherein the ion abundance of the m / z range of the original mass spectrometry data is summed over a second m / z range, providing a summed abundance value for the second m / z range.

85. The method of claim 80, wherein the ion abundance of the m / z value of the original mass spectrometry data is summed over a time range, providing a summed abundance value for the time range with respect to the m / z value.

86. The method of claim 85, wherein the ion abundance of the m / z value is further summed over an m / z range, providing a summed abundance value for the m / z range with respect to the time range.

87. The method of claim 86, wherein the ion abundance of the m / z range is further summed over a second m / z range, providing a summed abundance value for the second m / z range with respect to the time range.WSGR Docket No. 70030-701.60188. The method of claim 86, wherein the ion abundance of the time range is further summed over a second time range, providing a summed abundance value for the second time range with respect to the m / z range.