Novel systems and methods for early detection of multiple cancers

A metabolomics and AI/ML-integrated LC-MS system addresses the limitations of current cancer detection methods by offering a non-invasive, accurate, and cost-effective method for simultaneous detection of multiple cancers, enhancing global accessibility and survival rates.

JP2026505709APending Publication Date: 2026-02-18プレドミックス ヘルス サイエンシーズ プライベート リミテッド
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
JP2025540804
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-11
Filing Date
2024-01-11
Publication Date
2026-02-18

AI Technical Summary

Technical Problem

Current methods for early cancer detection are inadequate in accuracy, cost-effective, and invasive, lacking effective screening tests for multiple cancers, particularly for ovarian, endometrial, and breast cancer, and are not accessible across diverse economic spectrums.

Method used

A comprehensive metabolomics approach integrated with advanced artificial intelligence (AI) and machine learning (ML) processes using liquid chromatography-mass spectrometry (LC-MS) to analyze biological fluids, employing AI models like Cancer Detection AI (CDAI) and Tissue of Origin Identification (TOOAI) for accurate and non-invasive simultaneous detection of multiple cancers.

Benefits of technology

The system provides high-accuracy, non-invasive, and cost-effective early detection of multiple cancers, improving survival chances by minimizing invasive procedures and enhancing accessibility across various income levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505709000001_ABST
    Figure 2026505709000001_ABST
Patent Text Reader

Abstract

The present invention describes a comprehensive system and method for the simultaneous early detection of multiple cancers in a single analysis. The system involves a liquid chromatography-mass spectrometry (LC-MS) instrument coupled with a processor and AI / ML algorithms. The LC-MS instrument analyzes metabolite ions from dried extracts of biological fluid samples and aligns and normalizes the data while minimizing errors. Quality control processes, including neural network models and critical ion monitoring, ensure accurate detection. The system employs an AI / ML process to create two models: a Cancer Detection AI (CDAI) model for identifying cancer samples and a Tissue of Origin Identification (TOOAI) model for distinguishing specific cancer types. These models are applied to test samples and provide a score based on tissue of origin probability. The present invention aims to revolutionize early cancer detection through advanced analytical and machine learning techniques.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of clinical metabolomics and the use of metabolite biosignatures captured by machine learning for the detection of multiple early-stage cancers in adult male and female mammals. [Background technology]

[0002] Cancer is a leading cause of death worldwide, and the burden of disease is expanding in countries of all income levels due to growth and aging. In India, the estimated number of people living with the disease is approximately 2.25 million. Approximately 1.1 million cancer cases occur annually, with a mortality rate of approximately 700,000 per year. The risk of developing cancer before age 75 in men and women is 9.81% and 9.42%, respectively (http: / / cancerindia.org.in / cancer-statistics / ). In the United States, 1,700 people are expected to die from cancer every day (American Cancer Society. Economic impact of Cancer. Page revised January 3, 2018. Accessed July 16, 2020. cancer.org / cancer / cancer-basics / economic-impact-of-cancer.html). The American Cancer Society (ACS) projections for 2020 include just over 1.8 million new cancer cases and approximately 600,000 cancer deaths (Siegel RL, Miller KD, Jemal A. Cancer statistics, 2020. CA Cancer J Clin. 2020;70(1):7-30. doi:10.3322 / caac.21590). Disease Control and Prevention (CDC) provided similar estimates of new cancer cases and deaths (CDC. Expected New Cancer Cases and Deaths in 2020. Page outlined August 16, 2018. Accessed July 16, 2020. cdc.gov / cancer / dcpc / research / articles / cancer_2020.htm). Despite Europe having only one-tenth the world's population density, one-quarter of all cancer diagnoses occur in the region. Breast, prostate, lung and colorectal cancers account for more than half of all cancer diagnoses in Europe (https: / / canceratlas.cancer.org / the-burden / europe / ).

[0003] It is widely recognized that identifying cancer at an early stage is paramount for optimal prognosis, as identifying cancer at an early stage can significantly reduce mortality in the long term. Unfortunately, effective methods for early cancer detection are either lacking or not sensitive enough for many cancers of major public health relevance. For example, reliable methods for detecting early-stage ovarian cancer and asymptomatic endometrial cancer do not yet exist. Similarly, for breast cancer, existing detection methods suffer from limitations, including high cost, time consumption, and / or insufficient effectiveness. For example, while the prostate-specific antigen (PSA) test is widely recommended for prostate cancer screening, this test is considered inaccurate and fails to identify 8 out of 10 men under the age of 60 who are subsequently diagnosed with prostate cancer (BMJ. 2003 Aug. 2;327(7409):249). In another example, mammography is a commonly recommended method for early detection, but its relatively high false-positive and false-negative rates pose a problem, especially in patients with dense breasts. The efficacy of biomarker-based approaches employing either DNA or protein markers is similarly compromised due to either poor penetration (DNA) or low circulating concentrations (protein) in risk groups. In another example, lung cancer screening with low-dose CT scans (LDCT) has been shown to reduce the risk of death from lung cancer. However, this screening platform is based on radiation therapy, and its long-term use, including multiple patient visits to control tumor growth, may further lead to several other health problems. In addition, this test has been shown to have a false-positive rate of 12–14%, which is too high to qualify as a screening test. In yet another example, there are no widely recommended screening tests available for patients at average risk for cancer (e.g., leukemia, thyroid cancer, melanoma, kidney cancer, lymphoma, pancreatic cancer, liver cancer, and bile duct cancer). However, for some cancers, such as colorectal cancer, gastric cancer, and head and neck cancer, no blood-based screening tests are currently available.These cancers (e.g., colorectal, gastric, and head and neck cancers) can only be detected early using either physical examination or inherently invasive procedures. Finally, screening strategies for early-stage cervical cancer exist, but their impact is limited to less developed areas of the world, where approximately 85% of new cases occur. Even though early detection of disease can lead to lead-time bias (when patients live longer due to early detection) and length bias (when early detection tests preferentially detect slower-growing cancers, creating a false impression of longer survival), there remains a significant opportunity to reduce the cancer burden with effective early detection. The true benefits of early detection are realized only if effective early treatment leads to better patient outcomes and should not be confused with these biases (THE AMERICAN JOURNAL OF MANAGED CARE® Supplement Vol. 26, No. 14, S293). These factors therefore highlight the need to develop new methods that can detect early-stage cancers with high accuracy and are affordable across the economic spectrum. In this context, an integrated test that can simultaneously screen for multiple cancers would offer a clear advantage.

[0004] Metabolomics is an emerging field broadly defined as the comprehensive measurement of all metabolites and small molecules in a biological sample. Metabolomics allows for the profiling of a much larger number of metabolites than currently covered by standard clinical laboratory techniques. Thus, metabolomics facilitates comprehensive coverage of biological processes and metabolic pathways. For this reason, metabolomics holds promise as an essential objective in precision molecular microscopy. This is particularly relevant because metabolites have been described as proximal reporters of disease, as their abundance in biological samples is often directly related to pathogenic mechanisms.

[0005] The idea that the metabolite composition of biological fluids reflects an individual's health has existed for many years. Credibility in this hypothesis comes from recent applications to find early metabolic indicators of disease in longitudinal cohorts, for example, in pancreatic cancer, type 2 diabetes, cardiovascular disease, memory impairment, and many other conditions, several years before symptoms become clinically apparent. Metabolomics research has also stimulated studies that reveal novel insights into the relationship between diet and disease, such as the observations linking elevated branched-chain amino acids and obesity to insulin resistance. Therefore, such studies strongly support the idea that metabolomics, combined with multivariate statistical analysis, provides a relatively simple and efficient method for identifying risk factors and / or biomarkers of disease.

[0006] Metabolomics is a particularly relevant technology for cancer detection. Cancer cells have significantly altered metabolism, and therefore, the patterns of metabolites produced can provide a "signature" indicative of the presence or behavior of cancer. Importantly, in contrast to gene expression profiling as a risk stratifier, this signal is derived not from characteristics of the primary tumor but from micrometastatic disease, either directly or indirectly. As a result, metabolome-derived signatures provide highly accurate risk stratification factors for disease, with an accuracy that can far exceed that of methods based on DNA or protein markers. However, untargeted metabolomic profiles are inherently complex and multivariate and cannot be accurately analyzed by linear analytical methods. However, such data are readily amenable to the application of AI-based methodologies. By exploring nonlinear variables in the data that correlate with defined clinical conditions, metabolite signatures characteristic of a given disease state can potentially be extracted.

[0007] Metabolomics is now frequently used in oncology research, with particular emphasis on early cancer diagnosis, surveillance, and prognosis. For example, several studies have utilized metabolomics analysis for both breast cancer diagnosis and prognosis. However, overall, these studies suffer from variable results and limited accuracy. Similarly, the application of metabolomics to endometrial cancer has led to the identification of metabolites that can predict the presence of cancer, tumor behavior, and even pathological features. However, these findings await validation. Recent analyses have identified metabolite signatures for cervical intraepithelial neoplasia and cervical cancer. However, sample sizes were relatively small, and the discriminatory power of the tests was suboptimal. Metabolomic approaches for the diagnosis of ovarian cancer have recently been reviewed. The inference was that metabolomics offers important new opportunities for ovarian cancer diagnosis, but further research is needed.

[0008] U.S. Patent No. 9,459,255 discloses amino acids useful for distinguishing between individuals with and without breast cancer. Multivariate discriminants were found that included the concentrations of the identified amino acids as explanatory variables and were significantly correlated with breast cancer status. However, the sensitivity of this method was only about 87%, and the specificity was about 85%.

[0009] US Patent No. 5,162,504 (1992) discloses the use of monoclonal antibodies that target prostate-specific membrane antigen (PSMA) and act as cytogenic imaging agents for prostate cancer.

[0010] U.S. Patent Application Publication No. 2011 / 0143444 discloses a method for assessing female reproductive cancer using amino acid concentrations in blood collected from a subject. This method assesses the status of female reproductive cancer, including at least one of cervical cancer, endometrial cancer, and ovarian cancer, in a subject. However, the total number of subject samples tested was small, and the discriminatory power of this method was weak, ranging from 55% to 81% for individual cancers.

[0011] US Patent Application Publication No. 20120100558A1 demonstrated the development of lung cancer by screening biological fluids from patients. This invention relies on the presence of autoantibodies specific for one or more pre-diagnostic lung cancer indicator proteins, such as LAMR1, and additionally or alternatively, Annexin I and / or 14-3-3-theta.

[0012] U.S. Patent Application Publication No. 2017 / 0003291 is directed to a method for diagnosing endometrial cancer by detecting changes in the concentrations of specific lipids and several small metabolites in patient-derived biological samples. Using metabolomic analysis based on a combination of NMR and mass spectrometry (MS), statistically significant changes were found in the serum of endometrial cancer patients compared with unaffected controls. However, despite employing two separate metabolomic analysis techniques, the resulting sensitivity and specificity of this method were only in the range of 70% to 80%.

[0013] US Patent Application Publication No. 2017 / 0097355 describes a method for measuring metabolic changes useful for distinguishing between ovarian cancer and benign ovarian tumors. Two independent LC-MS-based metabolomics platforms (including a comprehensive lipidomics approach) were used to screen for differentially abundant plasma metabolites between cases with serous ovarian cancer and controls with benign serous ovarian tumors. Combining small molecules with lipidomic profiling resulted in a test with good sensitivity (95%); however, specificity was less than 50%. This limits the usefulness of the test for patient screening.

[0014] Therefore, from all of these studies, it is clear that better methods with higher fidelity are needed for the early diagnosis of multiple cancers simultaneously. Furthermore, it is clear that screening for early cancers greatly benefits the patient's chances of survival. Therefore, it is highly desirable to develop a single non-invasive test that can efficiently screen for multiple cancers simultaneously using small amounts of biological fluid. [Prior art documents] [Patent documents]

[0015] [Patent Document 1] U.S. Patent No. 9,459,255 [Patent Document 2] U.S. Patent No. 5,162,504 [Patent Document 3] US Patent Application Publication No. 2011 / 0143444 [Patent Document 4] U.S. Patent Application Publication No. 20120100558A1 [Patent Document 5] US Patent Application Publication No. 2017 / 0003291 [Patent Document 6] US Patent Application Publication No. 2017 / 0097355 [Non-patent literature]

[0016] [Non-Patent Document 1] American Cancer Society. Economic Impact of Cancer. Page revised January 3, 2018. Accessed July 16, 2020. cancer.org / cancer / cancer-basics / economic-impact-of-cancer.html [Non-patent document 2] Siegel RL, Miller KD, Jemal A. Cancer statistics, 2020. CA Cancer J Clin. 2020;70(1):7-30. doi:10.3322 / caac.21590 [Non-patent document 3] CDC. Expected New Cancer Cases and Deaths in 2020. Page outlined August 16, 2018. Accessed July 16, 2020. cdc.gov / cancer / dcpc / research / articles / cancer_2020.htm [Non-patent document 4] BMJ.August 2, 2003;327(7409):249 [Non-Patent Document 5] THE AMERICAN JOURNAL OF MANAGED CARE® Supplement Volume 26, Number 14, S293 Summary of the Invention [Problem to be solved by the invention]

[0017] The objective of the present invention is to revolutionize the early detection of multiple cancers in both adult males and females. The innovation focuses on utilizing clinical metabolomics and machine learning to capture metabolite biosignatures for accurate and comprehensive cancer detection. Keeping in mind the global significance of cancer as a leading cause of mortality, the present invention aims to address the burden of this disease globally and improve early detection rates, targeting countries of various income levels, including India and the United States, and various other countries.

[0018] It is yet another object of the present invention to overcome limitations associated with existing early cancer detection methods, particularly for cancers such as ovarian, endometrial, and breast cancer. The present invention addresses issues related to the accuracy, cost, time consumption, and effectiveness of current detection methods. To achieve this, a comprehensive metabolomics approach is employed, utilizing metabolomics as a precision risk stratifier to gain essential insight into metabolic pathways associated with cancer.

[0019] Another objective of the present invention is the integration of advanced artificial intelligence (AI) and machine learning (ML) processes, which play a key role in improving the accuracy of cancer detection. Emphasis is placed on the potential of AI / ML to analyze complex and multivariate metabolomic profiles for efficient disease identification. The present invention seeks to develop a noninvasive test that can simultaneously screen for multiple cancers through a single analysis, minimizing the need for invasive procedures and providing an efficient screening approach for various cancer types. The core technology used to resolve metabolites in biological fluid samples is liquid chromatography-mass spectrometry (LC-MS). The focus is on optimizing LC-MS to ensure accurate measurement of metabolite ion masses and obtaining ion spectra. A robust quality control process is established to identify and correct errors in the detection of multiple cancers. This includes the implementation of sequential neural network models, monitoring critical ions, and assessing matrix occupancy to improve data accuracy.

[0020] A further objective of the present invention includes the creation of specific AI models, namely, a Cancer Detection AI (CDAI) model for discriminating cancerous samples from normal samples, and a Tissue of Origin Identification (TOOAI) model for identifying specific cancer types based on their tissue origin using multi-class classification. Aiming for global impact and affordability, the present invention aims to contribute to reducing cancer-related mortality across diverse populations, prioritizing accessibility across a diverse economic spectrum. Validation and accuracy assessment are of paramount importance, with emphasis placed on validating the AI ​​model through logistic regression, class balancing, and optimization processes. Systematic evaluation using training and test datasets ensures the accuracy of the developed methodology. Finally, the present invention aims to revolutionize early cancer detection by combining advanced technologies, comprehensive metabolomics, and AI / ML methodologies, making a significant and impactful contribution to the field of precision medicine. [Means for solving the problem]

[0021] The present invention relates to a system and method for the simultaneous early detection of multiple cancers in a single analysis. The system involves a liquid chromatography-mass spectrometry (LC-MS) instrument, a processor for data analysis and quality control, and an AI / ML process for cancer detection. The LC-MS instrument resolves metabolites in a biological fluid sample, and the system aligns and normalizes the mass data while minimizing errors. The quality control process includes building a neural network, monitoring critical ions, and assessing matrix occupancy. The AI / ML process creates a cancer detection AI (CDAI) model and a tissue of origin identification (TOOAI) model to distinguish cancerous samples from normal samples and identify specific cancer types. The LC-MS instrument also includes sample collection, extraction, and reconstitution components. The method includes analyzing metabolite ions, applying quality control, and using the AI / ML process for cancer detection and tissue origin identification. The AI ​​models are created using logistic regression and multiclass classification. The accuracy of these models is evaluated using training and test datasets. The system aims to revolutionize early cancer detection through comprehensive analysis and advanced machine learning techniques.

[0022] Another embodiment of the present invention is a system for the simultaneous detection of multiple cancers at an early stage in a single analysis, comprising at least one liquid chromatography (LC) apparatus with a mass spectrometer (MS) (hereinafter abbreviated as LC-MS) for analyzing one or more obtained reconstituted metabolites using LC-MS techniques to measure masses of metabolite ions in the one or more obtained reconstituted metabolites, wherein the one or more obtained reconstituted metabolites are obtained after reconstitution of one or more dried metabolite extracts extracted from one or more biological fluid samples; at least one processor / computing device for aligning masses obtained from a metabolomic profile of metabolite ions in the one or more obtained reconstituted metabolites and minimizing possible errors in the measurement of the masses of the metabolite ions; and at least one processor / computing device for applying one or more quality control processes to the aligned and normalized dataset to identify possible errors, if any, in the detection of multiple cancers in the biological fluid sample, wherein the at least one processor / computing device for performing the quality control process comprises the steps of: (a) constructing a sequential neural network model to detect errors based on variations in chromatogram profiles of defective sample extractions or due to errors in mass spectrometry; (b) monitoring the presence of at least six of nine critical ions having m / z values ​​in the range of 100 to 800, wherein the presence of six or more of the nine critical ions in a sample chromatogram is interpreted as a threshold criterion for passing the quality control step; and (c) assessing matrix occupancy, including calculating a minimum matrix occupancy threshold using multiple runs of low quality samples or inappropriate mass spectrometry run conditions. a processor / computing device configured to execute at least one of the following: at least one processor that executes one or more AI / ML processes on the measured metabolite ions, the one or more AI / ML processes for: creating a first AI model (Cancer Detection AI (CDAI) model) for identifying diseased cancer samples and distinguishing them from non-diseased normal samples, wherein the measured metabolite ions are randomly divided into a training dataset and a test dataset; creating a second AI model (Tissue of Origin Identification (TOOAI) model) for further identifying individual diseased cancer samples and distinguishing them from other diseased cancer samples and non-diseased normal samples; and applying the TOOAI model to the cancer-positive samples determined by the CDAI model to obtain a score assigned to each individual diseased cancer sample, thereby identifying the individual diseased cancer samples and distinguishing them from other diseased cancer samples and non-diseased normal samples, one score for each cancer type defined by its tissue of origin and representing the probability that the sample belongs to the respective cancer type; and The present invention provides a system including:

[0023] Yet another embodiment of the present invention provides a system wherein the LC-MS device is configured to resolve the one or more obtained reconstituted metabolites by ultra-high performance liquid chromatography using the LC device, obtain an ion spectrum of the one or more obtained reconstituted metabolites through the MS device, and measure the masses of metabolite ions present in the ion spectrum of the one or more obtained reconstituted metabolites based on their mass-to-charge ratio or m / z using the MS device.

[0024] A further embodiment of the present invention provides a system wherein, to create the first AI model (Cancer Detection AI (CDAI) model), the at least one processor / computing device is configured to: apply a logistic regression function by running the AI / ML process on a training dataset of metabolite ions to find a functional mapping between a dependent / target variable and independent variables that separates diseased cancer samples from non-diseased normal samples; configure one or more class balancing parameters for the target variable to balance class imbalance in the training dataset; apply an optimization process to handle data complexity, thereby creating a first AI model from the training dataset, wherein the first AI model is a Cancer Detection AI (CDAI) model; and apply the CDAI model to a test dataset of metabolite ions to identify diseased cancer samples and distinguish them from non-diseased normal samples based on the function that separates diseased cancer samples from non-diseased normal samples. Further, to create the second AI model (tissue of origin identification (TOOAI model)), the at least one processor / computing device is configured to: use the training dataset to construct a classifier multi-class classification model, thereby creating a second AI model from the training dataset, wherein the second AI model is the tissue of origin identification (TOOAI model), and the classifier multi-class classification model includes at least one of a support vector machine, a logistic one versus rest, or a stochastic gradient descent process; and apply the TOOAI model to the cancer-positive samples determined by the CDAI model to obtain a score assigned to each individual diseased cancer sample, thereby identifying individual diseased cancer samples and distinguishing them from other diseased cancer samples and non-diseased normal samples, wherein one score for each cancer type is defined by its tissue of origin and represents the probability that the sample belongs to the respective cancer type.The system further comprises at least one sample collection device for collecting one or more biological fluid samples from one or more living mammals, at least one precipitation device for extracting one or more metabolite extracts from the one or more biological fluid samples by precipitation of proteins present in the biological fluid samples with chilled alcohol comprising at least methanol, at least one phase separation device for drying the one or more metabolite extracts extracted from the at least one precipitation device, and at least one device for reconstituting the one or more dried metabolite extracts in an aqueous solution in a mobile phase.

[0025] A further embodiment of the present invention provides a system wherein the at least one processor / computing device for creating a first AI model (Cancer Detection AI (CDAI)) is further configured to use a learned (trained) model / process derived from the CDAI model to find a score for each sample in the training set and evaluate a test set to determine an accuracy rate for applying the learned CDAI model.

[0026] In a further embodiment of the present invention, the at least one processor / computing device performing the AI / ML process to create a second AI model (tissue of origin identification (TOOAI model)) is further configured to determine the accuracy rate of the multi-class classification model in the TOOAI model in distinguishing individual diseased cancer samples from other diseased cancer samples and non-diseased normal samples, and in determining the accuracy rate, the at least one processor / computing device performing the AI / ML process: obtaining and providing a probability score for each cancer subclass for any given sample of individual diseased cancer samples, wherein the probability score for one individual diseased cancer sample is distinct from the probability scores for other diseased cancer samples and other diseased cancer samples; The top two scoring cancer subclasses are taken as the model results and matched with the true labels; and Constructing a final confusion matrix based on the model's dual-class accuracy rate The system is further configured to:

[0027] Yet another embodiment of the present invention is a method for simultaneous detection of multiple cancers at an early stage in a single analysis, comprising the steps of: analyzing one or more obtained reconstituted metabolites using LC-MS techniques to measure masses of metabolite ions in the one or more obtained reconstituted metabolites, wherein the one or more obtained reconstituted metabolites are obtained after reconstitution of one or more dried metabolite extracts extracted from one or more biological fluid samples; aligning masses obtained from metabolomic profiles of metabolite ions in the one or more obtained reconstituted metabolites to minimize possible errors in measuring the masses of the metabolite ions; and applying one or more quality control processes to the aligned and normalized dataset to identify possible errors, if any, in the detection of multiple cancers, wherein the one or more quality control processes comprise the steps of: (a) constructing a sequential neural network model to detect errors based on variations in chromatogram profiles of defective sample extractions or due to errors in mass spectrometry; (b) monitoring the presence of at least six of nine critical ions having m / z values ​​in the range of 100 to 800, wherein the presence of six or more of the nine critical ions in a sample chromatogram is interpreted as a threshold criterion for passing the quality control step; and (c) assessing matrix occupancy, including calculating a minimum matrix occupancy threshold using multiple runs of low quality samples or inappropriate mass spectrometry run conditions. and a step configured to perform at least one of the following: performing one or more AI / ML processes on the measured metabolite ions to create a first AI model (Cancer Detection AI (CDAI) model) for identifying diseased cancer samples and distinguishing them from non-diseased normal samples, wherein the measured metabolite ions are randomly divided into a training dataset and a test dataset; creating a second AI model (Tissue of Origin Identification (TOOAI) model) for further identifying individual diseased cancer samples and distinguishing them from other diseased cancer samples and non-diseased normal samples; and applying the TOOAI model to the cancer-positive samples determined by the CDAI model to obtain a score assigned to each of the individual diseased cancer samples, thereby identifying and distinguishing each diseased cancer sample from other diseased cancer samples and non-diseased normal samples, one score for each cancer type defined by its tissue of origin and representing the probability that the sample belongs to the respective cancer type; The present invention provides a method comprising:

[0028] A further embodiment of the present invention provides a method wherein the LC-MS technique further comprises resolving the one or more obtained reconstituted metabolites by ultra-high performance liquid chromatography using the LC instrument, obtaining an ion spectrum of the one or more obtained reconstituted metabolites through an MS instrument, and measuring the masses of metabolite ions present in the ion spectrum of the one or more obtained reconstituted metabolites based on their mass-to-charge ratio or m / z using the MS instrument.

[0029] Yet another embodiment of the present invention provides a method wherein, to create the first AI model (Cancer Detection AI (CDAI) model), the one or more AI / ML processes are further performed to: apply a logistic regression function to a training dataset of metabolite ions to find a functional mapping between a dependent / target variable and independent variables that separates diseased cancer samples from non-diseased normal samples; configure one or more class balancing parameters for the target variable to balance class imbalance in the training dataset; apply an optimization process to handle data complexity, thereby creating a first AI model from the training dataset, wherein the first AI model is a Cancer Detection AI (CDAI) model; and apply the CDAI model to a test dataset of metabolite ions to identify diseased cancer samples and distinguish them from non-diseased normal samples based on the function that separates diseased cancer samples from non-diseased normal samples.

[0030] Yet another embodiment of the present invention provides a method further comprising: using the training dataset to construct a classifier multi-class classification model to create the second AI model (tissue of origin identification (TOOAI model)), thereby creating a second AI model from the training dataset, wherein the second AI model is the tissue of origin identification (TOOAI model), and the classifier multi-class classification model includes at least one or more of a support vector machine, a logistic one-to-many, or a stochastic gradient descent process; and applying the TOOAI model to the cancer-positive samples determined by the CDAI model to obtain a score assigned to each individual diseased cancer sample, thereby identifying individual diseased cancer samples and distinguishing them from other diseased cancer samples and non-diseased normal samples, wherein one score for each cancer type is defined by its tissue of origin and represents the probability that the sample belongs to the respective cancer type.

[0031] Another embodiment of the present invention provides the method further comprising the steps of collecting the one or more biological fluid samples from one or more living mammals, extracting one or more metabolite extracts from the one or more biological fluid samples by precipitation of proteins present in the biological fluid samples with chilled alcohol containing at least methanol, drying the one or more metabolite extracts extracted from at least one precipitation device, and reconstituting the one or more dried metabolite extracts in an aqueous solution in a mobile phase. Further, to create the first AI model (Cancer Detection AI (CDAI)), the one or more AI / ML processes are further performed to find a score for each sample in a training set using a trained model / process obtained from the CDAI model, and to evaluate a test set to determine the accuracy rate of applying the trained CDAI model. Furthermore, to create the second AI model (tissue of origin identification (TOOAI model)), the one or more AI / ML processes are further performed to determine the accuracy rate of the multi-class classification model in the TOOAI model in distinguishing individual diseased cancer samples from other diseased cancer samples and non-diseased normal samples, and in determining the accuracy rate, the one or more AI / ML processes are further performed to obtain and provide a probability score for each cancer subclass for any given individual diseased cancer sample, wherein the probability score of one individual diseased cancer sample distinguishes the probability scores from other diseased cancer samples and other diseased cancer samples; the top two scoring cancer subclasses are taken as model results and matched with the true labels; and construct a final confusion matrix based on the dual-class accuracy rate of the model.

[0032] A further embodiment of the present invention provides a system in which a metabolomic profile of metabolite ions is generated using an automated platform that includes at least a compound discoverer. [Brief explanation of the drawings]

[0033] [Figure 1] Figure 1 shows the workflow of the overall steps involved in cancer detection. [Figure 2]Figure 2 shows a schematic diagram of the entire process under study. This process involves sample preparation, which primarily relies on protein precipitation to extract metabolites. The extracted metabolites were phase-separated, and the extract was dried under vacuum. Finally, UHPLC-HRMS was used to separate metabolites based on their retention capacity on a Waters Acquity UPLC HSS T3 column (1.8 microns (1.8 μm), dimensions - 2.1 × 100 mm, part number 186003539). Prior to the AI / ML workflow, the sample was subjected to quality check modules QC, QC2, and QC3 for sample extraction and chromatogram authentication. These separated features were then subjected to AI / ML-based analysis for pattern recognition. [Figure 3] Figure 3 shows the age distribution of samples across healthy and cancer individuals. A total of 8971 cancer serum samples of the 33 mentioned cancers were collected, with 3914 samples serving as a normal control set. [Figure 4] Figure 4 shows the number of metabolites present across normal control and 33 mentioned cancer samples. Cancer and normal controls are grouped based on age intervals: <40, 40-60, and >60. [Figure 5] Figure 5 shows the mass and retention time index for each ion box. The figure shows the mass error for each metabolite (Figure 3A) and the retention time variability for each mass box / metabolite (Figure 3B). [Figure 6] Figure 6 shows the quality checks for sample validation. These are QC1, QC2, and QC3. QC1 determines the spectral quality of a chromatogram based on its accepted criteria. QC2 was employed to find >5 of the 9 specified critical masses in the spectrum. Meanwhile, QC3 accepted a matrix occupancy of 0.2 in the sample as a correct chromatogram. [Figure 7] FIG. 7 shows a PLS DA plot of the matrix of samples and metabolites versus metabolite intensities, which shows a clear separation of the samples based on their clinical information. [Figure 8]Figure 8 shows the AI ​​workflow: Multiple Cancer Detection Platform / Tissue of Primary Detection. This workflow shows three main compartments common to hierarchical models: from left to right, data processing, train-test split, and model building-test. [Figure 9] Figure 9 shows testing of the trained tier I model on cancer versus normal and disease controls, demonstrating a clear separation of cancer versus controls based on model scores. The y-scores for each of the 33 cancers are shown separately. The confusion matrix obtained when applying a threshold of 0 shows high accuracy, sensitivity, and specificity. [Figure 10] Figure 10 shows the testing of the multi-class trained model Tier 2 model for tissue identification. The resulting confusion matrix is ​​generated after applying the dual-class predictions from Tier 2 model, i.e., tissue identification for cancer-positive samples. [Figure 11] FIG. 11 shows the coefficient / weight of each metabolite involved in the signature of cancer distinguishing from normal controls.

[0034] Table 1: Distribution of 33 cancer and normal control samples based on parameters such as age interval, BMI, ethnicity, and cancer stage. DETAILED DESCRIPTION OF THE INVENTION

[0035] For clarity, specific terms are used in the following description, but these terms are intended to refer only to specific structures of the invention selected for illustration in the drawings, and are not intended to define or limit the scope of the invention.

[0036] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, characteristic, or function described in connection with an embodiment is included in at least one embodiment of the invention. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.

[0037] The present invention discloses embodiments that allow for the simultaneous screening of multiple cancers, such as endometrial cancer, breast cancer, cervical cancer, lung cancer, prostate cancer, and ovarian cancer (but not limited to the names identified herein), in a single analysis. The present invention relates to systems and methods that may integrate comprehensive (global) metabolomic profiling with machine learning-driven data analysis to capture disease-specific signatures.

[0038] In one embodiment, the present invention may provide an integrated method for the simultaneous detection of multiple cancers, which may further elaborate a non-targeted metabolomics process to detect and measure metabolic changes that are useful not only in the broad differentiation between cancers and healthy individuals, but also effectively and simultaneously distinguish each individual cancer from normal controls and other cancers.

[0039] However, although the detailed description herein describes and relates to a number of cancers, which are endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, stomach cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumors, squamous cell carcinoma, and cancer of unknown primary origin, the methods described herein may not be limited to the detection of only these cancers, but may also be applied to the isolation and detection of other cancers in living mammalian specimens from normal controls.

[0040] As described in some of the following examples, a liquid chromatography-mass spectrometry (LC-MS)-based non-targeted metabolomics approach may be used to screen differentially abundant serum metabolites from control cases (normal and disease controls) and test cases (i.e., endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, gastric cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and cancer of unknown primary). In the specific exemplary study conducted and presented in this invention, a total of 8,971 serum samples were collected from participants. Of these, 3,914 were selected as normal controls, and 5,057 were selected as test cases. In the test cases, the distribution of cancer serum samples was as follows: endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, stomach cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and cancer of unknown primary cancer. The metabolite profiles for the 16 cancers in the 16 cancer groups were 445, 652, 458, 488, 307, 157, 169, 151, 296, 136, 97, 134, 147, 279, 20, 52, 566, 143, 122, 32, 42, 18, 9, 4, 5, 18, 2, 35, 45, 14, 8, and 6, respectively, while "other" cancers had a total sample distribution of 16 (Table 1) (shown in Figure 3). In an exemplary study of the present invention, the potential utility of the derived metabolite profiles for discriminating between cases and controls was investigated by constructing and evaluating a multivariate classification matrix.

[0041] Before describing exemplary embodiments in more detail, the following definitions are set forth to illustrate and define the meaning and scope of terms used in the description. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Singleton et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY, 2D ED., John Wiley and Sons, New York (1994), and Hale and Markham, THE HARPER COLLINS DICTIONARY OF BIOLOGY, Harper Perennial, New York (1991), provide those of ordinary skill in the art with the general meaning of many of the terms used herein. Additionally, certain terms are defined below for clarity and ease of reference.

[0042] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a sample" refers to one or more samples, i.e., a single sample and multiple samples. Accordingly, this statement is intended to serve as a predicate (antecedent basis) for using exclusive terminology such as "solely," "only," or a "negative" limitation in connection with recitation of the claim recitation.

[0043] The term "sample," as used herein, refers to a material or mixture of materials, typically, though not necessarily, in liquid form, containing one or more analytes of interest. In one embodiment, the term, when used in its broadest sense, refers to any mammalian material containing cells or producing cellular metabolic products, such as tissues or bodily fluids isolated from an individual (including, but not limited to, plasma, serum, cerebrospinal fluid, lymph, tears, saliva, and tissue sections), or tissues or bodily fluids isolated from in vitro cell culture components, as well as samples from the environment. The term "sample" may also refer to a "biological sample." As used herein, the term "biological sample" refers to a whole organism or a subset of its tissues, cells, or components (e.g., bodily fluids, including, but not limited to, blood, mucus, lymph, synovial fluid, cerebrospinal fluid, saliva, amniotic fluid, umbilical cord blood, urine, vaginal fluid, and semen). A "biological sample" may also refer to a homogenate, lysate, or extract prepared from a whole organism or a subset of its tissues, cells, or components, or a fraction or portion thereof, including, but not limited to, plasma, serum, spinal fluid, lymphatic fluid, external sections of skin, respiratory, intestinal, and genitourinary tracts, tears, saliva, milk, blood cells, tumors, organs. In certain embodiments, the sample is removed from an animal. Biological samples of the present invention comprise cells.

[0044] The metabolite profile used in the present invention should be understood to be any defined set of quantitative metabolite result values ​​that can be used for comparison with reference values ​​or reference profiles derived from another sample or group of samples. For example, the metabolite profile of a sample from a diseased patient may be significantly different from the metabolite profile of a sample from a similarly matched healthy patient. The metabolites may be, but are not limited to, amino acids, peptides, acylcarnitines, monosaccharides, lipids and phospholipids, prostaglandins, steroids, bile acids, and glycols, and phospholipids may be detected and / or quantified.

[0045] As used herein, non-targeted metabolomics studies are characterized by the simultaneous measurement of many metabolites from biological samples. This strategy, known as a top-down strategy, avoids the need for a specific hypothesis about a specific set of metabolites in advance and instead analyzes the entire metabolomic profile. Therefore, these studies are characterized by the generation of large amounts of data. This data is characterized not only by its quantity but also by its complexity, and therefore requires high-performance bioinformatics tools.

[0046] As used herein, the term chromatography refers to a process by which a liquid or gas-borne chemical mixture is separated into components as a result of the differential distribution of chemicals as they flow around or over a stationary liquid or solid phase.

[0047] As used herein, the term high performance liquid chromatography or HPLC (sometimes known as high pressure liquid chromatography) refers to liquid chromatography in which the degree of separation is increased by passing a mobile phase under pressure through a stationary phase, typically a tightly packed column. As used herein, the term ultra performance liquid chromatography or UPLC or UHPLC (sometimes known as ultra high pressure liquid chromatography) refers to HPLC performed at pressures much higher than conventional HPLC techniques.

[0048] As used herein, the term sample injection refers to the introduction of a single sample aliquot into an analytical instrument, such as a mass spectrometer. This introduction may occur directly or indirectly. Indirect sample injection may be achieved, for example, by injecting the sample aliquot into an HPLC or UPLC analytical column connected in an on-line manner to the mass spectrometer.

[0049] As used herein, the term mass spectrometry or MS refers to an analytical technique for identifying compounds by their mass. MS refers to a method of filtering, detecting, and measuring ions based on their mass-to-charge ratio or m / z.

[0050] As used herein, the term operating in positive ion mode refers to a mass spectrometry method in which positive ions are generated and detected.

[0051] As discussed herein, the term electron ionization or EI refers to a method in which an analyte of interest in the gas or vapor phase interacts with a stream of electrons. Collisions between the electrons and the analyte produce analyte ions, which may then be subjected to mass spectrometry techniques.

[0052] As used herein, the term electrospray ionization, or ESI, refers to a method in which a solution is passed along a short length of capillary tube and a high positive or negative potential is applied to the end of the tube. The solution reaching the end of the tube is vaporized (atomized) into a jet or droplets of very small droplets of the solution in solvent vapor. This mist of droplets flows through an evaporation chamber, which is slightly heated to prevent condensation and evaporate the solvent. As the droplets become smaller, the electrical surface charge density increases until natural repulsion between like charges causes the release of ions and neutral molecules.

[0053] As used herein, data processing typically involves a data reduction step called filtering. A noise filter reduces data based on a calculated noise threshold. At this point, data below a certain signal-to-noise ratio is filtered. Content-based filtering of the results leverages, for example, disease-specific knowledge to focus on relevant metabolite aspects of the disease under consideration.

[0054] After the pre-processed data from mass spectrometry has been technically validated, statistical analysis can proceed. Depending on the design of the metabolite profiling study, samples or several samples from healthy controls and patients are compared to identify differences, i.e., biomarkers that can be used to characterize disease at the molecular level. In another embodiment, the samples are from patients participating in clinical trials in which new drug compounds are being investigated and compared with approved drugs.

[0055] At its core, artificial intelligence is a new technological field that studies and develops theories, methods, techniques, and application systems to simulate the extension and expansion of human intelligence. The use of AI in research is likely to perform some complex tasks that require human cognitive abilities. The main core concepts of AI are machine learning and deep learning. However, machine learning is a research technique for algorithms that learn from examples and experiences. Additionally, machine learning is based on the idea that there are patterns in data that can be identified and used to make future predictions. On the other hand, deep learning uses different layers to learn from data. The depth of a model is represented by the number of layers in the model. In deep learning, the learning stage is carried out through a neural network. A neural network is an architecture in which layers are stacked on top of each other.

[0056] Reference is made to FIG. 1 , which shows a schematic diagram of a system for performing a metabolomics process to distinguish cancer types (e.g., endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, stomach cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and cancers of unknown primary and additional cancers classified in the “other” category) from normal controls, further distinguish cancer types within a group of cancers, and also perform QC to improve the accuracy rate of predictions, in accordance with one embodiment of the present invention. FIG. 1 shows a metabolomics system 100 that may include at least one or more components that perform one or more of the following functions: 1. At least one sample collection device 102 for collecting one or more biological fluid samples from one or more living mammals. 2. At least one precipitation device 104 for the extraction of one or more metabolites from one or more biological fluid samples by precipitation of proteins present in the biological fluid with chilled alcohol, including at least methanol. 3. At least one vacuum dryer 106 for drying the one or more metabolite extracts from the at least one precipitation device. 4. At least one device 108 for reconstituting one or more dried metabolite extracts in an aqueous solution. 5. At least one liquid chromatography (LC) device 110 equipped with a mass spectrometer (MS) (hereinafter abbreviated as LC-MS) for analyzing one or more of the resulting reconstituted metabolites. 6. At least one computational device 112 for aligning masses obtained from the metabolomic profile generated using an automated platform, i.e., compound discoverer software, which extracts data about metabolite ions and their associated features. 7. At least one device for minimizing possible errors in measuring the mass of the ions. 8. At least one computing device 114 for subjecting the aligned and normalized ion spectra to three quality controls: a) Faulty chromatogram profiles are identified using a sequential neural network model, which ensures that errors due to either poor sample extraction or errors in mass spectrometry are eliminated. b) Monitoring the presence of at least six of nine critical ions with m / z values ​​in the range of 100 to 800. This confirms a high probability of accurate identification of cancer samples. c) Matrix occupancy determines the percentage of features that match the matrix size, which ensures detection robustness and accuracy rate. 9. At least one computing device 116 that may execute one or more AI / ML algorithms for AI-based pattern recognition to ultimately identify, differentiate, and present cancer samples from normal control samples, and further identify, differentiate, and present individual cancer samples within the identified cancer samples. Thus, not only can detection and differentiation between cancer and healthy individuals be achieved from the present system 100, but the present system 100 may also effectively and simultaneously distinguish each individual cancer from normal controls and other cancer samples.

[0057] It should again be noted that Figures 1-11 are described by way of example and therefore should not be considered limiting to only these specific examples. For example, Figures 1-11 are described herein considering a sample size of 8971 taken from both male and female adult volunteers.

[0058] In one embodiment, the system 100 may be implemented to distinguish multiple cancers from normal controls, and subsequently differentiate between endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, gastric cancer, head and neck cancer, esophageal cancer, and prostate cancer.

[0059] They were either free of any cancer (normal controls) (n=3914) or had endometrial cancer (n=445), breast cancer (n=652), cervical cancer (n=458), ovarian cancer (n=488), lung cancer (n=307), leukemia (n=157), thyroid cancer (n=169), melanoma (n=151), colorectal cancer (n=296), kidney cancer (n=136), lymphoma (n=97), pancreatic cancer (n=134), liver and bile duct cancer (n=147), stomach cancer (n=279), laryngeal cancer (n=20), pharyngeal cancer (n=52), oral cancer (n= Numerous samples were obtained from adult volunteers with cancers including esophageal cancer (n=143), prostate cancer (n=122), bladder cancer (n=32), brain and central nervous system tumors (n=42), multiple myeloma (n=18), anal cancer (n=9), testicular cancer (n=4), vulvar cancer (n=5), penile cancer (n=18), vaginal cancer (n=2), gallbladder cancer (n=35), sarcoma cancer (n=45), germ cell tumor (n=14), squamous cell carcinoma (n=8), cancer of unknown primary (n=6), and other cancers (n=16) (Table 1). The samples were collected and stored in a sample collection device 102. In one embodiment, the sample collection device 102 may be a test tube.

[0060] Additionally, the system 100 may include metabolite extraction, which may be achieved by precipitating serum proteins using chilled methanol, according to one embodiment. Thus, a precipitation device 104 may be used in the present invention to extract metabolites from a collected sample by precipitating serum proteins using chilled methanol. In one embodiment, the precipitation device 104 may be a test tube.

[0061] The supernatant may be collected as a metabolite extract and may be further dried before use. For the drying process, in one embodiment, a phase separator or vacuum dryer 106 may be used, which may use a high speed vacuum to dry the metabolite extract.

[0062] Further, in one embodiment, the dried extract may be reconstituted in an aqueous solution in a mobile phase using a reconstitution device 108. An ion spectrum of the resulting sample from the reconstitution step may then be generated by LCMS, in which the sample may first be resolved by a liquid chromatography (LC) equipped with a mass spectrometry (MS) device (hereafter abbreviated as LCMS) 110. Using device 110, ions in the metabolite extract may be measured, and the masses of the ions may be determined based on their mass-to-charge ratio or m / z (FIG. 1).

[0063] The ion spectral features accumulated in the metabolic profile may then be extracted using compound finder software 112 (e.g., Thermo Fisher Scientific Compound Finder Software). The masses obtained for the ions in the metabolomic profile using the LCMS instrument 110 may be aligned across all samples. This may be done to allow for comparison of the peak intensity of each ion across all samples. For example, a pool of known internal standards may be used for RT alignment with an error window of ±0.02 minutes, followed by peak picking and metabolite identification.

[0064] The system 100 may also include functionality for minimizing errors generated in measuring the mass of ions. To normalize for unavoidable but minor mass (m / z) variations, an advanced approach using a parts-per-million (ppm) error-based approach may be used, according to one embodiment. In another embodiment, a modified virtual lock mass-based approach may also be used. This is based on the principle that mass error is known to increase with mass. This modified virtual lock mass-based approach may be used and adapted according to the data set in the present example. This may be done by combining a traditional virtual lock mass approach with metabolite identification from the Human Metabolome Database (HMDB). Specifically, a virtual lock mass box may be defined using the masses of metabolites identified by HMDB database searches across multiple samples. Metabolite ions may then be filtered based on their frequency of presence in the sample, which may be used for metabolite ion filtering. This means that ions present in more than 15% of the samples may be used in subsequent analyses.

[0065] To improve the overall accuracy rate of predictions using one or more AI / ML algorithms, system 114 was introduced, including steps QC1, QC2, and QC3. System 114 was applied to the aligned and normalized dataset to establish confirmation of samples processed according to the optimized protocol. System 114 identifies any errors that may have occurred during sample processing in any of the steps of system 100. The implementation of system 114 is critical to improving accuracy rates at both the CDAI and TOOAI prediction levels.

[0066] An AI / ML model is then applied to the acquired, measured, aligned, corrected, and characterized metabolite ions measured and aligned as described above for statistical analysis of the sample. The computing device 116 may be capable of executing one or more AI / ML algorithms for applying the AI / ML model for statistical analysis of the sample.

[0067] One or more first AI / ML models may be generated to initially distinguish cancer samples (endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, stomach cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and cancer of unknown primary, and "other") from normal controls by executing one or more AI / ML algorithms using one or more processors in the computing device 116. Then, in another embodiment, one or more additional AI / ML algorithms may be executed using one or more processors in the computing device 116 to further distinguish individual cancers (e.g., lung cancer and the remaining 18 cancers) (Table 2).

[0068] While generating the AI / model, the computing device 116 may follow one or more of the following steps (FIG. 8). i. During the development of the AI ​​model, a functional mapping is established between the dependent / target variable learning and the independent variable learning in the training dataset that can distinguish cancer samples from normal control samples based on the y-score. ii. To overcome class imbalance in the training data, class weights of the objective variable were set in the AI ​​model. iii. Optimization algorithms have been set up in the AI ​​model to handle the complexity of the data and make it faster.

[0069] This may generate a Cancer Detection AI (CDAI) model that can distinguish cancer samples from normal controls.

[0070] Samples identified as cancer-positive by the CDAI algorithm are then subjected to analysis by a second AI model for tissue of origin identification (the TOOAI model) to distinguish individual cancers (e.g., lung cancer from the remaining 18 cancers). The TOOAI model may include either a support vector machine, logistic one-to-many, or stochastic gradient descent algorithm to serve as the classifier model for training the cancer samples.

[0071] Therefore, in one embodiment, a two-step modeling scheme may be applied to the test set. First, a CDAI model may be applied to the test set to distinguish cancer samples from normal samples. Then, a TOOAI model may be applied to the resulting predicted cancer samples to discriminate between 18 individual cancers and a group of cancers called "others." The 18 individual cancers are endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, gastric cancer, head and neck cancer, esophageal cancer, and prostate cancer. The TOOAI model may generate 18 scores for each sample, each score defining the probability that the respective sample belongs to one of the 18 classes.

[0072] In a specific example, the above process as implemented by system 100 is performed. Of the total 8,971 samples, 5,057 samples were from endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, stomach cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and the remaining cancers classified as unknown primary cancer or "other." In addition, there were 3,914 non-cancer controls. This data was randomly split equally into a training dataset and a test dataset. This resulted in 2479 cancer samples and 1957 non-cancer controls in the training set, and 2594 cancer samples and 1957 non-cancer controls in the test set. The CDAI model was applied to the training set (see, e.g., Table 3) and tested on the test set to obtain accuracy, sensitivity, and specificity values. While applying the CDAI model, a PLS DA regression function may be applied to the training dataset to find a function that separates cancer samples from normal control samples (Figure 7).

[0073] Furthermore, to overcome class imbalance in the training data, the class weights of the objective variable were set in the AI ​​model, while an optimization algorithm was set in the AI ​​model to handle data complexity and make it faster. Thus, the CDAI may first be trained using a training dataset of samples. The resulting trained model / algorithm may find a score for each sample. The trained CDAI model may then be evaluated on a test set to determine the accuracy rate. The sensitivity, specificity, and accuracy rate obtained in this example were 99.26%, 99.64%, and 99.8%, respectively.

[0074] In yet another exemplary embodiment, the TOOAI model may be applied to cancer-positive samples determined by the CDAI model. The TOOAI model operates on the predicted cancer samples from the CDAI model, generating a multi-class score for each sample. One score for each cancer type, defined by the tissue of origin of the cancer, indicates the probability that the sample belongs to that cancer type. Here, of the total 8,971 cancer samples, 445 samples were endometrial cancer, 652 were breast cancer, 458 were cervical cancer, 488 were ovarian cancer, 307 were lung cancer, 157 were leukemia, 169 were thyroid cancer, 151 were melanoma, 296 were colorectal cancer, 136 were kidney cancer, 97 were lymphoma, 134 were pancreatic cancer, 147 were liver and bile duct cancer, 279 were gastric cancer, 638 were head and neck cancer, 143 were esophageal cancer, 122 were prostate cancer, and 254 were in the "other" category (Table 2). The data were randomly split into training and test datasets in equal proportions.This resulted in 222 endometrial cancers, 326 breast cancers, 229 cervical cancers, 244 ovarian cancers, 153 lung cancers, 78 leukemias, 84 thyroid cancers, 75 melanoma cancers, 148 colorectal cancers, 68 kidney cancers, 48 ​​lymphomas, 67 pancreatic cancers, 73 liver and bile duct cancers, 139 gastric cancers, 0 laryngeal cancers, 26 pharyngeal cancers, 283 oral cancers, and 71 thyroid cancers in the training set. The following cases were obtained: esophageal cancer, 61 prostate cancer, 16 bladder cancer, 21 brain and central nervous system tumors, 0 multiple myeloma, 0 anal cancer, 0 testicular cancer, 0 vulvar cancer, 0 penile cancer, 0 vaginal cancer, 17 gallbladder cancer, 22 sarcoma cancer, 0 germ cell tumor, 0 squamous cell carcinoma, 0 cancer of unknown primary, 1957 normal controls, and 8 samples in the "other" category. In the study, there were 223 cases of endometrial cancer, 326 cases of breast cancer, 229 cases of cervical cancer, 244 cases of ovarian cancer, 154 cases of lung cancer, 79 cases of leukemia, 85 cases of thyroid cancer, 76 cases of melanoma, 148 cases of colorectal cancer, 68 cases of kidney cancer, 49 cases of lymphoma, 67 cases of pancreatic cancer, 74 cases of liver and bile duct cancer, 140 cases of stomach cancer, 20 cases of laryngeal cancer, 26 cases of pharyngeal cancer, 283 cases of oral cancer, 72 cases of esophageal cancer, and 61 cases of pancreatic cancer. The samples yielded cases of prostate cancer, 16 bladder cancer, 21 brain and central nervous system tumors, 18 multiple myeloma, 9 anal cancer, 4 testicular cancer, 5 vulvar cancer, 18 penile cancer, 2 vaginal cancer, 18 gallbladder cancer, 23 sarcoma cancers, 14 germ cell tumors, 8 squamous cell carcinoma, 6 cancers of unknown primary, 1957 normal controls, and 8 "other" category samples (Table 3) (shown in Figure 3). Support vector machine, logistic one-to-many, and stochastic gradient descent algorithms were then used as classifier models on the training samples to obtain the TOOAI model. A two-step modeling scheme (CDAI model followed by TOOAI model) was then applied to the test set. That is, the CDAI model first distinguished cancer from non-cancer samples in the test set. The TOOAI model was then applied to the resulting predicted cancer samples. This results in 18 scores for each sample, each score defining the probability that the respective sample belongs to one of the 18 classes (Table 2).

[0075] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing endometrial cancer from the remaining cancers within the 18-cancer group. Endometrial cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given endometrial cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's double-class accuracy. The endometrial cancer tissue identification accuracy rate was calculated to be 92.6% (see, e.g., Figure 10, Table 4).

[0076] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing breast cancer from the remaining cancers within the 18-cancer group. Breast cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given breast cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The breast cancer tissue identification accuracy rate was calculated to be 93% (see, e.g., Figure 10, Table 4).

[0077] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing cervical cancer from the remaining cancers within the 18-cancer group. Cervical cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 99.6%. The TOOAI model provides a probability score for each cancer subclass for any given cervical cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The cervical cancer tissue identification accuracy rate was calculated to be 96.6% (see, e.g., Figure 10, Table 4). In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing ovarian cancer from the remaining cancers within the 18-cancer group. Ovarian cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given ovarian cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for ovarian cancer tissue identification was calculated to be 91% (see, e.g., Figure 10, Table 4).

[0078] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing lung cancer from the remaining cancers within the 18-cancer group. Lung cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given lung cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The lung cancer tissue identification accuracy rate was calculated to be 93% (see, e.g., Figure 10, Table 4).

[0079] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing leukemia from the remaining cancers within the 18-cancer group. Leukemia cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given leukemia sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for identifying leukemia tissues was calculated to be 83.3% (see, e.g., Figure 10, Table 4).

[0080] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing thyroid cancer from the remaining cancers within the 18-cancer group. Thyroid cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given thyroid cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for thyroid cancer tissue identification was calculated to be 87.5% (see, e.g., Figure 10, Table 4).

[0081] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing melanoma from the remaining cancers within the 18-cancer group. Melanoma cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given melanoma sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for melanoma tissue identification was calculated to be 92.8% (see, e.g., Figure 10, Table 4).

[0082] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing colorectal cancer from the remaining cancers within the 18-cancer group. Colorectal cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given colorectal cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for colorectal cancer tissue identification was calculated to be 92.5% (see, e.g., Figure 10, Table 4).

[0083] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing kidney cancer from the remaining cancers within the 18-cancer group. Kidney cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given kidney cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The kidney cancer tissue identification accuracy rate was calculated to be 86% (see, e.g., Figure 10, Table 4).

[0084] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing lymphoma from the remaining cancers within the 18-cancer group. Lymphoma cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given lymphoma sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The lymphoma tissue identification accuracy rate was calculated to be 89% (see, e.g., Figure 10, Table 4). In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing pancreatic cancer from the remaining cancers within the 18-cancer group. Pancreatic cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given pancreatic cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for pancreatic cancer tissue identification was calculated to be 100% (see, e.g., Figure 10, Table 4).

[0085] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing liver cancer from the remaining cancers within the 18-cancer group. Liver cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given liver cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for liver cancer was calculated to be 80% (see, e.g., Figure 10, Table 4).

[0086] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing gastric cancer from the remaining cancers within the 18-cancer group. Gastric cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given gastric cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for gastric cancer tissue identification was calculated to be 82% (see, e.g., Figure 10, Table 4).

[0087] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing head and neck cancer from the remaining cancers within the 18-cancer group. Head and neck cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given head and neck cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for head and neck cancer tissue identification was calculated to be 94% (see, e.g., Figure 10, Table 4).

[0088] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing esophageal cancer from the remaining cancers within the 18-cancer group. Head and neck cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given head and neck cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for head and neck cancer tissue identification was calculated to be 87% (see, e.g., Figure 10, Table 4).

[0089] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing prostate cancer from the remaining cancers within the 18-cancer group. Head and neck cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 100%. The TOOAI model provides a probability score for each cancer subclass for any given head and neck cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for head and neck cancer tissue identification was calculated to be 87.5% (see, e.g., Figure 10, Table 4).

[0090] In some embodiments, the system 100 may be further implemented to determine the accuracy rate of the TOOAI model in specifically distinguishing "other" cancers from the remaining cancers within the 18-cancer group. Head and neck cases were initially distinguished from normal control samples with a specificity of 99.64% and a sensitivity of 99.09%. The TOOAI model provides a probability score for each cancer subclass for any given head and neck cancer sample. This score can distinguish the cancer subclass of the sample. The top two scoring cancer subclasses were taken as the model results and matched with the true labels. A final confusion matrix was constructed based on the model's dual-class accuracy rate. The accuracy rate for head and neck cancer tissue identification was calculated to be 94% (see, e.g., Figure 10, Table 4).

[0091] Reference is now made to FIG. 2, which illustrates a flowchart for implementing a metabolomics process for distinguishing cancer samples (e.g., endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, stomach cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and cancer of unknown primary, as well as additional cancers classified in the "other" category) from normal controls in accordance with one embodiment of the present invention, and further identifying each specific cancer type in each sample from those belonging to other cancer types. FIG. 2 should be read and understood in conjunction with FIGS. 1 and 3-11, and may also include at least one or more embodiments of FIGS. 1 and 3-11 without departing from the meaning and scope of the present invention.

[0092] Additionally, the method 200 may include at least one or more of steps 202-218, individually or in combination.

[0093] Additionally, the method 200 is illustrated by taking several examples of cancers, including endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, gastric cancer, head and neck cancer, esophageal cancer, prostate cancer, and additional cancers that fall under the category of "other," which should not be construed as limiting the meaning and scope of the present invention.

[0094] Figure 2 shows the results of the 2014-nCoV study in which patients were either free of any cancer (normal controls) (n=3914) or had the following cancers: endometrial cancer (n=445), breast cancer (n=652), cervical cancer (n=458), ovarian cancer (n=488), lung cancer (n=307), leukemia (n=157), thyroid cancer (n=169), melanoma (n=151), colorectal cancer (n=296), kidney cancer (n=136), lymphoma (n=97), pancreatic cancer (n=134), liver and bile duct cancer (n=147), stomach cancer (n=279), laryngeal cancer (n=20), pharyngeal cancer (n=52), oral cancer (n=566), esophageal cancer (n=143), 1 illustrates a metabolomics process 200 that may include step 202 for collecting and storing multiple samples from male and female volunteers with prostate cancer (n=122), bladder (n=32), brain and central nervous system tumors (n=42), multiple myeloma (n=18), anal cancer (n=9), testicular cancer (n=4), vulvar cancer (n=5), penile cancer (n=18), vaginal cancer (n=2), gallbladder cancer (n=35), sarcoma cancer (n=45), germ cell tumor (n=14), squamous cell carcinoma (n=8), cancer of unknown primary (n=6), and other cancers (n=16) (Table 1). The samples are collected and stored in a sample collection device 102.

[0095] The method further includes step 204 of extracting a metabolite extract, which may be achieved by precipitating serum proteins using chilled methanol. In one embodiment, the precipitation device 104 may be a test tube. The supernatant may be collected as the metabolite extract. Then, in step 206, the metabolite extract may be dried before use. For the drying process, in one embodiment, a phase separation device 106 may be used, which may dry the metabolite extract using a high-speed vacuum. In step 208, in one embodiment, the dried extract may be reconstituted in an aqueous solution in a mobile phase using the device 108. Then, in step 210, analysis of the resulting sample from the reconstitution step may be performed by an LCMS 110. In step 210, the reconstituted sample may first be resolved by a liquid chromatography (abbreviated as LC) device 110, and then an ion spectrum may be subsequently acquired through a high-resolution mass spectrometer (abbreviated as MS). Using the LCMS device 110, ions in the metabolite extract may be measured, and the masses of the ions may be measured based on their mass-to-charge ratio or m / z.

[0096] The ion spectral features stored in the metabolic profile may then be extracted using a computing device 112, which may use one or more processors to run compound discoverer software. Additionally, in one embodiment, the method 200 may include aligning 212 the masses obtained for the ions in the metabolomic profile across all samples using LCMS. This may be done to allow for comparison of the peak intensities of each ion across all samples.

[0097] Furthermore, in one embodiment, an additional optional step 214 is included to minimize possible errors in measuring the mass of ions. To normalize for unavoidable but minor mass (m / z) variations, an advanced approach using a parts per million (ppm) error-based approach may be used, according to one embodiment. In another embodiment, a simplified modified virtual lock mass-based approach may also be used. This is based on the principle that mass error is known to increase with mass. This modified virtual lock mass-based approach may be used and adapted according to the data set in the present example. This may be done by combining a traditional virtual lock mass approach with metabolite identification from the Human Metabolome Database (HMDB). Specifically, a virtual lock mass box may be defined using the masses of metabolites identified by a HMDB database search across multiple samples (FIG. 5).

[0098] Step 216 is then applied to the aligned and normalized ion spectra to further improve the overall accuracy rate of the algorithm used for prediction. Step 216 includes three quality checks (QCs), which are described as follows:

[0099] Step QC1 of the system 114 involves chromatogram profile matching. Defective chromatograms may be the result of faulty sample extraction or errors in the mass spectrometry setup. These errors affect the quality of the data and subsequently impair the overall prediction accuracy of the algorithm. To detect these defects based on the chromatogram profile variability, a sequential neural network model was constructed. The chromatograms obtained from the mass spectrometer for each sample were first converted to JPEG format. The images were then scaled to an appropriate width and length for model training. The images were binarized to separate the chromatograms, facilitating more efficient analysis. The Keras sequential neural network model was then used to train, test, and validate the model. We employed three Keras Conv2D two-dimensional convolutional layers. These generate convolution kernels wrapped around the layered input, which serve to generate the output tensor. An 80-20 training-testing split was employed. The Adam optimizer was then used to iteratively update the network weights based on the training data. The threshold used for correct detection was 0.5, and samples that showed a quality control (QC) score <0.5 were designated as having passed QC1. Samples with a QC score >0.5 were rejected as having failed the QC1 step. We obtained 100% detection accuracy for defective chromatograms, as shown in Figure 6A.

[0100] Step QC2 of system 114 monitors for the presence of critical m / z ions. The second step of system 114, called QC2, monitored for the presence of at least six of nine critical ions with m / z values ​​ranging from 100 to 800. The intensity and RT distributions of these nine ions are shown in Figure 6B. The presence of six or more of the nine critical ions in a sample chromatogram is the threshold criterion for passing the QC2 step. Samples with fewer than six of the nine critical masses are rejected as having failed the QC2 step. This QC2 step is important for the accurate identification of cancer samples because samples that do not pass this QC2 step are likely to be misclassified by the CDAI algorithm (Figure 6B).

[0101] Step QC3 of system 114 monitors matrix occupancy. Another layer of quality check, introduced alongside the two quality checks above, included assessing matrix occupancy. This layer, step QC3, relies on the percentage of features that match the matrix size. The threshold was optimized through multiple runs of low-quality samples or inappropriate mass spectrometry run conditions. Based on these studies, the minimum matrix occupancy threshold was set at 15%. This threshold was confirmed through multiple validation runs and subsequently found to improve the robustness and accuracy of the predictive CDAI algorithm. As shown in Figure 6C, using the trained model for QC3 resulted in 100% accuracy in capturing defective samples.

[0102] Additionally, in one embodiment, method 200 may include step 218, which may use one or more AI / ML algorithms for AI-based pattern recognition to ultimately identify, differentiate, and represent cancer samples from normal control samples, and further to identify, differentiate, and represent individual cancer samples within the identified cancer samples.

[0103] Method 200 may further include step 218 of applying an AI / ML model / algorithm to the acquired, measured (e.g., aligned and corrected), and characterized metabolite ions measured and aligned as described above. Step 218 may include applying the AI / ML model for statistical analysis of the sample. Computing device 116 may be capable of executing one or more AI / ML algorithms for applying the AI / ML model for statistical analysis of the sample.

[0104] Applying 218 the AI / ML models / algorithms may include creating and applying at least two AI models: first a CDAI model, followed by a TOOAI model. By executing one or more AI / ML algorithms using one or more processors in the computing device 116, one or more first AI / ML models may be generated to distinguish cancer samples (endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, stomach cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and cancer of unknown primary and additional cancers classified as "other") from normal controls. Additionally, in another embodiment, the computing device 116 may use one or more processors to execute one or more AI / ML algorithms, followed by a second AI / ML model to further discriminate and identify the cancer type as defined by tissue of origin (e.g., colorectal cancer from the remaining cancer types) (Table 2).

[0105] Step 218 may optionally be included in method 200. Furthermore, the sequence of steps 202-218 may be varied and need not be limited to that shown in method 200.

[0106] While generating the AI / model in step 218, it may follow one or more of the following steps. i. During the development of the AI ​​model, a functional mapping is established between the dependent / target variable learning and the independent variable learning in the training dataset that can distinguish cancer samples from normal control samples based on the y-score. ii. To overcome class imbalance in the training data, class weights of the objective variable were set in the AI ​​model. iii. Optimization algorithms have been set up in the AI ​​model to handle the complexity of the data and make it faster.

[0107] This may generate a CDAI model in step 218 that may separate normal control samples from cancer samples. Another AI model, called a TOOAI model, may be generated in step 218 and applied to the cancer-positive samples identified by the CDAI model to distinguish individual cancer types (e.g., colorectal cancer) from the remaining cancer types (e.g., endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, gastric cancer, head and neck cancer, esophageal cancer, prostate cancer, and cancers in the "other" group) from the normal controls (Table 2). In one embodiment, the TOOAI model may be generated in a similar manner to the CDAI and may further include a support vector machine, logistic one-to-many, or stochastic gradient descent algorithm classifier that acts as a classification model that may be created using the training samples to provide a second TOOAI (FIG. 8).

[0108] Therefore, in one embodiment, a two-step modeling scheme may be applied to the test set. That is, first, a CDAI model for distinguishing cancer samples from normal samples may be applied to the test set. Then, TOOAI may be applied to the resulting predicted cancer samples. Here, if 18 types of cancer are taken, such as endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, colorectal cancer, gastric cancer, head and neck cancer, esophageal cancer, prostate cancer, and an additional cancer classified as "other," this two-step modeling scheme may result in 18 scores for each sample, each score defining the probability that the respective sample belongs to one of the 18 classes.

[0109] The present invention will now be described by way of examples. The examples are for illustrative purposes only and should not be construed as limiting. The following examples will be described in detail while implementing the systems and methods shown in Figures 1 and 2, respectively. [Example]

[0110] We described an approach for the early differentiation of multiple cancers in both men and women from control cases using liquid chromatography-mass spectrometry for untargeted metabolomics of serum. Samples were collected from subjects with these cancers, i.e., those without any cancer (normal controls) (n=3914) or from those with endometrial cancer (n=445), breast cancer (n=652), cervical cancer (n=458), ovarian cancer (n=488), lung cancer (n=307), leukemia (n=157), thyroid cancer (n=169), melanoma (n=151), colorectal cancer (n=296), kidney cancer (n=136), lymphoma (n=97), pancreatic cancer (n=134), liver and bile duct cancer (n=147), gastric cancer (n=279), laryngeal cancer (n=20), pharyngeal cancer (n=10), and rectal cancer (n=10). Samples were collected from male and female volunteers with cancers of the following cancer types: 52), oral cavity (n = 566), esophageal (n = 143), prostate (n = 122), bladder (n = 32), brain and central nervous system (n = 42), multiple myeloma (n = 18), anal (n = 9), testicular (n = 4), vulvar (n = 5), penile (n = 18), vaginal (n = 2), gallbladder (n = 35), sarcoma (n = 45), germ cell (n = 14), squamous cell (n = 8), cancer of unknown primary (n = 6), and other cancers (n = 16) (Table 1). A total of 3,914 normal control samples were collected.

[0111] [Table 1(1)] [Table 1(2)] [Table 1(3)]

[0112] The untargeted metabolomics approach (see, for example, Figure 2) generated a large list of metabolites in female cases and correlated them with 17 subsets of normal controls, endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, leukemia, thyroid cancer, melanoma, colorectal cancer, kidney cancer, lymphoma, pancreatic cancer, liver and bile duct cancer, gastric cancer, laryngeal cancer, pharyngeal cancer, oral cancer, esophageal cancer, prostate cancer, bladder cancer, brain and central nervous system tumors, multiple myeloma, anal cancer, testicular cancer, vulvar cancer, penile cancer, vaginal cancer, gallbladder cancer, sarcoma cancer, germ cell tumor, squamous cell carcinoma, and unknown primary cancer. The database was further divided into subclasses: 04, 1821, 1766, 1762, 1846, 1481, 1725, 1605, 1780, 1578, 1613, 1655, 1826, 1770, 1164, 1408, 1845, 1940, 2095, 1968, 2016, 1973, 1933, 2025, 1954, 1948, 2014, 1877, 1959, 2027, 1911, and 1903, while "other" cancers each had a total of 1915 metabolite unions across all samples in the subclasses (see, e.g., Figure 4). The total number of unique metabolites identified in this study was 2709. Plant and drug metabolites were removed from this database.

[0113] The data was then passed through our data processing pipeline (see, e.g., Figure 8). Briefly, here, samples were first aligned using a combination of the VLM approach and the identified metabolites to generate a matrix of 8971 samples and 2709 metabolites with corresponding intensity information. The intensity values ​​were transformed to a log10 scale. In one embodiment, metabolite ion filtering was then performed to remove metabolites with weights below a threshold obtained from the PLS-DA regression mapping of cancer versus control samples. In one embodiment, data normalization and missing value imputation were then performed on the data. This resulted in a matrix of a total of 2709 metabolites across the 8971 samples. Of the 5057 samples, 445 samples were from endometrial cancer, 652 from breast cancer, 458 from cervical cancer, 488 from ovarian cancer, 307 from lung cancer, 157 from leukemia, 169 from thyroid cancer, 151 from melanoma, 296 from colorectal cancer, 136 from kidney cancer, 97 from lymphoma, 134 from pancreatic cancer, 147 from liver and bile duct cancer, 279 from gastric cancer, 638 from head and neck cancer, 143 from esophageal cancer, 122 from prostate cancer, and 3914 were normal control samples (Table 2).

[0114] The matrix generated above was used to determine whether there were any differences between these samples based on their metabolite profiles. A PLS-DA plot was created using the matrix shown in Figure 7. Figure 7 clearly shows that each cancer could be distinguished from healthy samples based on their metabolic data in both male and female cases. To quantify how well they could be distinguished, AI analysis (see, e.g., Figures 1-2, 9, 10, and Table 4) was performed on the data as described below to find common patterns in metabolite variations within cancer samples that differed from control samples. Furthermore, a classification model was constructed for the detected metabolite ions using a random distribution of samples into test and training sets (see, e.g., Figures 1-2, 9, 10, and Table 4). The first such model (CDAI model) was constructed to distinguish between cancer samples and normal control samples. For this run, 5,057 cancer samples and 3,914 normal control cases were considered. Furthermore, these study samples (n=8971) were randomly split (50%) into a training set and a test set (see, e.g., Table 3). A multivariate classifier was derived on the training set and evaluated on the test set, generating a confusion matrix with predicted and true labels. This ultimately led to the discrimination of cancer samples from controls with a sensitivity, specificity, and accuracy of 100%, 99.64%, and 100%, respectively (see, e.g., Figure 9).

[0115] Additionally, a multiclass classifier was constructed to distinguish cancers from each other. Here, a model (TOOAI model) was constructed using a total of 445 endometrial cancer samples, 652 breast cancer samples, 458 cervical cancer samples, 488 ovarian cancer samples, 307 lung cancer samples, 157 leukemia samples, 169 thyroid cancer samples, 151 melanoma samples, 296 colorectal cancer samples, 136 renal cancer samples, 97 lymphoma samples, 134 pancreatic cancer samples, 147 liver and bile duct cancer samples, 279 gastric cancer samples, 638 head and neck cancer samples, 143 esophageal cancer samples, 122 prostate cancer samples, and 3914 normal control samples (Table 2). These study samples were randomly divided (50%) into a training set and a test set. This resulted in 222 endometrial cancers, 326 breast cancers, 229 cervical cancers, 244 ovarian cancers, 153 lung cancers, 78 leukemias, 84 thyroid cancers, 75 melanomas, 148 colorectal cancers, 68 kidney cancers, 48 ​​lymphomas, 67 pancreatic cancers, 73 liver and bile duct cancers, 139 gastric cancers, 0 laryngeal cancers, 26 pharyngeal cancers, 283 oral cancers, and 71 The following cases were obtained: 10 cases of esophageal cancer, 61 cases of prostate cancer, 16 cases of bladder cancer, 21 cases of brain and central nervous system tumors, 0 cases of multiple myeloma, 0 cases of anal cancer, 0 cases of testicular cancer, 0 cases of vulvar cancer, 0 cases of penile cancer, 0 cases of vaginal cancer, 17 cases of gallbladder cancer, 22 cases of sarcoma cancer, 0 cases of germ cell tumor, 0 cases of squamous cell carcinoma, 0 cases of cancer of unknown primary, 1957 normal controls, and 8 cases of "other" category samples. In the set, there were 223 endometrial cancers, 326 breast cancers, 229 cervical cancers, 244 ovarian cancers, 154 lung cancers, 79 leukemias, 85 thyroid cancers, 76 melanoma cancers, 148 colorectal cancers, 68 kidney cancers, 49 lymphomas, 67 pancreatic cancers, 74 liver and bile duct cancers, 140 gastric cancers, 20 laryngeal cancers, 26 pharyngeal cancers, 283 oral cancers, and 72 esophageal cancers. This resulted in 61 cases of prostate cancer, 16 bladder cancer, 21 brain and central nervous system tumors, 18 multiple myeloma, 9 anal cancer, 4 testicular cancer, 5 vulvar cancer, 18 penile cancer, 2 vaginal cancer, 18 gallbladder cancer, 23 sarcoma cancer, 14 germ cell tumor, 8 squamous cell carcinoma, 6 cancers of unknown primary, 1957 normal controls, and 8 samples in the "other" category (Table 3).A set of 1,957 normal samples was retained as a test set to test the accuracy of first applying the cancer-versus-normal model and then applying the TOOAI model to discriminate between multiple cancers. Multivariate classifiers were derived on the training set and evaluated on the test set. The TOOAI model assigned each sample 18 scores corresponding to endometrial cancer score, breast cancer score, cervical cancer score, ovarian cancer score, lung cancer score, leukemia cancer score, thyroid cancer score, melanoma cancer score, colorectal cancer score, kidney cancer score, lymphoma cancer score, pancreatic cancer score, liver and bile duct cancer score, gastric cancer score, head and neck cancer score, esophageal cancer, prostate cancer score, and "other" cancer score.

[0116] To test the accuracy of the dual-class prediction obtained from the TOOAI model for endometrial cancer, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for endometrial cancer, which ultimately resulted in a 93.3% discrimination between endometrial cancer candidates and others (see, for example, Figure 10, Table 4).

[0117] To test the accuracy of the dual-class prediction obtained from the TOOAI model for breast cancer, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for breast cancer, which ultimately resulted in a 94.1% discrimination between breast cancer candidates and others (see, e.g., Figure 10, Table 4).

[0118] To test the accuracy of the dual-class prediction obtained from the TOOAI model for cervical cancer, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for cervical cancer, which ultimately resulted in a 96% discrimination between cervical cancer candidates and others (see, for example, Figure 10 and Table 4).

[0119] To test the accuracy of the dual-class prediction obtained from the TOOAI model for ovarian cancer, a confusion matrix with predicted and true labels was generated based on the predicted probability scores for ovarian cancer, which ultimately resulted in a 90% discrimination between ovarian cancer candidates and others (see, e.g., Figure 10, Table 4).

[0120] To test the accuracy of the dual-class prediction obtained from the TOOAI model for lung cancer, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for lung cancer, which ultimately resulted in a 95.4% discrimination between lung cancer candidates and others (see, for example, Figure 10, Table 4).

[0121] To test the accuracy of the dual-class prediction obtained from the TOOAI model for leukemia, we generated a confusion matrix using predicted and true labels based on the predicted probability scores of leukemia cancer, which ultimately resulted in a 91.1% discrimination between leukemia candidates and others (see, for example, Figure 10 and Table 4).

[0122] To test the accuracy of the dual-class prediction obtained from the TOOAI model for thyroid cancer, we generated a confusion matrix with predicted and true labels based on the predicted probability scores for thyroid cancer, which ultimately resulted in 90% discrimination between thyroid cancer candidates and others (see, e.g., Figure 10, Table 4).

[0123] To test the accuracy of the dual-class prediction obtained from the TOOAI model for melanoma, we generated a confusion matrix with predicted and true labels based on the predicted probability scores for melanoma, which ultimately resulted in a 94.7% discrimination between melanoma cancer candidates and others (see, e.g., Figure 10, Table 4).

[0124] To test the accuracy of the dual-class prediction obtained from the TOOAI model for colorectal cancer, a confusion matrix with predicted and true labels was generated based on the predicted probability scores for colorectal cancer, which ultimately resulted in a 95.2% discrimination between colorectal cancer candidates and others (see, e.g., Figure 10, Table 4).

[0125] To test the accuracy of the dual-class prediction obtained from the TOOAI model for kidney cancer, a confusion matrix with predicted and true labels was generated based on the predicted probability scores for kidney cancer, which ultimately resulted in an 86% discrimination between kidney cancer candidates and others (see, for example, Figure 10, Table 4).

[0126] To test the accuracy of the dual-class prediction obtained from the TOOAI model for lymphoma, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for lymphoma cancer, which ultimately resulted in 89.7% discrimination between non-Hodgkin's lymphoma candidates and others (see, e.g., Figure 10, Table 4).

[0127] To test the accuracy of the dual-class prediction obtained from the TOOAI model for pancreatic cancer, we generated a confusion matrix with predicted and true labels based on the predicted probability scores for endometrial cancer, which ultimately resulted in 98.5% discrimination between pancreatic cancer candidates and others (see, e.g., Figure 10, Table 4).

[0128] To test the accuracy of the dual-class prediction obtained from the TOOAI model for liver cancer, a confusion matrix using predicted and true labels was generated based on the predicted probability scores for liver cancer, which ultimately resulted in a 91.89% discrimination between liver cancer candidates and others (see, for example, Figure 10 and Table 4).

[0129] To test the accuracy of the dual-class prediction obtained from the TOOAI model for gastric cancer, a confusion matrix using predicted and true labels was generated based on the predicted probability scores for gastric cancer, which ultimately resulted in a 92.85% discrimination between gastric cancer candidates and others (see, for example, Figure 10, Table 4).

[0130] To test the accuracy of the dual-class prediction obtained from the TOOAI model for head and neck cancer, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for endometrial cancer, which ultimately resulted in a 94.69% discrimination between head and neck cancer candidates and others (see, for example, Figure 10, Table 4).

[0131] To test the accuracy of the dual-class prediction obtained from the TOOAI model for esophageal cancer, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for esophageal cancer, which ultimately resulted in a 95.83% discrimination between esophageal cancer candidates and others (see, for example, Figure 10 and Table 4).

[0132] To test the accuracy of the dual-class prediction obtained from the TOOAI model for prostate cancer, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for prostate cancer, which ultimately resulted in a 98.3% discrimination between prostate cancer candidates and others (see, e.g., Figure 10, Table 4).

[0133] To test the accuracy of the dual-class predictions obtained from the TOOAI model for "other" cancers, we generated a confusion matrix using predicted and true labels based on the predicted probability scores for "other" cancers, which ultimately resulted in a 94.4% accuracy in distinguishing "other" cancer candidates from others (see, e.g., Figure 10, Table 4).

[0134] Thus, as described above, system 100 and associated method 200 may efficiently detect and discriminate cancer samples from normal controls using a first CDAI model, and may further efficiently detect and discriminate each individual cancer sample from other cancer samples by using a TOOAI model on samples identified as cancer positive by the CDAI model.

[0135] The following is a description of exemplary processes and equipment that may be used in the system 100 for performing metabolomics processes and that was used in the studies conducted.

[0136] Subjects and methods Serum samples were obtained from biobanks in the US and Europe or collected from various clinical sites / hospitals in India. The demographic and ethnic distribution of the specimens is shown in Table 1. Controls and disease cases were classified according to age group, BMI, ethnicity, and cancer stage. All diagnoses were performed according to uniform histological and pathological guidelines.

[0137] Serum samples Blood samples were collected and processed according to a standardized protocol. Each sample was assigned a unique laboratory identification number that identified laboratory personnel blinded to the processing sequence and sample origin. Samples were stored at -80°C until use.

[0138] Sample preparation Metabolite extraction from serum was performed as previously described. Briefly, all serum samples were thawed on ice and mixed appropriately. 10 μl of each serum sample was placed in a microcentrifuge tube (1.5 ml) (Genaxy, catalog number GEN-MT-150-CS), and then the sample was added with 30 μl of chilled methanol (Merck, catalog number 1.06018.1000), vortexed briefly, and then kept at -20°C for 60 minutes.

[0139] The samples were then centrifuged at 10,000 rpm for 10 minutes (Sorvall Legend Micro17, Thermo Fisher Scientific, catalog number Legend Micro17). After centrifugation, 27 μl of the supernatant was collected into a separate microfuge tube without disturbing the pellet and dried for 30–35 minutes at low energy using a Speed ​​Vacuum (ThermoFisher Scientific, catalog number SPD1030-230). The sample pellet was then resuspended using 50 μl of a 1:1 mixture of methanol:water for injection. Alternatively, the sample can be stored at −20°C without resuspension. or, Ten microliters of each serum sample was placed in a 1.5 ml microcentrifuge tube (Genaxy, catalog number GEN-MT-150-CS), and then 30 μl of chilled methanol (Merck, catalog number 1.06018.1000) was added to the sample, which was then vortexed briefly and maintained at -80°C for 15 minutes. The sample was then centrifuged at 10,000 rpm for 10 minutes (Sorvall Legend Micro17, Thermo Fisher Scientific, catalog number Legend Micro17). After centrifugation, 27 μl of the supernatant was collected in a separate microcentrifuge tube without disturbing the pellet and dried at low energy for 30–35 minutes using a Speed ​​Vacuum (ThermoFisher Scientific, catalog number SPD1030-230). The sample pellet was then resuspended in 50 μl of a 1:1 mixture of methanol:water for injection. Alternatively, the sample can be stored at -20°C without resuspension. or, Twenty microliters of each serum sample was placed in a 1.5 ml microcentrifuge tube (Genaxy, catalog number GEN-MT-150-CS). 40 μl of chilled methanol (Merck, catalog number 1.06018.1000) was then added to the sample, which was then briefly vortexed. 200 μl of MTBE (methyl tertiary butyl ether) (catalog number 306975-1L) was added to the sample tube, which was then placed on a shaker at room temperature for 1 hour. 50 μl of water was added, followed by brief vortexing. The sample was then centrifuged at 3000 g for 10 minutes (Sorvall Legend Micro17, Thermo Fisher Scientific, catalog number Legend Micro17). After centrifugation, organic and aqueous phases were formed, and each phase was carefully collected into a separate microcentrifuge tube without disturbing the pellet or interface. 25 μl of each phase was collected, added to a fresh microcentrifuge tube, and dried for 25–30 minutes at low energy using a speed vacuum (ThermoFisher Scientific, catalog number SPD1030-230). The sample pellet was then resuspended using 50 μl of a 1:1 methanol:water (water:methanol) mixture for injection. Alternatively, the sample can be stored at -20°C without resuspension. or Ten microliters of each serum sample was placed in a 1.5 ml microcentrifuge tube (Genaxy, catalog number GEN-MT-150-CS), and 400 μl of chilled chloroform / methanol / water (1:3:1) (Merck, catalog number 1.06018.1000; Merck, catalog number C2432-1L; Merck, catalog number 1.15333.1000) was added to the sample and vortexed for 1 minute. The sample was then centrifuged at 13,000 g for 3 minutes (Sorvall Legend Micro17, Thermo Fisher Scientific, catalog number Legend Micro17). After centrifugation, 80 μl of the supernatant was collected into a separate microcentrifuge tube without disturbing the pellet and dried for 25–30 minutes at low energy using a high-speed vacuum (ThermoFisher Scientific, catalog number SPD1030-230). The sample pellet was then resuspended using 50 μl of a 1:1 water:methanol mixture for injection. Alternatively, the sample can be stored at −20° C. without resuspending it.

[0140] LC-MS / MS analysis Untargeted LC-MS / MS metabolomics experiments were performed using a Dionex LC system (Ultimate 3000) online coupled with a QExactive Plus (Thermo Scientific). Each extracted metabolite sample was injected (10 μl for positive ESI ionization) onto a Waters Acquity UPLC HSS T3 (1.8 microns (1.8 μm), dimensions - 2.1 × 100 mm, part number 186003539) heated to 40 °C. The flow rate was 0.3 ml / min. Mobile phase A was (water + 0.1% formic acid), and mobile phase B was (methanol + 0.1% formic acid). The mobile phase was kept isocratic at 5% B for 1 min, increased to 95% B in 7 min, maintained at 95% B for an additional 2 min, and the mobile phase composition was returned to 5% B in 14 min. The ESI voltage was 4 kV. The mass accuracy of the QExactive mass spectrometer was less than 5 ppm and was calibrated according to the recommended schedule before each batch run. The mass scan range was 66.7–1000 Da, and the resolution was set to 35,000. The maximum injection time for the Orbitrap was 100 ms, and the AGC target was optimized at 1e6.

[0141] Optimization and validation of liquid chromatography and mass spectrometry methods To obtain reliable and consistent results of serum metabolite profiles from mass spectrometry, we optimized several parameters to counter imperfect data recording. Among the many steps taken into consideration, our main focus was on matching chromatogram profiles and the quality of data obtained each time a sample was run. We refer to these steps as quality checks (QCs). We designated three main QCs, detailed descriptions of which are provided below.

[0142] Step QC1 of the system 114 involves chromatogram profile matching. Defective chromatograms may be the result of faulty sample extraction or errors in the mass spectrometry setup. These errors affect the quality of the data, which then impairs the overall prediction accuracy of the algorithm. To detect these defects based on the chromatogram profile variability, a sequential neural network model was constructed. The chromatograms obtained from the mass spectrometer for each sample were first converted to JPEG format. The images were then scaled to an appropriate width and length for model training. The images were binarized to separate the chromatograms, facilitating more efficient analysis. The Keras sequential neural network model was then used to train, test, and validate the model. We employed three Keras Conv2D two-dimensional convolutional layers. These generate convolution kernels wrapped around the layered input, which serve to generate the output tensor. An 80-20 training-testing split was employed. The Adam optimizer was then used to iteratively update the network weights based on the training data. The threshold used for correct detection was 0.5, and samples that showed a quality control (QC) score <0.5 were designated as having passed QC1. Samples with a QC score >0.5 were rejected as having failed the QC1 step. We obtained 100% detection accuracy for the defective chromatograms, as shown in Figure 6A.

[0143] Step QC2 of system 114 monitors for the presence of critical m / z ions. The second step of system 114, designated QC2, monitored for the presence of at least six of nine critical ions with m / z values ​​ranging from 100 to 800. The intensity and RT distributions of these nine ions are shown in Figure 6B. The presence of six or more of the nine critical ions in a sample chromatogram is the threshold criterion for passing the QC2 step. Samples with fewer than six of the nine critical ions are rejected as failing the QC2 step. Samples that do not pass this QC2 step are likely to be misclassified by the CDAI algorithm (Figure 6B), making the QC2 step critical for accurate identification of cancer samples.

[0144] Step QC3 of system 114 monitors matrix occupancy. Another layer of quality check, introduced alongside the two quality checks above, included assessing matrix occupancy. This layer, step QC3, relies on the percentage of features that match the matrix size. The threshold was optimized using multiple runs of low-quality samples or inappropriate mass spectrometry run conditions. Based on these studies, we set the minimum matrix occupancy threshold at 15%. This threshold was confirmed through multiple validation runs and subsequently found to improve the robustness and accuracy of the predictive CDAI algorithm. As shown in Figure 6C, using the trained model for QC3 resulted in 100% accuracy in capturing defective samples.

[0145] result The demographic and ethnic distributions of controls and cancer patients (Table 1) were balanced for frequency-matching variables, including age, race, BMI, and cancer stage. None of the observed variations in the distribution of these variables between control and disease cases achieved statistical significance in the training or test sets. Approximately 80% of the cancer cases were stage I. Figure 2 shows a schematic of the complete procedure, along with illustrations of key steps at each stage. A Dionex LC system interfaced online with a QExactive Plus mass spectrometer received an infusion of metabolites isolated from serum. Data preprocessing was first outlined in Figure 2. The following list includes various data preprocessing steps:

[0146] 1. Extraction of metabolic feature nodes Data from metabolomics are known to contain mass imprecision. As a result, the mass of the same identified metabolite in several samples varies slightly. This makes it difficult to compare the intensities of the same metabolite across samples, which impacts downstream AI-based analyses that require robust feature intensity values ​​and dimensionality as the basis for this intensity comparison. To overcome this, we employed various simultaneous and parallel approaches. Metabolic feature nodes were identified using fixed mass boxes that covered the entire mass range. The thresholds for these mass boxes were defined using different techniques, such as mass KNN clustering, uniform linear separation, and virtual locked mass (VLM)-based strategies. By exhaustively searching our techniques across a large number of samples in combination with the HMDB database, we were able to identify robust mass ranges, which we used to align any new datasets.

[0147] 2. Data Filtering. The presence of noise in a dataset can increase the complexity and training time of the model, which reduces the performance of the training algorithm. Data filtering is a process of noise reduction and dimensionality reduction whereby the initial set of raw data is reduced to a more manageable data format that contains target-specific attributes.

[0148] 3. Data normalization / standardization. Because metabolic data vary under different mass spectrometer parameters, normalization techniques are required to reduce data variability. We tried various normalization methods, including Quantile Normalization, Variance Stabilization Normalization, Best Normalization, and Probabilistic Quotient Normalization. Data standardization is a data processing workflow that converts the structure of different datasets into one common format of data. Data standardization deals with the transformation of datasets after data is collected from different sources and before it is loaded into the target system. Various data standardization methods such as standard normalization, L1 and L2 norm standardization were employed on the datasets. The two-stage algorithm used a combination of standardization and normalization, and this method was further adapted for our dataset to allow for normalization of new samples with respect to the training dataset and testing one sample at a time.

[0149] 4. Missing Value Imputation. It is well recognized that missing values ​​in untargeted metabolomics data can be troublesome. In large metabolite panels, measurements are often missing and, if ignored or suboptimally imputed, can lead to biased study results. We employed various supervised and unsupervised multiple imputation techniques, such as Iterative Imputer, misforest, simple impute, and KNN impute, to evaluate the effects of sample size, missing percentage, and correlation structure on the accuracy of the imputation methods.

[0150] 5. Feature Reduction. Dimensionality reduction is the process of reducing the number of random variables under consideration by obtaining a set of master variables. It is an important step in high-dimensional data as it addresses the curse of dimensionality, multicollinearity, noise, computational cost, and visualization. Feature extraction can be unsupervised (PCA) or supervised (LDA, PLS-DA, etc.). We evaluated various feature reduction techniques based on data variance capture and class separation, namely PLSDA R2 maximization, RFE, PCA, non-negative matrix factorization, and LDA.

[0151] 6. Machine learning model development. After going through the above pipeline, the data was fed into an AI machine. An AI model was created to distinguish between cancer and normal, and then between individual cancers.

[0152] With the clinical application of the AI ​​model in mind, we used a stepwise approach in which an AI model was first developed for cancer signal detection (CDAI model), and then a next model, the TOOAI model, classified the tissue of origin of cancer-positive samples. The total of 8,971 samples included endometrial cancer (n = 445), breast cancer (n = 652), cervical cancer (n = 458), ovarian cancer (n = 488), lung cancer (n = 307), leukemia (n = 157), thyroid cancer (n = 169), melanoma (n = 151), colorectal cancer (n = 296), kidney cancer (n = 136), lymphoma (n = 97), pancreatic cancer (n = 134), liver and bile duct cancer (n = 147), gastric cancer (n = 279), head and neck cancer (n = 638), esophageal cancer (n = 143), and prostate cancer (n = 122), as well as other cancers (n = 254), and 3,914 normal control samples. Using the matrix constructed above, we investigated whether there were any differences between these samples based on their metabolic data. As seen in Figure 7, a PLS DA plot was generated using the 18 cancer classes and normal controls. The graph clearly demonstrates how cancer samples can be distinguished from normal control samples using their metabolic signatures. To uncover common patterns in metabolite variation within cancer samples that differ from normal control samples, AI analysis was performed on the data, as detailed below, to measure how well they could be distinguished.

[0153] Development of algorithms for the CDAI model Of the total 8971 cancer samples, 5057 samples were from the 32 cancer classes mentioned in Table 1, and 3914 were normal controls. Normal controls were samples from cancer-free volunteers. The data were randomly split into training and testing datasets in equal proportions, resulting in 2479 cancer samples and 1957 controls in the training set and 2594 cancer samples and 1957 controls in the testing set (Table 3). A complete schematic of the cancer detection process is shown in Figure 1, and the models were evaluated using the parameters log loss, accuracy, sensitivity, and specificity.

[0154] A parametric machine learning model was applied to the training data to obtain a score function according to the intensity value of the features. A class balancing parameter was configured in the model to address the imbalance between cancer and control samples in the training dataset. The final trained model was used to evaluate the score of each sample using the following formula: y_score=x0+x1×I1+x2×I2+x3×I3+......+x n ×I n

[0155] In this formula, x0 is a constant and I i (1≦i≦n) is the intensity of metabolite i present in each sample. The total number of metabolites is represented by the symbol n (n∈[1000,8300]). Figure 11 shows the coefficient x for each metabolite. i Give a value of (1≦i≦n).

[0156] For example, a y-score plot of the trained model applied to the test set for a single split of data containing 32 cancer classes and normal controls is shown in Figure 9. The scatter plot shows the model scores for the control and cancer cases. The model scores are clearly seen to differ between the control and cancer samples, where a y-score threshold of zero is applied to distinguish the two types of results in the confusion matrix as shown. Sensitivity, specificity, and accuracy can be calculated from the following formulas:

number

[0157] Development of algorithms for the TOOAI model To address the clinical implications of cancer-positive samples, i.e., the cancer tissue of origin of the cancer signal detected in the CDAI model, we developed the TOOAI model. Briefly, the TOOAI model is a multi-class algorithm that evaluates a probability score for a cancer-positive sample, suggesting the tissue from which the cancer-positive signal originated.

[0158] To develop the tissue-of-origin algorithm (TOOAI model), a dataset containing cancer samples was first processed according to the steps described in the previous section. Here, out of a total of 5,057 cancer samples, the samples were endometrial cancer, breast cancer, cervical cancer, ovarian cancer, lung cancer, kidney cancer, thyroid cancer, acute myeloid lymphoma, non-Hodgkin's lymphoma, pancreatic cancer, colorectal cancer, liver cancer, gastric cancer, melanoma, head and neck cancer, esophageal cancer, prostate cancer, and "other" (Table 2). The data were randomly split into equal proportions into a training dataset and a test dataset. The complete training and test distributions for this stratum are shown in Table 3.

[0159] The machine learning environment was set to Python 3.10.4. Various algorithms were used to obtain predictive probability functions for cancer samples, with each probability score indicating the occurrence of that cancer type. To achieve this, we used support vector machines, logistic one-to-many, and stochastic gradient descent algorithms. The optimal set of hyperparameters for these parameters was obtained using exhaustive training and testing with the Python Grid search CV package. This resulted in 18 probability scores for each sample, each defining the probability that the sample belongs to one of the 18 cancer tissue types. The trained algorithm found the tissue of origin probability for each sample according to the following formula:

number

[0160] In this formula, a0, a1, a2, ..., a n is a constant, and I i(1≦i≦8000) is the normalized intensity of metabolite i present in each sample. N is the number of cancer type classes included in the training set.

[0161] The final model with the highest dual-class prediction accuracy on the test set was selected for further evaluation, where dual-class prediction accuracy refers to the occurrence of correct predictions in the top two predictions from the model using the probability function defined above. The dual-class prediction accuracy was evaluated for the above single test dataset as an example. The confusion matrix for the final prediction is shown in Figure 10. Table 4 shows the dual-class prediction accuracy for the same. The prediction accuracy of dual-class prediction from this model was evaluated using the following formula:

number

[0162] Ranking of Cancer-Specific Features. The features derived for model prediction include metabolites from the HMDB database. Feature ranking helps identify important metabolites contributing to the model accuracy and also broadens the scope of predictions made by the model in terms of the molecular translation of the resulting cancer signature. Using various feature ranking methods, both parametric and nonparametric-based approaches, we obtained the top 100 metabolites obtained for the cancer signal detection process relevant to all cancer types, as shown in Table 5.

[0163] [Table 3]

[0164] [Table 4(1)] [Table 4(2)]

[0165] [Table 5]

[0166] Table 6(1)

Table 6(2)

Claims

1. 1. A system for the simultaneous detection of multiple early stage cancers in a single analysis, comprising: at least one liquid chromatography (LC) apparatus equipped with a mass spectrometer (MS) (hereinafter abbreviated as LC-MS) for analyzing one or more obtained reconstituted metabolites using LC-MS techniques to measure the masses of metabolite ions in the one or more obtained reconstituted metabolites, said one or more obtained reconstituted metabolites being obtained after reconstitution of one or more dried metabolite extracts extracted from one or more biological fluid samples; at least one processor / computing device for aligning masses obtained from the metabolomic profile of said metabolite ions in said one or more obtained reconstructed metabolites to minimize possible errors in measuring the masses of said metabolite ions; at least one processor / computing device that applies one or more quality control processes to the aligned and normalized data set to identify possible errors in the detection of multiple cancers in the biological fluid sample, wherein the at least one processor / computing device that performs the quality control processes comprises the following steps: (a), (b) and (c) (a) constructing a sequential neural network model to detect errors based on variations in chromatogram profiles of faulty sample extractions or due to errors in mass spectrometry; (b) monitoring the presence of at least six of nine critical ions having m / z values ​​in the range of 100 to 800, wherein the presence of six or more of the nine critical ions in a sample chromatogram is interpreted as a threshold criterion for passing the quality control step; and (c) assessing matrix occupancy, including calculating a minimum matrix occupancy threshold using multiple runs of poor quality samples or inappropriate mass spectrometry run conditions. a processor / computing device configured to execute at least one or more of the following: at least one processor that performs one or more AI / ML processes on the measured metabolite ions; generating a first AI model (Cancer Detection AI (CDAI) model) for identifying diseased cancer samples and distinguishing them from non-diseased normal samples, wherein the measured metabolite ions are randomly divided into a training dataset and a test dataset; Creating a second AI model (tissue of origin identification (TOOAI model)) to further identify and distinguish each individual diseased cancer sample from other diseased cancer samples and the non-diseased normal samples; and applying the TOOAI model to cancer-positive samples determined by the CDAI model to obtain a score assigned to each individual cancer sample, thereby identifying and distinguishing said individual cancer samples from said other cancer samples and said non-diseased normal samples, one score for each cancer type defined by its tissue of origin and representing the probability that the sample belongs to the respective cancer type; a processor and A system including:

2. The LC-MS device comprises: resolving the one or more resulting reconstituted metabolites by ultra-performance liquid chromatography using the LC device; obtaining an ion spectrum of the one or more resulting reconstituted metabolites through the MS instrument; and measuring the masses of metabolite ions present in the ion spectrum of the one or more obtained reconstituted metabolites based on their mass-to-charge ratio or m / z with the MS instrument; The system of claim 1 configured to:

3. To create the first AI model (Cancer Detection AI (CDAI) model), the at least one processor / computing device: applying a logistic regression function by running the AI / ML process on the training dataset of the metabolite ions to find a functional mapping between dependent / target variables and independent variables that separates diseased cancer samples from non-diseased normal samples; configuring one or more class balancing parameters for the response variable to balance class imbalance in the training dataset; applying an optimization process to handle data complexity, thereby creating the first AI model from the training dataset, wherein the first AI model is the Cancer Detection AI (CDAI) model; and applying the CDAI model to the test data set of metabolite ions to identify diseased cancer samples and distinguish them from the non-diseased normal samples based on a function that separates the diseased cancer samples from the non-diseased normal samples. The system of claim 1 configured to:

4. To create the second AI model (Tissue of Origin Identification (TOOAI) model), the at least one processor / computing device: constructing a classifier multi-class classification model using the training dataset, thereby creating the second AI model from the training dataset, the second AI model being a tissue of origin identification (TOOAI model), and the classifier multi-class classification model including at least one or more of a support vector machine, a logistic one-to-many, or a stochastic gradient descent process; and applying the TOOAI model to cancer-positive samples determined by the CDAI model to obtain a score assigned to each individual cancer sample, thereby identifying and distinguishing said individual cancer samples from said other cancer samples and said non-diseased normal samples, one score for each cancer type defined by its tissue of origin and representing the probability that the sample belongs to the respective cancer type; The system of claim 1 configured to:

5. The system comprises: at least one sample collection device for collecting said one or more biological fluid samples from one or more living mammals; at least one precipitation device for extraction of said one or more metabolite extracts from said one or more biological fluid samples by precipitation of proteins present in said biological fluid samples with chilled alcohol comprising at least methanol; at least one phase separator for drying the one or more metabolite extracts extracted from the at least one precipitation device; at least one device for reconstituting said one or more dried metabolite extracts in an aqueous solution in a mobile phase; The system of claim 1 further comprising:

6. 2. The system of claim 1, wherein the at least one processor / computing device for creating the first AI model (Cancer Detection AI (CDAI)) is further configured to use a trained model / process derived from the CDAI model to find a score for each sample in the training set and evaluate the test set to determine an accuracy rate for applying the trained CDAI model.

7. The at least one processor / computing device executing the AI / ML process to generate the second AI model (tissue of origin identification (TOOAI model)) is further configured to determine an accuracy rate of a multi-class classification model in the TOOAI model in distinguishing individual diseased cancer samples from the other diseased cancer samples and the non-diseased normal samples, and in determining the accuracy rate, the at least one processor / computing device executing the AI / ML process: obtaining and providing a probability score for each cancer subclass for any given sample of individual diseased cancer samples, wherein the probability score for one individual diseased cancer sample is distinct from the probability scores for other diseased cancer samples and other diseased cancer samples; The top two scoring cancer subclasses are taken as the model results and matched with the true labels; and constructing a final confusion matrix based on the dual-class accuracy rate of the model; The system of claim 6 , further configured to:

8. 1. A method for simultaneous detection of multiple early stage cancers in a single assay, comprising: analyzing one or more obtained reconstituted metabolites using LC-MS techniques to determine the masses of metabolite ions in said one or more obtained reconstituted metabolites, said one or more obtained reconstituted metabolites being obtained after reconstitution of one or more dried metabolite extracts extracted from one or more biological fluid samples; aligning the masses obtained from the metabolomic profile of the metabolite ions in the one or more obtained reconstituted metabolites to minimize possible errors in the determination of the masses of the metabolite ions; applying one or more quality control processes to the aligned and normalized dataset to identify possible errors in the detection of multiple cancers, the one or more quality control processes comprising the following steps: (a) constructing a sequential neural network model to detect errors based on variations in chromatogram profiles of faulty sample extractions or due to errors in mass spectrometry; (b) monitoring the presence of at least six of nine critical ions having m / z values ​​in the range of 100 to 800, wherein the presence of six or more of the nine critical ions in a sample chromatogram is interpreted as a threshold criterion for passing the quality control step; and (c) assessing matrix occupancy, including calculating a minimum matrix occupancy threshold using multiple runs of poor quality samples or inappropriate mass spectrometry run conditions. and performing one or more AI / ML processes on the measured metabolite ions; generating a first AI model (Cancer Detection AI (CDAI) model) for identifying diseased cancer samples and distinguishing them from non-diseased normal samples, wherein the measured metabolite ions are randomly divided into a training dataset and a test dataset; Creating a second AI model (tissue of origin identification (TOOAI model)) to further identify and distinguish each individual diseased cancer sample from other diseased cancer samples and the non-diseased normal samples; and applying the TOOAI model to cancer-positive samples determined by the CDAI model to obtain a score assigned to each individual cancer sample, thereby identifying and distinguishing said individual cancer samples from said other cancer samples and said non-diseased normal samples, one score for each cancer type defined by its tissue of origin and representing the probability that the sample belongs to the respective cancer type; A process of performing A method comprising:

9. The LC-MS technique resolving the one or more resulting reconstituted metabolites by ultra-performance liquid chromatography using the LC device; obtaining an ion spectrum of the one or more obtained reconstituted metabolites through an MS instrument; and measuring the masses of metabolite ions present in the ion spectrum of the one or more obtained reconstituted metabolites based on their mass-to-charge ratio or m / z with the MS instrument; The method of claim 8 further comprising:

10. To create the first AI model (Cancer Detection AI (CDAI) model), the one or more AI / ML processes: applying a logistic regression function to the training dataset of metabolite ions to find a functional mapping between dependent / target variables and independent variables that separates diseased cancer samples from non-diseased normal samples; configuring one or more class balancing parameters for the response variable to balance class imbalance in the training dataset; applying an optimization process to handle data complexity, thereby creating the first AI model from the training dataset, wherein the first AI model is the Cancer Detection AI (CDAI) model; and applying the CDAI model to the test data set of metabolite ions to identify diseased cancer samples and distinguish them from the non-diseased normal samples based on a function that separates the diseased cancer samples from the non-diseased normal samples. The method of claim 8 , further comprising:

11. To create the second AI model (tissue of origin identification (TOOAI model)), the one or more AI / ML processes: constructing a classifier multi-class classification model using the training dataset, thereby creating the second AI model from the training dataset, the second AI model being a tissue of origin identification (TOOAI model), and the classifier multi-class classification model including at least one or more of a support vector machine, a logistic one-to-many, or a stochastic gradient descent process; and applying the TOOAI model to cancer-positive samples determined by the CDAI model to obtain a score assigned to each individual cancer sample, thereby identifying and distinguishing said individual cancer samples from said other cancer samples and said non-diseased normal samples, one score for each cancer type defined by its tissue of origin and representing the probability that the sample belongs to the respective cancer type; The method of claim 8 , further comprising:

12. The method comprises: collecting said one or more biological fluid samples from one or more living mammals; extraction of said one or more metabolite extracts from said one or more biological fluid samples by precipitation of proteins present in said biological fluid samples with chilled alcohol comprising at least methanol; drying the one or more metabolite extracts extracted from at least one precipitation device; reconstituting said one or more dried metabolite extracts in an aqueous solution in a mobile phase; The method of claim 8 further comprising:

13. 9. The method of claim 8, wherein to create the first AI model (Cancer Detection AI (CDAI)), the one or more AI / ML processes are further executed to use a trained model / process derived from the CDAI model to find a score for each sample in the training set and evaluate the test set to determine an accuracy rate for applying the trained CDAI model.

14. To generate the second AI model (tissue of origin identification (TOOAI model)), the one or more AI / ML processes are further performed to determine an accuracy rate of a multi-class classification model in the TOOAI model in distinguishing individual diseased cancer samples from the other diseased cancer samples and the non-diseased normal samples, and in determining the accuracy rate, the one or more AI / ML processes: obtaining and providing a probability score for each cancer subclass for any given sample of individual diseased cancer samples, wherein the probability score for one individual diseased cancer sample is distinct from the probability scores for other diseased cancer samples and other diseased cancer samples; The top two scoring cancer subclasses are taken as the model results and matched with the true labels; and constructing a final confusion matrix based on the dual-class accuracy rate of the model; The method of claim 8 , further comprising:

15. 10. The system of claim 1, wherein the metabolome profile of metabolite ions is generated using an automated platform including at least a compound discoverer module that extracts data about metabolite ions and their associated features.

Citation Information

Patent Citations

  • Molecular attribute prediction method and device, intelligent equipment and terminal

    CN112420125A

  • Automated boundary detection in mass spectrometry data

    JP2022525427A

  • Monoclonal antibodies to a new antigenic marker in epithelial prostatic cells and serum of prostatic cancer patients

    US5162504A

  • Method for evaluation of breast cancer, breast cancer evaluation apparatus, breast cancer evaluation method, breast cancer evaluation system, breast cancer evaluation program, and recording medium

    WO2008075662A1

  • Female genital cancer evaluation method

    WO2009154296A1