Method for diagnosing early stage non-small cell lung cancer
By measuring specific metabolites in serum and plasma, combined with logistic regression models and smoking history, the problem of insufficient sensitivity and accuracy of existing lung cancer detection methods has been solved. A high-performance combination of plasma metabolite biomarkers has been developed, achieving highly sensitive and specific detection of non-small cell lung cancer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BIOMARK CANCER SYST INC
- Filing Date
- 2020-10-17
- Publication Date
- 2026-08-04
AI Technical Summary
Existing lung cancer detection methods are not sensitive and accurate enough, and the widespread implementation of low-dose computed tomography (LDCT) is hindered by technical and socioeconomic challenges, resulting in a lack of low-cost, minimally invasive early lung cancer detection methods.
By measuring the concentrations of specific metabolites such as β-hydroxybutyric acid, LysoPC 20:3, PC ae C40:6, citric acid, and fumaric acid in serum and plasma, and combining logistic regression models and smoking history, a high-performance plasma metabolite biomarker combination was developed for the diagnosis of non-small cell lung cancer, especially stage I and II.
The model achieved high sensitivity and specificity for the detection of non-small cell lung cancer. The diagnostic model with AUC>0.9 can effectively distinguish early-stage lung cancer patients from healthy controls. The combination of smoking history further improved the performance of the model.
Smart Images

Figure CN115023609B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims U.S. Patent Application No. 62 / 916,486, the contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to a method for diagnosing cancer, and more particularly to a method for diagnosing early-stage non-small cell lung cancer by measuring metabolite biomarkers in serum and plasma. Technical Background
[0004] Lung cancer is the leading cause of cancer-related deaths worldwide. Sensitive and accurate strategies for early lung cancer detection are crucial for improving lung cancer survival statistics. Unfortunately, current methods for detecting or screening for lung cancer are less than ideal. Although low-dose computed tomography (LDCT) has been shown to reduce lung cancer mortality, widespread clinical implementation is hampered by various technical and socioeconomic challenges. Therefore, developing a low-cost, minimally invasive detection method for early lung cancer detection would significantly improve the current situation.
[0005] International Patent Application Publication No. WO2016 / 205960, published on December 29, 2016, discloses a set of biomarkers for a serum test to detect lung cancer, wherein the biomarkers are selected from valine, arginine, ornithine, methionine, spermidine, spermine, diacetylspermine, 00:2, PC aa C32:2, PC ae C36:0 and PC ae C44:5; and lysoPC a 08:2, or a combination thereof.
[0006] The contents of this disclosure
[0007] This disclosure relates to a method comprising determining the concentration of each metabolite in a group of metabolites in a biological sample from a subject, wherein the group of metabolites includes: β-hydroxybutyric acid, LysoPC 20:3, PC ae C40:6, citric acid, carnitine, and fumaric acid; β-hydroxybutyric acid, LysoPC 20:3, PC ae C40:6, and fumaric acid; or β-hydroxybutyric acid, PC ae C40:6, citric acid, and carnitine. In various embodiments, the disclosed method is a method for diagnosing non-small cell lung cancer, and in a particular embodiment, a method for diagnosing stage I or II non-small cell lung cancer.
[0008] This disclosure relates to a method comprising determining the concentration of each metabolite in a group of metabolites in a biological sample from a subject, wherein the group of metabolites includes β-hydroxybutyric acid, LysoPC 20:3, fumaric acid, and spermine. In various embodiments, the disclosed method is a method for diagnosing non-small cell lung cancer, and in a particular embodiment, it is a method for diagnosing stage I non-small cell lung cancer.
[0009] This disclosure relates to the treatment of patients diagnosed with non-small cell lung cancer according to the methods described herein.
[0010] Brief description of the attached figures
[0011] Figure 1 a is a two-dimensional partial least squares discriminant analysis (PLS-DA) plot, which shows a comparison between plasma metabolite data collected from healthy controls (shown in the shaded area on the left) and stage I NSCLC patients (shown in the shaded area on the right);
[0012] Figure 1 b is a projected importance (VIP) plot, which shows the most differentiating metabolites between healthy controls and stage I NSCLC patients. Boxes indicate whether metabolite concentrations are increased (circled) or decreased (not circled) in controls and cases;
[0013] Figure 2 Figure a shows a two-dimensional partial least squares discriminant analysis (PLS-DA) plot, which compares plasma metabolite data collected from healthy controls (shown in the shaded area on the left) and NSCLC patients at all stages (shown in the shaded area on the right). PLS-DA results for healthy controls and NSCLC at all stages;
[0014] Figure 2 b is a projected importance (VIP) plot, which shows the most differentiating metabolites between healthy controls and NSCLC patients at all stages. Boxes indicate whether metabolite concentrations are increased (circled) or decreased (not circled) in controls and cases;
[0015] Figure 3 'a' is the receiver operating characteristic (ROC) curve generated by a metabolite-only logistic regression model used to diagnose stage I NSCLC patients. The ROC curve and its 95% CI on the discovery set are shown as a curve. The ROC curve obtained from the validation set is shown as a step-like line;
[0016] Figure 3 b is the receiver operating characteristic (ROC) curve generated by a logistic regression model of metabolites and smoking history used to diagnose stage I NSCLC patients. The ROC curve and its 95% CI on the discovery set are shown as a curve. The ROC curve obtained from the validation set is shown as a step-like line;
[0017] Figure 4 'a' is a receiver operating characteristic (ROC) curve generated by a random forest exploration model for stage I NSCLC patients with varying numbers of metabolite features. The number of metabolite features in each model is represented by Var. See the bottom left box;
[0018] Figure 4 b is a projected importance (VIP) plot showing the most frequently selected metabolites in healthy controls and stage I NSCLC patients (feature count = 5). Boxes indicate whether metabolite concentrations are increased (circled) or decreased (not circled) in controls and cases;
[0019] Figure 5 a is a two-dimensional partial least squares discriminant analysis (PLS-DA) plot, which shows a comparison between plasma metabolite data collected from healthy controls (shown in the shaded area on the left) and stage II NSCLC patients (shown in the shaded area on the right);
[0020] Figure 5 b is a projected importance (VIP) plot, which shows the most differentiating metabolites between healthy controls and stage II NSCLC patients. Boxes indicate whether metabolite concentrations are increased (circled) or decreased (not circled) in controls and cases;
[0021] Figure 6 a is the receiver operating characteristic (ROC) curve generated by the random forest exploration model for stage II INSCLC patients;
[0022] Figure 6 b is a projection importance (VIP) plot showing the most frequently selected metabolites in stage II INSCLC patients (feature count = 5). Boxes indicate whether metabolite concentrations are increased (circled) or decreased (not circled) in controls and cases;
[0023] Figure 7 'a' is the receiver operating characteristic (ROC) curve generated by a metabolite-only logistic regression model used to diagnose stage II NSCLC patients. The number of metabolite features in each model is represented by Var, shown in the bottom left box. The ROC curve and its 95% CI on the discovery set are displayed as a curve. The ROC curve obtained from the validation set is shown as a step-like line.
[0024] Figure 7 b is the receiver operating characteristic (ROC) curve generated by a logistic regression model of metabolites and smoking history used to diagnose stage II NSCLC patients. The ROC curve and its 95% CI on the discovery set are shown as a curve. The ROC curve obtained from the validation set is shown as a step-like line;
[0025] Figure 8a is a two-dimensional principal component analysis (PCA) score plot, which shows a comparison between plasma metabolite data collected from healthy controls (shown as the shaded area at the bottom) and NSCLC patients at all stages (shown as the shaded area at the top);
[0026] Figure 8 b is a partial least squares discriminant analysis (PLS-DA) plot, which shows a comparison between plasma metabolite data collected from healthy controls (shown in the shaded area on the left) and NSCLC patients at all stages (shown in the shaded area on the right);
[0027] Figure 8 c is a projected importance (VIP) plot, which shows a comparison between plasma metabolite data collected from healthy controls and NSCLC patients at all stages. The most discriminative metabolites are shown in descending order of coefficient scores. Boxes indicate whether metabolite concentrations are increased (circled) or decreased (not circled) in controls and cases;
[0028] Figure 9 'a' is the receiver operating characteristic (ROC) curve generated by a metabolite-only logistic regression model used to diagnose patients with early-stage (stage I+II) NSCLC. The ROC curve and its 95% CI on the discovery set are shown as a curve. The ROC curve obtained from the validation set is shown as a step-like line;
[0029] Figure 9 b is the receiver operating characteristic (ROC) curve generated by a logistic regression model of metabolites and smoking history used to diagnose early-stage (stage I+II) NSCLC patients. The ROC curve and its 95% CI on the discovery set are shown as a curve. The ROC curve obtained from the validation set is shown as a line resembling a step function.
[0030] Figure 10 This is a partial least squares discriminant analysis (PLS-DA) plot, which shows a two-dimensional score plot of quantitative MS metabolite analysis of serum samples from stage I lung cancer patients compared with healthy controls;
[0031] Figure 11 This is a projected importance (VIP) plot, ranking serum metabolites in descending order of importance. The plot is from PLS-DA and ranks metabolites according to their importance in stage I cancer classification. VIP scores (x-axis) with coefficients above 85 indicate highly significant metabolites. The right-hand plot shows whether a specific metabolite is increased or decreased in lung cancer relative to healthy controls. Therefore, LysoPC-20:3 is increased in lung cancer, while spermine is associated with death in lung cancer.
[0032] Figure 12This is a receiver operating characteristic (ROC) analysis of lung cancer metabolites in the serum of stage I lung cancer patients, including those from... Figure 11 The four most important metabolites in the VIP analysis of the serum sample shown; and
[0033] Figure 13 It refers to lung cancer metabolites from stage I lung cancer patients who are smokers included in the model. Figure 11 The receiver operating characteristic (ROC) analysis of the four most important metabolites of VIP in the serum samples shown in the figure was performed. The permutation test of the ROC analysis (repeated 1000 times) showed that the results were significant, with a p-value < 0.001.
[0034] definition
[0035] As used in this article, “smoker” includes “current smoker” and “former smoker” as defined in the Tobacco Glossary of the National Center for Health Statistics (“NCHS”) of the Centers for Disease Control and Prevention (“CDC”).
[0036] The term "non-smoker" as used in this article does not refer to subjects who are "smokers" as defined above, but includes "never-smokers". As used in this article, "smoking frequency" is a value calculated by multiplying the smoking time (in days) by the number of cigarettes smoked per day.
[0037] Detailed description
[0038] A high-performance (AUC>0.9) plasma metabolite biomarker set for the detection of early-stage non-small cell lung cancer (NSCLC) is disclosed. Plasma samples were obtained from 156 biopsy-confirmed NSCLC patients and age- and sex-matched plasma samples from 60 healthy controls. Clinical data and smoking history were also available for all samples. Fully quantified targeted mass spectrometry (MS) analysis (direct injection / LC and tandem MS) was performed on all 216 plasma samples. Two-thirds of the samples were randomly selected for discovery, and one-third for validation. Metabolite concentration data, clinical data, and smoking history were used to determine the optimal biomarker set and the optimal regression model to identify different stages of NSCLC using the discovery set. The same biomarkers and regression model were used and evaluated on the validation model.
[0039] An average of 103 metabolites were quantified in these plasma samples. Univariate and multivariate statistical analyses identified significant differences in β-hydroxybutyrate, LysoPC 20:3, PC ae C40:6, citrate, and fumarate between healthy controls and stage I / II NSCLC. Robust predictive models with area under the curve (AUC) > 0.9 were developed and validated using these metabolites and other readily measurable clinical data to detect NSCLC at different stages.
[0040] Archived plasma samples were obtained from the IUCPQ (Institute of Cardiology and Pulmonary Medicine, University of Quebec) tissue bank, which is home to the Quebec-Santorini Research Foundation's Respiratory Health Network tissue bank in Quebec, Canada. Frozen (-80°C) aliquots of 200–400 μL of plasma were assembled and transported to the Metabolomics Innovation Centre (TMIC) at the University of Alberta, Canada, for quantitative metabolomics analysis. Plasma samples were collected from 156 patients with biopsy-confirmed and biopsy-graded NSCLC and 60 age- and sex-matched healthy controls. The healthy control group included smokers and nonsmokers. Cancer samples contain detailed data on cancer stage, lung cancer histology, age, weight, height, body mass index, smoking status (never / previous / current), smoking history (cigarettes / day and smoking time in years), sex, survival history, medical history, personal cancer history, lung disease status, treatment, tumor size (in millimeters), tumor grade, positive nodules, and data collected from each cancer patient via transthoracic biopsy, transbronchial biopsy, endobronchial biopsy, bronchoalveolar lavage, bronchial brushing, bronchial aspiration, endobronchial ultrasound, transesophageal echocardiography, bone scintigraphy, abdominal ultrasound, abdominal CT scan, chest CT scan, brain CT scan, chest X-ray, mediastinoscopy, chest MRI, brain MRI, and PET scan. Healthy controls contain data on age, weight, height, body mass index, smoking status (never / previous / current), smoking history (cigarettes / day and smoking time in years), and medical history. Patients with any history of liver or kidney disease (and the control group), as well as those who had previously received any anti-tumor drug treatment, were excluded from this cohort.
[0041] Optima TM LC / MS grade formic acid and HPLC grade water were purchased from Fisher Scientific (Ottawa, ON, CA). Sixty-eight pure reference standard compounds were purchased from Sigma-Aldrich (Oakville, ON, CA). Optima TMLC / MS grade ammonium acetate, phenyl isothiocyanate (PITC), 3-nitrophenylhydrazine (3-NPH), 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC), and butylated hydroxytoluene (BHT), HPLC grade pyridine, HPLC grade methanol, HPLC grade ethanol, and HPLC grade acetonitrile (ACN) were also purchased from Sigma-Aldrich (Oakville, ON, CA). Forty-four 2H, 13C, and 15N labeled compounds, used as internal quantitative standards for amino acids, biogenic amines, carnitine and its derivatives, and phosphatidylcholine and its derivatives, were purchased from Cambridge Isotope Laboratories, Inc. (Tewksbur, MA, USA). 3-(3-hydroxyphenyl)-3-hydroxypropionic acid (HPHPA) and 13C-labeled HPHPA were synthesized internally, as described by Khaniani et al. in “A Simple and Convenient Synthesis of Unlabeled and 13C-Labeled 3-(3-Hydroxyphenyl)-3-Hydroxypropionic Acid and Its Quantification in Human Urine Samples”, Metabolites, 2018, 8(4):80. All other standards, including lactic acid, β-hydroxybutyric acid, α-ketoglutarate, citric acid, butyric acid, isobutyric acid, propionic acid, p-hydroxyhippuric acid, succinic acid, fumaric acid, pyruvic acid, hippuric acid, methylmalonic acid, homovanillic acid, indole-3-acetic acid, uric acid, and their isotopically labeled standards, were purchased from Sigma-Aldrich (Oakville, California). Multiscreen “solvinert” filter plates (hydrophobic, PTFE, 0.45 pm, transparent, non-sterile) and 96 DeepWell TM Sigma-Aldrich (Oakville, ON, CA).
[0042] All solid chemicals were carefully weighed to an accuracy of 0.0001 g on a CPA225D semi-micro electronic balance (Sartorius, USA). Stock solutions for each compound were prepared by dissolving the precisely weighed solids in water. Calibration curve standards were obtained by mixing and diluting the corresponding stock solutions with water. Stock solutions for isotopically labeled compounds, including amino acids, biogenic amines, carbohydrates, carnitine and its derivatives, and phosphatidylcholine and its derivatives, were prepared in the same manner. A working internal standard (ISTD) solution mixture in water was also prepared by mixing all the prepared isotopically labeled stock solutions together. For organic acids, stock solutions for isotopically labeled compounds were prepared by dissolving the precisely weighed solids in a 75% aqueous methanol solution. A working internal standard (ISTD) solution mixture in 75% aqueous methanol was prepared by mixing and diluting all the isotopically labeled stock solutions. All standard solutions were aliquoted and stored at -80°C until further use.
[0043] A targeted, MS-based quantitative metabolomics approach is used for direct injection (DI) mass spectrometry (MS) and reversed-phase high-performance liquid chromatography (HPLC) tandem mass spectrometry (MS / MS / multiple sclerosis). This semi-automated 96-well plate assay, combined with an ABI 4000Q-Trap (Applied Biosystems / MDS Sciex) mass spectrometer, can be used to target the identification and quantification of up to 138 different endogenous metabolites, including amino acids, organic acids, biogenic amines, acylcarnitines, glycerophospholipids, sphingolipids, and sugars. The method combines derivatization and extraction of 138 analytes with selective mass spectrometry detection using multiple reaction monitoring (MRM) pairs. Isotope-labeled internal standards and other internal standards are integrated into a special filter cartridge placed within the 96-well plate for precise metabolite quantification. The assay uses an upper 96-well plate with a 96-well filter plate attached below using sealing tape. The first 14 wells of the upper plate are used for quality control and calibration. The first well is used as a double blank, three wells contain blank samples, seven wells contain reference compound standards, and three wells contain quality control samples.
[0044] In short, plasma samples were thawed on ice (in the dark) and vortexed and centrifuged at 18,000 ref (relative centrifugal force or ×g). 10 μL of each sample was loaded into the center of a cartridge on the upper 96-well kit plate and dried in a nitrogen stream. PITC was then added to each well for amine derivatization. After incubation, the cartridge was dried using an evaporator. Metabolites were then extracted by adding 300 μL of methanol containing 5 mM ammonium acetate. The extract was obtained by centrifugation (50 ref 5 min) using a dual-plate system. This allows the contents of the upper 96-well plate to flow into the lower 96-well plate. The extract was subsequently diluted with water for analysis of biogenic amines and amino acids. The extract was diluted with methanol for analysis of sugars, carnitines, and lipids. The extract was then prepared using a... Mass spectrometry analysis of diluted extracts was performed on an HPLC system (Agilent Technologies, Santa Clara, US) using an Applied Biosystems / MDS Analytical Technologies, Foster City, CA 4000 tandem mass spectrometer (Agilent 1100 HPLC, Agilent Technologies, Santa Clara, US) and an Agilent 1100 HPLC.
[0045] For organic acid analysis, 50 μL of plasma sample was thoroughly mixed with ISTD mixed solution and ice-cold methanol, and then incubated overnight at 20°C to precipitate proteins. After removing the sample from the refrigerator, all tubes were centrifuged at 18,000 rpm for 20 minutes (to rotate the protein precipitate). The supernatant was then transferred to each well of a 96-well plate system, and 25 μL of each of the following three reagents were added: 3-NPH (250 mM methanol solution), EDC (150 mM methanol solution), and pyridine for a 2-hour derivatization reaction. After the derivatization reaction was complete, water and BHT solution (2 mg / mL methanol solution) were added to dilute and stabilize the final solution. 10 μL was injected into a plate equipped with… LC-MS / MS analysis was performed using HPLC on a 4000 mass spectrometer.
[0046] Recommended statistical procedures for standard quantitative metabolomics analysis were followed. Specifically, metabolites with more than 50% missing values (across all groups) were removed from further analysis. For metabolites with less than 50% missing values, their values were estimated using half the minimum concentration value of that metabolite. Median normalization, logarithmic transformation, and automatic scaling (centered at the mean and divided by the standard deviation of each variable) were used for data scaling and normalization. Characteristic normality was checked by the Shapiro-Wilk test, with a p-value threshold set to 0.05. Univariate analyses of continuous and categorical data were performed using Student's t-test and Fisher's exact test, respectively. Principal component analysis (PCA) and partial least squares discriminant analysis (PLS-DA) were performed using MetaboAnalyst. A 1000-fold permutation test was performed to minimize the possibility that observed PLS-DA segregation was accidental.
[0047] Logistic regression using the Lasso feature selection algorithm was employed to develop predictive models for NSCLC staging using metabolite and clinical variables. For these regression studies, two-thirds of the sample (40 controls and 40–94 cancer samples, depending on stage) were randomly selected as the discovery set. 10-fold cross-validation was performed on all discovery / training set models. Once the optimal regression model for each cancer stage predictor was determined, the remaining one-third of the sample (20 controls and 20–62 cancer samples, used as the retention set) was used to validate each corresponding regression model. The area under the receiver operating characteristic (AUC), sensitivity / specificity, and 95% confidence intervals for all discovery and validation sets and all models were calculated using MetaboAnalyst.
[0048] A total of 138 different metabolites were tested using our quantitative LC-MS method. 35 metabolites were removed due to their low abundance and high proportion of missing values (>50%). Most of these missing values were due to plasma concentrations of the metabolites below the limit of detection (LOD). The number of samples in each group is summarized in Table 1 below.
[0049] Table 1. Summary of Sample Grouping
[0050]
[0051] Comparisons between cancer patients and healthy controls regarding age, sex, height, weight, and smoking history (yes = previous + current, no = never) were performed using a standard Student's t-test or Fisher's exact test to confirm their demographic comparability. The only significantly different variable was smoking history (p = 2.673 × 10⁻⁶). -13 The effects of multiple clinical variables, including age, sex, height, weight, and smoking history (yes = previous + current, no = never), on lung cancer incidence were further evaluated using logistic regression. The results are summarized in Table 2 below. As expected, only smoking history was identified as a clinical variable significantly associated with lung cancer incidence (p = 1.13 × 10⁻⁶). -13 Although the association between smoking history and lung cancer has been extensively studied and widely accepted, this model suggests that incorporating smoking history (including duration and amount of smoking) into any diagnostic model to identify early-stage lung cancer would be a good strategy.
[0052] Table 2. Correlation studies based on logistic regression: NSCLC and clinical variants
[0053] estimated value Standard error z value p-value (intercept) 5.9317 5.6113 1.0571 0.2905 age 0.0170 0.0232 0.7306 0.4650 gender -0.2930 0.5870 -0.4991 0.6177 height -4.6740 3.3846 -1.3810 0.1673 weight -0.0014 0.0138 -0.1041 0.9171 Smoking (YIN) 2.8079 0.4136 6.7883 <![CDATA[1.13*10 -11 ]]>
[0054] A simple Student's t-test was applied to the metabolomics dataset, revealing significant differences in metabolic profiles between healthy controls and lung cancer patients (all stages). Table 3 below lists 36 metabolites with significant FDR-adjusted p-values (q<0.05) identified by the t-test. In this study, phosphatidylcholines such as PC ae C40:6, PC aa C38:0, and PC aa C40:2 were the most downregulated metabolites in the plasma of NSCLC patients, while lysophosphatidylcholines (LysoPCs) such as LysoPC 20:3 and LysoPC 20:4 were significantly upregulated in cancer patients. Other significantly altered metabolites included β-hydroxybutyrate (increased in NSCLC), methionine sulfoxide (decreased), tryptophan (decreased), carnitine (both C0 and C2, increased), and members of the TCA cycle, such as citrate (decreased) and fumarate (increased).
[0055] Table 3. Metabolites that showed significant differences between normal cases and non-small cell lung cancer patients in univariate statistical analysis.
[0056]
[0057] Multivariate analysis was also performed to further reveal metabolite differences between healthy controls and NSCLC patients at all stages. Using PLS-DA, a significant separation was found between NSCLC patients and healthy controls. Figure 1 a). The permutation test showed that the observed segregation was not accidental (P<0.001). LysoPC 20:3, carnitine, β-hydroxybutyric acid, and PC ae C40:6 were found to have the highest overall coefficient scores driving the segregation ( Figure 1 b).
[0058] Biomarkers that can effectively diagnose lung cancer patients in the early stages of the disease are clearly more valuable than those used in later stages. Therefore, a series of statistical analyses were performed to identify plasma metabolites that could distinguish stage I NSCLC patients from healthy controls. Figure 1 As shown in Figure a, PLS-DA analysis revealed a clearly detectable separation between the stage I NSCLC group and healthy controls. The permutation test indicated that the observed separation between cases and controls was not accidental (p < 0.001). Figure 1b shows the results of the overall coefficient scores from PLS-DA. Based on this analysis, LysoPC 20:3, PC ae C40:6, PC aa C38:0, carnitine, and fumarate appeared to be the most important plasma metabolites distinguishing stage I NSCLC patients from healthy controls.
[0059] Logistic regression and exploratory ROC analysis based on random forest were performed using MetaboAnalyst to identify the optimal combination of metabolites to differentiate between stage I NSCLC and healthy controls. In this analysis, Monte Carlo cross-validation (MCCV) based on balanced subsampling was used to generate receiver operating characteristic (ROC) curves. Using a discovery cohort of plasma samples from 40 healthy controls and 47 patients with stage I NSCLC, the AUC of different ROC models with varying numbers of metabolite profiles ranged from 0.824 to 0.922. Figure 3 a). Figure 3 b shows the most frequently selected metabolites, with LysoPC 20:3, PC ae C40:6, PCaa C38:0, LysoPC 20:4, fumaric acid, carnitine, and β-hydroxybutyric acid identified as the top-ranked metabolites. A logistic regression model was then established to predict the probability of stage I NSCLC (P), as follows: log(P / (lP))=0.258-1.341×PCae C40:6+1.747×LysoPC 20:3+0.913×β-hydroxybutyric acid+0.939×fumaric acid, where the concentration of each named metabolite in the equation is in μM. The ROC curve with a 95% confidence interval (CI) is shown below. Figure 3 As shown in figure a. The AUC of the ROC curve and the AUC of 10-fold cross-validation were 0.931 (95% CI, 0.924–0.955) and 0.923 (95% CI, 0.866–0.980), respectively. The performance of the metabolite-only model was further examined on the validation set (consisting of 20 healthy controls and 23 patients with stage I cancer), and a slightly lower AUC (0.890) was obtained. The ROC curve obtained from the validation set is also shown in figure a. Figure 3 a. Other details of the model are listed in Table 4 below.
[0060] Table 4. Optimal model for detecting stage I NSCLC based on logistic regression: metabolites only.
[0061]
[0062]
[0063] When a patient's smoking history was added, the logical model of the discovery cohort was modified to logit(P) =
[0064] log(P / (lP))=0.311+0.641×smoking amount-1.372×PC ae C40:6+1.623×LysoPC20:3+0.882×β-hydroxybutyric acid+0.65×fumaric acid, where P is the probability of stage I NSCLC. As mentioned earlier, the concentration of each named metabolite in the equation is in μM. Here and in all other models below, smoking amount is calculated by multiplying smoking time (in days) by the number of cigarettes smoked per day. The ROC curves for the corresponding models are shown below. Figure 3 As shown in b. The AUC of the metabolite + smoking history model was 0.942 (95% CI, 0.926–0.957), and after 10-fold cross-validation, it was 0.922 (95% CI, 0.864–0.979). This is similar to the metabolite-only model. When the same metabolite + smoking history model was tested on the validation set, the AUC of the validation cohort was essentially the same as that of the metabolite-only model (0.920, ...). Figure 3 b). Interestingly, the model's sensitivity increased slightly when considering smoking history (Table 5 below).
[0065] Table 5. Optimal model for stage I NSCLC detection based on logistic regression: metabolites plus smoking history.
[0066]
[0067] A series of similar analyses were performed on patients with stage II lung cancer. The corresponding PLS-DA and VIP plots are shown below. Figure 5 As shown in a and 5b. The permutation test indicated that the separation of the observed cases from the normal group was not accidental (p < 0.001). Compared with stage I NSCLC patients, fumaric acid was no longer identified as one of the most important features in the PLS-DA VIP plot, while β-hydroxybutyrate was identified as one of the metabolites with the highest coefficient score.
[0068] Using a cohort of plasma samples from 40 healthy controls and 40 patients with stage II NSCLC, the AUC of different metabolite-only regression models with varying numbers of metabolite signatures ranged from 0.894 to 0.946. Figure 5 a). Figure 5 b shows the most frequently selected metabolites. LysoPC 20:3, tryptophan, β-hydroxybutyrate, PC ae C40:6, glutamate, and carnitine were identified as the most differentiated metabolites.
[0069] Then, a logistic regression model was established to predict the probability of stage II NSCLC (P) with the following equation: logit(P)=log(P / (1-P))=0.346+2.565×β-hydroxybutyrate-2.219×citric acid+2.904×carnitine-1.599×PC aeC40:6, where the concentration of each named metabolite in the equation is in μM. The ROC curve with 95% CI is shown below. Figure 7 As shown in Figure a. The AUC of the ROC curve and the 10-fold cross-validation AUC were 0.980 (95% CI, 0.973–0.987) and 0.952 (95% CI, 0.909–0.995), respectively. Further examination of the performance of the metabolite-only model on the retained validation set (consisting of 20 healthy controls and 20 stage II cancer patients) yielded a slightly lower AUC (0.922). The ROC curves obtained from the validation set are also shown in Figure a. Figure 7 As shown in a. Further details of the model are listed in Table 6 below.
[0070] Table 6. Optimal model for detecting stage II NSCLC based on logistic regression: metabolites only.
[0071]
[0072] When a patient's smoking history was added, the logical model of the discovery cohort was modified to logit(P) =
[0073] log(P / (lP))=0.098+1.489×smoking amount+2.911×β-hydroxybutyric acid-1.627×citric acid+2.605×carnitine-0.702×PC ae C40:6, where P is the probability of stage II non-small cell lung cancer, and the concentration of each named metabolite in the equation is in μM. The corresponding ROC curve of the model is as follows. Figure 7 As shown in b, the AUC of the ROC curve for the metabolite + smoking history model was 0.985 (95% CI, 0.979–0.991), and after 10-fold cross-validation, it was 0.948 (95% CI, 0.900–0.996). When the same metabolite + smoking history model was tested on the validation set, the AUC of the validation set was also close to that of the training set (0.940, ...). Figure 7 (b) Similar to the model for Phase I NSCLC, the sensitivity of the model and the overall model performance on the validation set are improved when smoking history is considered (Table 7 below).
[0074] Table 7. Optimal model for stage II NSCLC detection based on logistic regression: metabolites plus smoking history.
[0075]
[0076] The same method described above was used to obtain a predictive model for the co-diagnosis of stage I+II NSCLC patients (defined as early-stage NSCLC). Using a discovery cohort of plasma samples from 40 healthy controls and 87 patients with early-stage NSCLC, a logistic regression model was built to predict the probability of having early-stage NSCLC (P), as follows: logit(P)=log(P / (lP))=2.346-1.528×PC ae C40:6+1.429×β-hydroxybutyrate-2.481×citric acid+1.03×LysoPC 20:3+1.773×fumaric acid, where the concentration of each named metabolite in the equation is in μM. Figure 9 Figure a shows the ROC curve with 95% CI. The AUC of the ROC curve and the 10-fold cross-validation AUC were 0.974 (95% CI, 0.965–0.982) and 0.959 (95% CI, 0.923–0.995), respectively. The performance of the metabolite-only model was further examined on a validation set (consisting of 20 healthy controls and 43 early-stage patients), yielding a slightly lower AUC (0.898). The ROC curve obtained from the validation set and other details of the model are shown below. Figure 9 a and Table 8 (below) are shown.
[0077] Table 8. Optimal model for detecting stage I+II NSCLC based on logistic regression: metabolites only.
[0078]
[0079]
[0080] When a patient's smoking history was added, the logical model of the discovery cohort was modified to logit(P) =
[0081] log(P / (lP))=2.427+1.425×smoking amount-1.414×PC ae C40:6+1.414×β-hydroxybutyric acid-2.193×citric acid+1.738×LysoPC 20:3+1.44×fumaric acid, where P is the probability of stage II non-small cell lung cancer, and the concentration of each named metabolite in the equation is in μM. The corresponding ROC curve of the model is as follows. Figure 5 As shown in b, the AUC of the ROC curve for the metabolite + smoking history model was 0.982 (95% CI, 0.975–0.990), and after 10-fold cross-validation, it was 0.948 (95% CI, 0.930–1.000). When the same metabolite + smoking history model was tested on the validation set, the AUC of the validation set was quite close to that of the training set (0.933, ...). Figure 5b). Similarly, when smoking history was added to the model, both the model's sensitivity / specificity and model performance were improved (Table 9 below).
[0082] Table 9. Optimal model for detecting stage I+II NSCLC based on logistic regression: metabolites plus smoking history.
[0083]
[0084]
[0085] Compared to early-stage NSCLC, plasma metabolite analysis in patients with advanced NSCLC differed more significantly from that in healthy controls. Both PCA and PLS-DA showed clear separation (Figs. S4a and S4b). VIP data from PLS-DA analysis indicated that ketone body dysregulation appeared to be one of the most typical features of stage IIIB+IV NSCLC (Fig. S4c). Elevated cadaverine (a product of lysine decarboxylation) levels were also identified as one of the most important features distinguishing stage IIIB+IV NSCLC. In contrast, upregulation of LysoPC20:3, a characteristic of stage I / II NSCLC, was not prominent as a significant feature in stage II / III / IV NSCLC. Since the identification of biomarkers for advanced lung cancer was not a primary focus of this study (and due to the relatively small sample size), no logistic regression model for predicting stage IIIB / IV NSCLC was developed.
[0086] The aim of this study was to discover and validate combinations of plasma metabolite (and clinical) biomarkers for the early detection of non-small cell lung cancer (NSCLC). Specifically, plasma metabolite changes in NSCLC patients (at different stages) and healthy (age- and sex-matched) controls were investigated using MS-based quantitative metabolomics. Separate discovery and validation cohorts were used to prevent overtraining and any unintended biases in the results. Three different metabolite-only models and three different metabolite + smoking status models were developed and independently validated for the detection of stage I, II, and Eli NSCLC. Most of these models achieved AUC > 0.9.
[0087] A key advantage of developing blood-based metabolomics assays is their ease of translating into low-cost, high-throughput analysis that can be run in virtually any clinical laboratory equipped with a standard triple quadrupole mass spectrometer. Modified assays specific to the metabolites identified here can be run using as little as 10 μL of plasma at a rate of 4–5 minutes per sample. These promising results suggest the potential development of a minimally invasive, high-performance, high-throughput, low-cost lung cancer screening test for selecting patients for further follow-up and confirmation using LDCT or other lung imaging modalities.
[0088] Therefore, those skilled in the art will understand that this disclosure relates to a method, in certain embodiments, capable of, relating to a method for detecting non-small cell lung cancer (e.g., stage I or II non-small cell lung cancer). The method includes determining the concentration of each metabolite in a group of metabolites from a biological sample from a subject, wherein the group of metabolites includes: β-hydroxybutyric acid, LysoPC20:3, PC ae C40:6, citric acid, carnitine, and fumaric acid; β-hydroxybutyric acid, LysoPC20:3, PC ae C40:6, and fumaric acid; or β-hydroxybutyric acid, PC ae C40:6, citric acid, and carnitine.
[0089] In various embodiments, the metabolite comprises β-hydroxybutyric acid, LysoPC 20:3, PC ae C40:6, and fumaric acid. In various embodiments, the metabolite is essentially composed of β-hydroxybutyric acid, LysoPC 20:3, PC ae C40:6, and fumaric acid. In such embodiments, the method includes determining a probability score of the biological sample according to Formula 1:
[0090] logit(P)=log(P / (lP))=0.258-1.341×PC ae C40:6+1.747×LysoPC 20:3+0.913×β-hydroxybutyric acid+0.939×fumaric acid
[0091] (Formula 1)
[0092] The numerical values for each metabolite in the equation are the metabolite concentrations after median normalization, logarithmic transformation, and autoscaling, expressed in μM. A probability score reaching or exceeding the Phase I threshold indicates that the subject has Stage I non-small cell lung cancer.
[0093] In other embodiments, the subject is a smoker. In such embodiments, the method includes determining a probability score for the biological sample according to Formula 2:
[0094] logit(P)=log(P / (lP))=0.311+0.641×smoking amount-1.372×PC ae C40:6+1.623×LysoPC 20:3+0.882×P-hydroxybutyric acid+0.65×fumaric acid
[0095] (Formula 2).
[0096] The numerical values for each metabolite in the equation are the concentrations of the metabolite after median normalization, logarithmic transformation, and autoscaling, in μM. A probability score reaching or exceeding the stage I smoker threshold indicates that the subject has stage I non-small cell lung cancer.
[0097] In various embodiments, the metabolome includes: β-hydroxybutyrate; PC ae C40:6; citric acid; and carnitine. In some embodiments, the metabolome consists essentially of β-hydroxybutyrate, PC ae C40:6, citric acid, and carnitine. In such embodiments, particularly when the subject is a non-smoker, the method includes determining a Phase I probability score for the biological sample according to Formula 3:
[0098] logit(P)=log(P / (lP))=0.346+2.565×β-hydroxybutyric acid-2.219×citric acid+2.904×carnitine-1.599×PC ae C40:6;
[0099] (Formula 3).
[0100] The numerical values for each metabolite in the equation are the metabolite concentrations after median normalization, logarithmic transformation, and autoscaling, expressed in μM. A probability score reaching or exceeding the stage II threshold indicates that the subject has stage II non-small cell lung cancer.
[0101] In other embodiments, the subject is a smoker. In such embodiments, the method includes determining a Phase I probability score for the biological sample according to Formula 4:
[0102] logit(P)=log(P / (lP))=0.098+1.489×smoking amount+2.911×β-hydroxybutyric acid-1.627×citric acid+2.605×carnitine-0.702×PC ae C40:6
[0103] (Formula 4).
[0104] The numerical values for each metabolite in the equation are the concentrations of the metabolites after median normalization, logarithmic transformation, and autoscaling, in μM. A probability score reaching or exceeding the stage II smoker threshold indicates that the subject has stage II non-small cell lung cancer.
[0105] logit(P)=log(P / (lP))=2.346-1.528×PC ae C40:6+1.429×β-hydroxybutyric acid-2.481×citric acid+1.03×LysoPC 20:3+1.773×fumaric acid;
[0106] In other embodiments, the metabolome includes: β-hydroxybutyric acid; LysoPC 20:3; PC ae C40:6; citric acid; and fumaric acid. In various embodiments, the metabolome consists essentially of β-hydroxybutyric acid, LysoPC 20:3, PC ae C40:6, citric acid, and fumaric acid. In such embodiments, particularly when the subject is a non-smoker, the method includes determining a probability score for the biological sample according to Formula 5:
[0107] logit(P)=log(P / (lP))=2.346-1.528×PC ae C40:6+1.429×β-hydroxybutyric acid-2.481×citric acid+1.03×LysoPC 20:3+1.773×fumaric acid;
[0108] The numerical values for each metabolite in the equation are the concentrations of the metabolite after median normalization, logarithmic transformation, and autoscaling, in μM. A probability score that reaches or exceeds the stage I / II probability threshold indicates that the subject has stage I or II non-small cell lung cancer.
[0109] In other embodiments where the subject is a smoker, the method includes determining a probability score for the biological sample according to Formula 6:
[0110] logit(P)=log(P / (lP))=2.427+1.425×smoking amount-1.414×PC ae C40:6+1.414×β-hydroxybutyric acid-2.193×citric acid+1.738×LysoPC 20:3+1.44×fumaric acid
[0111] (Formula 6).
[0112] The numerical values for each metabolite in the equation are the concentrations of the metabolite after median normalization, logarithmic transformation, and autoscaling, in μM. A probability score that reaches or exceeds the stage I / II probability threshold indicates that the subject has stage I or II non-small cell lung cancer.
[0113] In various embodiments, the metabolite group essentially consists of β-hydroxybutyric acid, LysoPC20:3, PCaeC40:6, citric acid, carnitine, and fumaric acid. Those skilled in the art will understand that in such embodiments involving all six of these metabolites, the probability of a subject having stage I and stage II non-small cell lung cancer can potentially be analyzed simultaneously according to each formula. In such embodiments, particularly when the subject is a non-smoker, the method includes determining a stage I probability score for the biological sample according to Formula 1. A stage I probability score that meets or exceeds the stage I threshold of Formula 1 indicates that the subject has stage I non-small cell lung cancer.
[0114] Additionally, the method may include determining a stage II probability score for the biological sample according to Formula 3. A stage II probability score that meets or exceeds the stage II threshold of Formula 3 indicates that the subject has stage II non-small cell lung cancer.
[0115] However, this method may also include determining a phase I / II probability score for the biological sample according to Formula 5. A phase I / II probability score that meets or exceeds the phase I / II threshold indicates that the subject has phase I or phase II non-small cell lung cancer.
[0116] In an embodiment where the subject is a smoker, the method may include determining a Phase I probability score for the biological sample according to Formula 2. A Phase I probability score that meets or exceeds the Phase I threshold indicates that the subject has Phase I non-small cell lung cancer.
[0117] Additionally, the method may include determining a stage II probability score for the biological sample according to Formula 4. A stage II probability score that meets or exceeds the stage II threshold of Formula 4 indicates that the subject has stage II non-small cell lung cancer.
[0118] Simultaneously, the method may also include determining a phase I / II probability score for the biological sample according to Formula 6. A phase I / II probability score that meets or exceeds the phase I / II threshold of Formula 6 indicates that the subject has phase I or phase II non-small cell lung cancer.
[0119] Of course, those skilled in the art will understand that when determining the concentrations of all six metabolites, the analyses according to Formulas 1, 3, and 5 (or 2, 4, and 6 if the subject is a smoker) can be performed in any order. Alternatively, only one or three analyses may be performed.
[0120] Cancer detection using LYSO-PC 20:3 (lysophospholipids), β-hydroxybutyric acid, fumaric acid, and spermine.
[0121] This disclosure also relates to a group of four serum metabolite biomarkers for the diagnosis of early lung cancer, which exhibit an AUROC (area under the receiver operating characteristic curve) of 0.94 for stage I lung cancer, with 84% specificity and 90% sensitivity. Combined with easily measurable clinical data, namely past smoking history and smoking amount, the AUROC for stage I lung cancer slightly increases to 0.95, with a sensitivity and specificity of 91% and 92%, respectively. This is likely one of the highest AUROCs reported for all lung cancer tests, regardless of stage. The four serum biomarkers are LYSO-PC20:3 (a lysophospholipid), β-hydroxybutyrate, fumarate, and spermine.
[0122] Metabolomics analysis was performed on 216 serum samples from lung cancer patients (n=156) and healthy controls (n=60) using liquid chromatography-mass spectrometry (LC-MS). The lung cancer patient group included 70 patients with stage I lung cancer, 60 patients with stage II cancer, and 26 patients with stage III / IV cancer. All lung cancer patients were identified as having non-small cell lung cancer (NSCLC), the most common form of lung cancer.
[0123] Targeted LC-MS studies using TMIC-Prime TM The assay is performed using a targeted quantitative metabolomics assay kit developed and extensively validated by the Metabolomics Innovation Centre (TMIC) at the Department of Biological Sciences, BSBZ-824, University of Alberta, Edmonton T6G2R3, Alberta, Canada. TMIC-Prime TM The assay measures 143 different endogenous metabolites, including amino acids, acylcarnitines, organic acids, biogenic amines, uremic toxins, glycerophospholipids, sphingolipids, and sugars. TMIC-Prime TM The assay combines direct injection mass spectrometry (DIMS) with a custom-designed reversed-phase LC-MS / MS system optimized for the ABI 4000Q-Trap mass spectrometer from Applied Biosystems / MDS Sciex equipped with an Agilent 1100 series HPLC. This method combines analyte derivatization and extraction with selective MS detection using multiple reaction monitoring (MRM) pairs. Isotope-labeled internal standards are used for metabolite quantification.
[0124] The custom-designed assay kit includes a 96-well plate with a filter plate sealed with adhesive tape, and all reagents and solvents for preparing the plate assay. The first 14 wells of each plate are used for quality control (QC) and instrument calibration, consisting of one blank, three “zero” samples, seven calibration standards, and three QC samples. For all metabolite measurements except for organic acid measurements, serum samples are thawed on ice, then vortexed and centrifuged at 13,000 × g. 10 μL of each serum sample is loaded into the center of the filter on the upper 96-well plate and dried in a nitrogen stream. Phenyl isothiocyanate is then added to derive all amino groups. After incubation, the filter points are dried again using an evaporator. Metabolites are then extracted by adding 300 μL of extraction solvent (MeOH and FhO). The extracts are centrifuged into the lower 96-well plate and then obtained by running a solvent dilution step with MS. For organic acid analysis, 150 μL of ice-cold methanol and 10 μL of an isotopically labeled internal standard mixture are added to 50 μL of serum for overnight protein precipitation. The resulting sample was then centrifuged at 13000×g for 20 minutes. 50 μL of the supernatant was then loaded into the center of each well in a 96-well plate, followed by the addition of 3-nitrophenylhydrazine (NPH) to derive the carboxylate group. After incubation for 2 hours, BHT stabilizer and water were added before LC-MS injection.
[0125] In the LC-MS method, a total of 138 metabolites were quantified in each of 216 serum samples. Statistical preprocessing removed 35 metabolites due to 20% of the MS signal being below the MS detection limit. To identify potential diagnostic metabolites and generate a lung cancer detection model, a series of statistical and computational procedures were performed, as previously described in Wishart, DS (2010) Computational approaches to metabolomics. Methods Mol Biol. 593:283-313. A simple Student's t-test was applied to our metabolomics dataset, revealing significant differences in metabolic characteristics between healthy controls and lung cancer patients (all stages). Multivariate statistical and logistic regression analyses were performed to identify the minimum metabolite set required for accurate diagnosis of early-stage NSCLC. Partial least squares discriminant analysis (PLS-DA) was performed using MetaboAnalyst, as disclosed in Xia, J., et al., (2015) MetaboAnalyst 3.0 - making metabolomics more meaningful. Nucleic Acids Res. 43(W1): W251-W257. This resulted in good separation between NSCLC patients and healthy controls. The generated models were ranked according to their AUROC values (from highest to lowest). Using this protocol, we were able to identify metabolite biomarkers that could distinguish early-stage lung cancer (i.e., stage I lung cancer patients) from healthy controls with AUROC values higher than 0.90. 10-fold cross-validation was applied to validate the models. Sensitivity and specificity were calculated from ROC curves with 95% confidence intervals during the training and validation steps of model building.
[0126] Figure 10 The PLS-DA analysis shows that resulted in detectable separation between lung cancer patients with stage I lung cancer (shown in the shaded area on the right) and healthy controls (shown in the shaded area on the left). Figure 11 The VIP plot is shown. The permutation test showed that the separation of the observed cases from the normal group was highly unlikely to be accidental (P<0.001). The model used to diagnose stage I lung cancer consists of four serum metabolites, such as... Figure 12The ROC curves are shown in the table. The model is based on the levels of LYSO-PC 20:3, β-hydroxybutyric acid, fumaric acid, and spermine, and is represented by the probability of stage I NSCLC, where (P) is log(P / (1-P)=0.504+2.192*LYSO-PC20:3+1.252*β-hydroxybutyric acid+1.23*fumaric acid-1798*spermine). The AUROC values for the training set and the 10-fold cross-validation set were 0.95 (95% CI, 0.94–0.96) and 0.94 (95% CI, 0.90–0.98), respectively, with validation sensitivity and specificity of 0.84 and 0.90, respectively. These metrics indicate that the model is a highly significant predictor of stage I NSCLC. Detailed information on this stage I model is listed in Table 10. A similar analysis was performed for the diagnosis of stage II lung cancer, and a set of metabolites and AUROC values similar to those for the diagnosis of stage I lung cancer were generated.
[0127] Table 10. Detailed information on the logistic regression model used to diagnose stage I lung cancer.
[0128]
[0129] To improve the performance of the diagnostic model, logistic regression was used to assess the impact of multiple clinical variables, including age, sex, height, weight, and smoking history, on lung cancer incidence. Among these clinical parameters, only smoking history was found to be significantly associated with lung cancer incidence (p-value = 1.13 * 10^- ... -11 Further logistic regression analysis of lung cancer incidence and smoking history confirmed a significant positive correlation between former smokers and lung cancer incidence (p = 4.16 x 10^9). -10 The odds ratio was 9.82. Our results also indicate that the incidence of lung cancer is significantly higher among current smokers (p-value = 7.082 * 10^-10^-10). -11 Although the association between smoking history and lung cancer has been extensively studied and widely accepted, our analysis suggests that smoking history (including duration and amount of smoking) should be included in any lung cancer diagnostic model, as it can improve overall diagnostic performance. The ROC curves for models including smoking history are shown below. Figure 13As shown. The logical model established using four metabolites plus the period and amount of smoking is expressed as log(P / (1-P)=0.739+0.68*fumaric acid-1.861*spermine+5.248*smoking time-4.19*cigarettes / day+1.139*β-hydroxybutyric acid+1.776*LYSO-PC 20:3, where P is the probability of stage 1 NSCLC. The AUROC from the training set was 0.96 (95% CI, 0.95–0.97), and the AUROC from 10-fold cross-validation was 0.95 (95% CI, 0.903–0.985). The sensitivity and specificity of the validation set were 91% and 92%, respectively. Full details of the logistic regression model can be found in Table 11. We have identified four metabolite biomarkers for diagnosing stage 1 lung cancer that are present in serum and can be rapidly and simply tested in the blood. Another advantage of our early lung cancer detection is that it is, in fact, a multi-component test. The advantage of using a multi-component biomarker set is that the shape of the ROC curve can be adjusted to optimize sensitivity / specificity, thereby significantly reducing the number of false negatives at the cost of increased false positives (which is the preferred method for screening tests). ROC curve shape adjustment cannot be performed using a single biomarker set.
[0130] Table 11. Detailed information on the logistic regression model (including smoking history) used to diagnose stage I lung cancer.
[0131]
[0132] Based on the foregoing knowledge, those skilled in the art will understand that aspects of this disclosure relate to a method, which in various respects may be a method for diagnosing non-small cell lung cancer. The method includes determining the concentration of each metabolite in a group of metabolites from a biological sample of a subject, wherein the group of metabolites includes β-hydroxybutyric acid, LysoPC 20:3, fumaric acid, and spermine. In various embodiments, the metabolite group consists of β-hydroxybutyric acid, LysoPC 20:3, fumaric acid, and spermine.
[0133] This method may also include determining a probability score for the biological sample according to Formula 7:
[0134] logit(P)=log(P / (lP))=0.504+2.192×LysoPC 20:3+2.252×β-hydroxybutyric acid+1.23×fumaric acid-1.798×spermine
[0135] (Formula 7)
[0136] The numerical values for each metabolite in the equation are the metabolite concentrations after median normalization, logarithmic transformation, and autoscaling, expressed in μM. A probability score reaching or exceeding the Stage I threshold indicates that the subject has Stage I non-small cell lung cancer. This implementation is particularly predictive for non-smokers.
[0137] In other embodiments, the subject may be a smoker. In such embodiments, the method further includes determining a probability score for the biological sample according to Formula 8:
[0138] 0.739 + 0.68 × fumaric acid - 1.861 × spermine + 5.248 × smoking time - 4.19 × cigarettes / day + 1.139 × β-hydroxybutyric acid + 1.776 × LYSO-PC 20:3;
[0139] (Formula 8)
[0140] The numerical values for each metabolite in the equation are the concentrations of the metabolites after median normalization, logarithmic transformation, and autoscaling, in μM.
[0141] A probability score that meets or exceeds the stage I threshold indicates that the subject has stage I non-small cell lung cancer.
[0142] Treatment of non-small cell lung cancer
[0143] Those skilled in the art will understand that once a subject is diagnosed with stage I or II non-small cell lung cancer according to the methods disclosed herein, the subject can be treated with treatments known in the art.
[0144] Treating a subject's lung cancer may include administering a therapeutic agent to the subject. The therapeutic agent may include a variety of agents known or found to be effective in treating non-small cell lung cancer, including but not limited to: cisplatin; carboplatin; paclitaxel; albumin-bound paclitaxel; docetaxel; gemcitabine; vinorelbine; etoposide; pemetrexed; bevacizumab; ramucirumab; erlotinib; afatinib; gefitinib; osimertinib; dacomitinib; nutulimab; crizotinib; ceritinib; lorlatinib; entrectinib; dabrafenib; trametinib; serpatinib; prasacetinib; carmatinib; larotrectinib; entrectinib; nivolumab; pembrolizumab; atezolizumab; durvalumab; ipilimumab; or combinations thereof.
[0145] Therefore, those skilled in the art will understand that aspects of this disclosure relate to the use of therapeutic agents to treat a subject diagnosed with non-small cell lung cancer according to the methods described herein. The therapeutic agents may include any agent known to be used to treat non-small cell lung cancer, including but not limited to: cisplatin; carboplatin; paclitaxel; albumin-bound paclitaxel; docetaxel; gemcitabine; vinorelbine; etoposide; pemetrexed; bevacizumab; ramucirumab; erlotinib; afatinib; gefitinib; osimertinib; dacomitinib; nutulimab; crizotinib; ceritinib; lorlatinib; entrectinib; dabrafenib; trametinib; serpatinib; prasacetinib; carmatinib; larotrectinib; entrectinib; nivolumab; pembrolizumab; atezolizumab; durvalumab; ipilimumab; or combinations thereof.
[0146] Those skilled in the art will understand that many of the details provided above are merely examples and are not intended to limit the scope of the invention, which will be determined with reference to the appended claims.
Claims
1. Use of a reagent for determining a group of metabolites in the manufacture of a clinical test reagent for detecting stage I NSCLC in a subject, wherein the group of metabolites includes β-hydroxybutyric acid, PC ae C40:6, LysoPC 20:3 and fumaric acid.
2. The use as claimed in claim 1, wherein the group of metabolites further includes at least one of PC aa C38:0, carnitine, and LysoPC 20:
4.
3. Use of a reagent for determining a group of metabolites in the manufacture of a clinical test reagent for detecting stage II NSCLC in subjects, wherein the group of metabolites includes β-hydroxybutyric acid, citric acid, carnitine, and PC ae C40:
6.
4. The use as described in claim 3, wherein the group of metabolites further includes at least one of LysoPC 20:3, tryptophan, and glutamic acid.
5. Use of a reagent for determining a group of metabolites in the manufacture of a clinical test reagent for detecting early NSCLC in a subject, wherein the group of metabolites includes β-hydroxybutyric acid, citric acid, LysoPC 20:3, PC ae C40:6 and fumaric acid.
6. The use as described in claim 5, wherein the group of metabolites further includes at least one of PC aa C38:0, carnitine, LysoPC 20:4, LysoPC 20:3, tryptophan, and glutamic acid.
7. Use of a reagent for determining a group of metabolites in the manufacture of a clinical test reagent for detecting stage I NSCLC in a subject, wherein the group of metabolites includes β-hydroxybutyric acid, LysoPC 20:3, fumaric acid and spermine.