Methods for predicting risk of recurrence and metastases in lung cancer patients using circulating tumor DNA and spatial biomarkers

By detecting ctDNA levels and aerogenous tumor spread, the method accurately predicts lung cancer recurrence and metastases, enabling personalized surgical interventions to reduce recurrence rates.

WO2025212823A1PCT designated stage Publication Date: 2025-10-09MEMORIAL SLOAN KETTERING CANCER CENT +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/022829
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-03
Filing Date
2025-04-02
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current methods fail to accurately predict the risk of recurrence and metastases in lung cancer patients, particularly for small lung adenocarcinomas, leading to high recurrence rates and suboptimal treatment strategies.

Method used

The method involves detecting ctDNA levels at a variant allele fraction (VAF) of 0.1%-0.5% and performing lobectomy or segmentectomy based on the presence of aerogenous tumor spread in peritumoral lung parenchyma, using machine learning classifiers trained on features like ctDNA concentration, VAF, and STAS status to predict recurrence and metastases.

Benefits of technology

This approach enhances the accuracy of predicting recurrence and metastases, allowing for tailored surgical interventions like lobectomy or segmentectomy, thereby improving patient outcomes and reducing recurrence rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025022829_09102025_PF_FP_ABST
    Figure US2025022829_09102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for accurately predicting the risk of recurrence and metastases in lung cancer (e.g., lung adenocarcinoma) patients using ctDNA and 'spread through air spaces' (STAS) biomarkers.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS FOR PREDICTING RISK OF RECURRENCE AND METASTASES IN LUNG CANCER PATIENTS USING CIRCULATING TUMOR DNA AND SPATIAL BIOMARKERSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 573,709, filed April 3, 2024, the entire contents of which are incorporated herein by reference.TECHNICAL FIELD

[0002] The present technology relates generally to methods for accurately predicting the risk of recurrence and metastases in lung cancer (e.g., lung adenocarcinoma) patients using ctDNA and ‘spread through air spaces’ (STAS) biomarkers.STATEMENT OF GOVERNMENT SUPPORT

[0003] This invention was made with government support under grant numbers CA008748, CA236615, CA235667 and CA214195 awarded by National Institutes of Health. The government has certain rights in the invention.BACKGROUND

[0004] The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.

[0005] Despite early detection and curative-intent resection, 20%-30% of small lung adenocarcinomas (ADCs; <3 cm, stage IA) recur, with postrecurrence 5-year survival of 20%- 25%, underscoring the unmet need to develop strategies to overcome this resistance to therapy. The current treatment for small lung ADC is based on the results of a 1995 randomized trial that examined limited resection (LR, removal of a small portion of lung [2%-6% loss of lung function]) vs. lobectomy (LO, removal of a lobe [10%-25% loss of lung function]). This trial and others have shown that LR is associated with a higher incidence of locoregional recurrence(15%-25%). Recently published randomized studies showed that overall survival (OS) and disease-free survival (DFS) were noninferior after LR or segmentectomy vs. after LO; however, patients in these studies were selected (tumor <2 cm with <50% consolidation / tumor ratio on CT scan; lymph nodes negative for cancer intraoperatively), and recurrence rates remained high.

[0006] Accordingly, there is an urgent need for biomarkers that accurately predict the risk of recurrence and metastases in lung cancer patients so as to inform treatment selection.SUMMARY OF THE PRESENT TECHNOLOGY

[0007] In one aspect, the present disclosure provides a method for preventing risk of recurrence and / or metastases in a lung cancer patient in need thereof comprising (a) detecting ctDNA levels in a biological sample obtained from the lung cancer patient, wherein the ctDNA levels are detected at a variant allele fraction (VAF) detection limit of at least 0.1 %-0.5%; and (b) performing a lobectomy on the lung cancer patient, wherein the lung cancer patient comprises at least one lung tumor and exhibits an aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor.

[0008] In another aspect, the present disclosure provides a method for preventing risk of recurrence and / or metastases in a lung cancer patient in need thereof comprising performing a lobectomy on the lung cancer patient, wherein the lung cancer patient comprises detectable preoperative ctDNA levels, wherein the ctDNA levels are detected at a variant allele fraction (VAF) detection limit of at least 0.1%-0.5%; and wherein the lung cancer patient comprises at least one lung tumor and exhibits an aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor.

[0009] Additionally or alternatively, in some embodiments of the methods disclosed herein, the ctDNA levels are detected at a VAF detection limit of from about 0.1% to about 0.5%, from about 0.5% to about 2%, from about 2% to about 10% or from about 10% to about 99%. In certain embodiments, the ctDNA levels are detected at a VAF detection limit of about 0.1%, about 0.2%, about 0.3%, about 0.4%, about 0.5%, about 0.6%, about 0.7%, about 0.8%, about 0.9%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%,about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99%.

[0010] In some embodiments, the biological sample has a cfDNA concentration ranging from about 3 pg / pL to 5.5 ng / pL. In some embodiments, the biological sample has a cfDNA concentration of about 3 pg / pL, about 4 pg / pL, about 5 pg / pL, about 6 pg / pL, about 7 pg / pL, about 8 pg / pL, about 9 pg / pL, about 10 pg / pL, about 15 pg / pL, about 20 pg / pL, about 25 pg / pL, about 30 pg / pL, about 35 pg / pL, about 40 pg / pL, about 45 pg / pL, about 50 pg / pL, about 55 pg / pL, about 60 pg / pL, about 65 pg / pL, about 70 pg / pL, about 75 pg / pL, about 80 pg / pL, about 85 pg / pL, about 90 pg / pL, about 100 pg / pL, about 125 pg / pL, about 150 pg / pL, about 175 pg / pL, about 200 pg / pL, about 225 pg / pL, about 250 pg / pL, about 275 pg / pL, about 300 pg / pL, about 325 pg / pL, about 350 pg / pL, about 375 pg / pL, about 400 pg / pL, about 425 pg / pL, about 450 pg / pL, about 475 pg / pL, about 500 pg / pL, about 525 pg / pL, about 550 pg / pL, about 575 pg / pL, about 600 pg / pL, about 625 pg / pL, about 650 pg / pL, about 675 pg / pL, about 700 pg / pL, about 725 pg / pL, about 750 pg / pL, about 775 pg / pL, about 800 pg / pL, about 825 pg / pL, about 850 pg / pL, about 875 pg / pL, about 900 pg / pL, about 925 pg / pL, about 950 pg / pL, about 975 pg / pL, about 1 ng / pL, about 1.25 ng / pL, about 1.5 ng / pL, about 1.75 ng / pL, about 2 ng / pL, about 2.25 ng / pL, about 2.5 ng / pL, about 2.75 ng / pL, about 3 ng / pL, about 3.25 ng / pL, about 3.5 ng / pL, about 3.75 ng / pL, about 4 ng / pL, about 4.25 ng / pL, about 4.5 ng / pL, about 4.75 ng / pL, about 5 ng / pL, about 5.25 ng / pL, or about 5.5 ng / pL.

[0011] In yet another aspect, the present disclosure provides a method for preventing risk of recurrence and / or metastases in a lung cancer patient in need thereof comprising performing a limited resection or segmentectomy on the lung cancer patient, wherein the lung cancer patient comprises at least one lung tumor and exhibits an aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor; and wherein the lung cancer patient does not comprise any detectable pre-operative ctDNA levels.

[0012] In any of the preceding embodiments of the methods disclosed herein, the lung cancer is lung adenocarcinoma. The lung adenocarcinoma may comprise a histologic subtype selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. Additionally or alternatively, in some embodiments, the cancer is Stage 1, Stage 2, Stage 3, or Stage 4 cancer.

[0013] Additionally or alternatively, in certain embodiments, the lung cancer comprises a truncal loss or loss of heterozygosity (LOH) of 3p, a truncal loss or loss of heterozygosity (LOH) of 3q, a truncal gain of Iq, and / or a truncal gain of 8q. In other embodiments of the methods described herein, the lung cancer comprises a TP53 mutation, a KRAS mutation, a SMARCA4 mutation, a CCNE1 amplification, and / or a truncal loss or loss of heterozygosity (LOH) of chromosome 21q.

[0014] In any and all embodiments of the methods disclosed herein, the aerogenous spread of tumor cells in the lung parenchyma are colocalized with M2 macrophages and T regulatory cells. Additionally or alternatively, in some embodiments, the aerogenous spread of tumor cells in the peritumoral lung parenchyma beyond the edge of the at least one lung tumor is identified via histopathologic analysis of lung tissue sections obtained from the lung cancer patient. The lung tissue sections may be frozen sections or permanent sections. In any of the above embodiments of the methods disclosed herein, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0015] In any and all embodiments of the methods disclosed herein, the biological sample is whole blood, serum or plasma.

[0016] In one aspect, the present disclosure provides a method of training a machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, comprising: (a) receiving data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; (b) generating a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, wherein applying the machine learning method comprises: (i) applying a machine learning technique to the training dataset; (ii) performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the classifier; and (iii) determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In some embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura. In some embodiments, the STAS status is based on pathologist assessments or through machine learning predictions.

[0017] Additionally or alternatively, in some embodiments, the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4. Examples of surgery type include, but are not limited to, pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection. The adjuvant therapy may comprisechemotherapy or immunotherapy. In other embodiments, the lung cancer patient has not received adjuvant therapy.

[0018] In any of the preceding embodiments of the methods disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of clustercell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0019] In any of the preceding embodiments of the methods disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0020] Additionally or alternatively, in some embodiments, the methods of the present technology further comprise applying the classifier to data on a lung cancer patient to generate a predictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In certain embodiments, the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0021] Additionally or alternatively, in certain embodiments, the methods of the present technology further comprise performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments, the methods described herein further comprise performing asegmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

[0022] In one aspect, the present disclosure provides a method of estimating risk of recurrence and / or metastases in a lung cancer patient using a machine learning classifier, the method comprising: (a) receiving patient data corresponding to a plurality of features for the lung cancer patient; (b) applying the machine learning classifier to the patient data to generate a predictor; and (c) determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the machine learning classifier is trained by: (A) receiving cohort data on a cohort of lung adenocarcinoma (LUAD) subjects, each LUAD subject in the cohort comprising at least one lung tumor; (B) generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (C) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura. In some embodiments, the STAS status is based on pathologist assessments or through machine learning predictions.

[0023] In some embodiments, the methods of the present technology further comprise performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrenceand / or metastases based on the predictor and the operating-point threshold. In some embodiments, the methods disclosed herein further comprise administering an effective amount of an adjuvant therapy to the lung cancer patient, optionally wherein the adjuvant therapy comprises chemotherapy or immunotherapy. In other embodiments, the methods of the present technology further comprise performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold. The predictor may comprise a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0024] Additionally or alternatively, in some embodiments, the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

[0025] In any of the preceding embodiments of the methods disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of clustercell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0026] In any of the above embodiments of the methods disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In someembodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0027] Additionally or alternatively, in some embodiments of the methods described herein, one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA, and / or one or more of the plurality of features for each subject in the cohort are determined by assaying blood and / or sequencing tumor DNA.

[0028] In another aspect, the present disclosure provides a machine learning system for training a machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, the system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients; wherein applying the machine learning method comprises: (a) applying a machine learning technique to the training dataset; (b) performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and (c) determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura. In some embodiments, the STAS status is based on pathologist assessments or through machine learning predictions.

[0029] In any of the preceding embodiments of the machine learning systems disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0030] Additionally or alternatively, in some embodiments, the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4. Examples of surgery type include, but are not limited to, pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection. The adjuvant therapy may comprise chemotherapy or immunotherapy. In other embodiments, the lung cancer patient has not received adjuvant therapy.

[0031] In any of the above embodiments of the machine learning systems disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0032] Additionally or alternatively, in some embodiments of the machine learning systems of the present technology, the instructions further cause the processor to apply the classifier todata on a lung cancer patient to generate a predictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In certain embodiments, the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0033] Additionally or alternatively, in certain embodiments of the machine learning systems disclosed herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the machine learning systems disclosed herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

[0034] In yet another aspect, the present disclosure provides a computing system for estimating risk of recurrence and / or metastases in lung cancer patients, the computing system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: (a) receive patient data corresponding to a plurality of features for the lung cancer patient; (b) apply a machine learning classifier to the patient data to generate a predictor; and (c) determine whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the classifier is trained by: (A) receiving cohort data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; (B) generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (C) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models withan accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura. In some embodiments, the STAS status is based on pathologist assessments or through machine learning predictions.

[0035] In any of the above embodiments of the computing systems disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0036] Additionally or alternatively, in some embodiments of the computing system described herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the computing system described herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold. The predictor may comprise a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0037] Additionally or alternatively, in some embodiments, the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STASfrom the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

[0038] In any of the preceding embodiments of the computing systems disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0039] In any and all embodiments of the computing systems described herein, one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA.

[0040] In one aspect, the present disclosure provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a machine learning system, configure the machine learning system to train a machine learning classifier to estimate risk of recurrence and / or metastases in lung cancer patients, the instructions configured to cause the processor to: (a) receive data on a cohort of lung adenocarcinoma (LUAD) subjects, each LUAD subject in the cohort comprising at least one lung tumor; (b) generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (c) apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients; wherein applying the machine learning method comprises: (A) applying a machine learning technique to the training dataset; (B) performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and (C)determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura. In some embodiments, the STAS status is based on pathologist assessments or through machine learning predictions.

[0041] In any of the above embodiments of the computer-readable storage medium disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0042] Additionally or alternatively, in some embodiments, the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4. Examples of surgery type include, but are not limited to, pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection. The adjuvant therapy may comprise chemotherapy or immunotherapy. In other embodiments, the lung cancer patient has not received adjuvant therapy.

[0043] In any of the above embodiments of the computer-readable storage medium disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non- circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0044] Additionally or alternatively, in some embodiments of the computer-readable storage medium of the present technology, the instructions further cause the processor to apply the classifier to data on a lung cancer patient to generate a predictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In certain embodiments, the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0045] Additionally or alternatively, in some embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

[0046] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a computing system, configure the computing system to estimate risk of recurrence and / or metastases in lung cancer patients, the instructions configured to cause the processor to: (a) receive patient data corresponding to a plurality of features for the lung cancer patient; (b) apply a machine learning classifier to the patient data to generate a predictor; and (c) determine whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-pointthreshold, wherein the classifier is trained by: (A) receiving cohort data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; (B) generating a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (C) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. Tn certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura. In some embodiments, the STAS status is based on pathologist assessments or through machine learning predictions.

[0047] In any of the above embodiments of the computer-readable storage medium disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performingan exhaustive grid search technique. In some embodiments, the STAS status is based on pathologist assessments or through machine learning predictions.

[0048] Additionally or alternatively, in some embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold. The predictor may comprise a cumulative incidence function (OF) for lung recurrence and / or metastases.

[0049] Additionally or alternatively, in some embodiments, the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

[0050] In any of the foregoing embodiments of the computer-readable storage medium disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non- circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0051] In any of the preceding embodiments of the computer-readable storage medium disclosed herein, one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA.BRIEF DESCRIPTION OF THE DRAWINGS

[0052] FIGs. 1A-1G show determinants of inter-tumoral growth pattern heterogeneity. FIG. 1A: Overview of TRACERx 421 LUAD cohort (n = 244 tumors). Each column represents one tumor. Invasive mucinous adenocarcinomas (IMAs) were included. Fetal adenocarcinoma, colloid adenocarcinoma, and two tumors from a collision tumor determined by genomic analysis were excluded from the analysis. The proportion of each growth pattern based on diagnostic sectional area, the growth pattern per region, and basic clinical information are summarized. FIGs. 1B-1C: Correlation of genomic variables and (FIG. IB) proportion of high-grade patterns and (FIG. 1C) proportion of each growth pattern within each tumor, with high-grade patterns indicated in bold. Color scale reflects Spearman’s rank correlation coefficient (rho). Correlation P values were corrected for multiple testing according to Benjamini -Hochberg (BH) and and asterisks indicate q value ranges * q < 0.05, ** q <0.01, *** q < 0.001, **** q < 0.0001. TMB, tumor mutational burden; ITH, intra-tumor heterogeneity; LOH, loss of heterozygosity; SCNA, somatic copy number alteration. FIG. ID: Adjusted odds ratios (OR) of truncal genomic alterations associated with the predominantly high-grade pattern tumors. Genomic alterations selected by the model simplification are shown, with statistically significant alterations indicated in bold. The OR and P values (type II ANOVA) in the figure come from single multivariable logistic regression analysis. Asterisks indicate P value ranges * P < 0.05, ** P <0.01, *** P < 0.001. Color represents the type of genomic alteration. FIG. IE: Mutual exclusivity and cooccurrence of truncal driver gene mutations and chromosome arm somatic copy number alterations in predominantly low / mid-grade tumors (n = 116). Color of the edge represents the relationship (mutual exclusivity vs co-occurrence). The negative log of the q value (BH method) is represented in colour scale within each tile. Relationships with q < 0.1 are shown and asterisks indicate q value ranges * q < 0.05, ** q <0.01. FIG. IF: Comparison of ploidy adjusted mean copy number of chromosomal arm and driver genes between high-grade and low / mid-grade predominant tumors. Fixed effect coefficients of the linear mixed effect model with each tumor as a random effect are displayed on the x-axis, and the negative log of the q value (BH method) is displayed on the y-axis. Color represents the sign of the mean ploidy adjusted copy number, stratified with predominance of high-grade and low / mid-grade patterns.Data points with q value > 0.05 are colored in grey. Horizontal red dashed line represents q = 0.05. FIG. 1G: Gene set enrichment analysis of Hallmark gene sets between predominantly high- and low / mid-grade tumors. The normalized enrichment score is displayed on the x-axis and indicates the enrichment for a given gene set. Gene sets with q < 0.25 (BH method) are shown.

[0053] FIGs. 2A-2F demonstrate morphological intra-tumoral heterogeneity reflects genomic heterogeneity. FIG. 2A: Schematic illustrating regions with different or the same growth patterns within each tumor. FIG. 2B: Genomic distance between regions calculated by presence of somatic mutation (left, n = 51 tumors) and LOH (right, n = 51 tumors). Genomic distances between identical (same) growth pattern regions and different growth pattern regions were compared. Each point represents a distance between a pair of regions in a tumor, and the number of regional pairs is shown under each group. Tumors with regions containing both different growth pattern pair(s) and same growth pattern pair(s) are included in the analysis. Centre line, median; box limits, upper and lower quartiles; whiskers, 1.5x interquartile range. P values were calculated using a linear mixed effects model, with each tumor as a random effect. FIG. 2C: Schematic illustrating inference of ancestor-like and descendant-like regional pairs using shared and private LOH profdes per cytoband. After building a LOH tree, if a branch length of Regionl (Rl) is shorter than 2% of the trunk length, namely if R1 has private LOH burden less than 2% of the shared LOH burden, then Rl is inferred as a common ancestor-like region. Conversely, if a branch length of Region2 (R2) is longer than 10% of trunk length, namely if R2 has private LOH burden more than 10% of shared LOH burden, then R2 is inferred as a descendant-like region. FIG. 2D: Comparison of growth pattern (grade) between inferred ancestral-like and descendant-like regions. Only tumors with mixed pattern grades are included in the analysis. Color represents the transition of grade from ancestor-like to descendant-like region. Empirical P value was calculated using a permutation test (1000 permutations, randomizing growth patterns within each tumor, Monte-Carlo procedure). FIGs. 2E-2F: Proportion of tumors which are purely solid, mixed pattern with solid component, and without any solid component, compared between the tumors with and without (FIG. 2E) truncal gain of arm or focal 3q (3q21.3-3q29) and (FIG. 2F) truncal SMARCA4 mutation and / or LOH.

[0054] FIGs. 3A-3G show evolution of LUAD growth patterns through metastasis. FIG. 3A: Overview of metastasis samples in the TRACERx LUAD cohort (n = 65 tumors). Growth pattern and the presence of seeding clones in primary tumor, growth pattern and the site of metastasis samples, timing of divergence of the metastasizing clone, and presence of the tumor spread through air space (STAS) in the primary tumor are shown. FIG. 3B: Frequency of metastasis samples analysed according to the growth pattern at metastasis. The y-axis represents the proportion of the metastatic samples, with the color representing the predominant subtype of the primary tumor. Multiple metastasis samples from the same primary tumor are counted independently. FIG. 3C: Schematic showing the calculation of mean grade scores of nonseeding regions and seeding regions in the primary tumor, as well as metastatic samples. Grade scores of 1, 2, and 3 were given for low-, mid- and high-grade patterns respectively, and mean scores per group were calculated for each tumor. FIG. 3D: Comparison of growth patterns between seeding and non-seeding regions in primary tumors. Growth patterns were transformed into scores (1 : low-grade, 2 : mid-grade and 3: high-grade) and mean scores of non-seeding region(s) and seeding region(s) were calculated for each tumor, as described in FIG. 3C. Tumors harboring at least one seeding and non-seeding region with growth pattern annotation were included in the analysis (n = 30). Mean scores of growth patterns in seeding and nonseeding regions were calculated. The median is indicated by the red horizontal line. A two- sided Wilcoxon signed-rank test was used. FIG. 3E: Comparison of growth pattern between metastasis and the primary tumor seeding regions (n = 60). The median is indicated by the red horizontal line. A two-sided Wilcoxon signed-rank test was used. FIG. 3F: Example of phylogenetic tree (CRUK0543) including multiple metastases to lymph nodes resected at surgery. Each node in the tree represents a mutational cluster and their color indicates the following: blue, mutational cluster only seen in papillary region (primary tumor regions R2, 3, 4, 5, 6, 7); pink, mutational clusters only seen in micropapillary region (primary tumor region Rl); green, mutational clusters only seen in cribriform regions (metastatic LN #8); yellow, mutational clusters seen in regions with different patterns. LN# 10 (acinar) and LN#7 (cribriform) were predicted to have identical mutational clones. Asterisks represent most recent common ancestors of primary tumor regions and metastases (seeding clones). Terminal clusters of each branch andseeding clones are annotated with the region name where the cluster is harbored and with the growth pattern of the region in the brackets. FIG. 3G: Representative hematoxylin and eosin staining images from CRUK0543. Rl, primary region with micropapillary pattern; R5 and R7, primary regions with papillary pattern; LN# 10, metastatic lymph node with acinar pattern; LN#7 and #8, metastatic lymph nodes with cribriform pattern.

[0055] FIGs. 4A-4G show the impact of tumor morphology upon site and risk of recurrence. FIG. 4A: Overview of the TRACERx 421 LU AD cohort with STAS assessment and preoperative ctDNA data (n = 136 patients), excluding the patients with synchronous primary lung cancers. Each column represents each patient. IMA, invasive mucinous adenocarcinoma. Tumors that did not relapse before death or the development of a new primary cancer are treated as no recurrence (No rec). FIG. 4B: Frequency of STAS and pre-operative ctDNA positivity across predominant subtypes (left) and grades of the predominant subtype (right) of primary tumor. FIGs. 4C-4D: Relapse-site specific (sub -distribution) hazard ratio (HR) for (FIG. 4C, left) the presence of micropapillary pattern and (FIG. 4C, right) the presence of solid and / or cribriform patterns (n = 215), and (FIG. 4D, left) the positivity of STAS and (FIG. 4D, right) pre-operative ctDNA detection in patients with pre-operative ctDNA data (n = 131). HR were adjusted for age, stage, pack-years, surgery type, and adjuvant therapy using Fine-Gray regression model. P < 0.05 are shown in red (unadjusted for FDR). FIG. 4E: Frequency of the relapse site (intra- and / or extra-thoracic), stratified by the positivity of STAS and pre-operative ctDNA detection. Tumors that did not relapse before death or the development of a new primary cancer are treated as no recurrence (No rec). The order of relapse site for each individual stratified bar graph (from bottom to top) depicted herein is as follows: extra-thoracic, intra & extra, intra-thoracic and no rec. FIG. 4F: Kaplan-Meier curves of disease-free survival, split by the positivity of STAS and pre-operative ctDNA detection. Hazard ratios were adjusted for age, stage, pack-years, surgery type, and adjuvant therapy. The number of patients in each group for every time point is indicated below the time point. FIG. 4G: Summary of the findings related to high-grade patterns, pathologic and genomic features, and relapse site. Factors with prognostic impact investigated in the study are highlighted in bold.

[0056] FIGs. 5A-5L show histopathological assessment of the TRACERx 421 LU AD cohort. FIG. 5A: Definition and categorization of LU AD growth patterns. FIG. 5B: Representative hematoxylin and eosin (H&E) photo of invasive mucinous adenocarcinoma (IMA). FIGs. 5C-5H: Representative H&E photos of lepidic (FIG. 5C), papillary (FIG. 5D), acinar (FIG. 5E), cribriform (FIG. 5F), micropapillary (FIG. 5G), and solid (FIG. 5H) pattern observed in LU AD. FIG. 51: Number of each predominant subtype tumor in the TRACERx 421 cohort. FIG. 5 J: Number of regions with growth pattern assessment. FIG. 5K: Schematic of histological assessment in the TRACERx study. Proportions of each subtype in the diagnostic slides were reported, and the predominant subtype was used to label each tumor. Multi-region sampling specimens were processed for whole exome sequencing, and each region was annotated with the representative growth pattern. FIG. 5L: Overview of TRACERx 421 LU AD cohort including invasive mucinous adenocarcinoma. Each column represents one tumor (n = 244). The proportion of each growth pattern based on diagnostic sectional area, genomic variables, and Ki-67 fraction by immunohistochemical staining are summaried. WGD, whole genome doubling; TMB, tumor mutational burden; ITH, intra-tumor heterogeneity; Mut, mutational; wGII, weighted genome instability index; FLOH, fraction of the genome subject to loss of heterozygosity; SCNA, somatic copy number alteration.

[0057] FIGs. 6A-6I show genomic correlates of LU AD predominant subtypes. FIGs. 6A- 6C: Frequency of truncal driver mutations (FIG. 6A), truncal driver gene somatic copy number alterations (SCNAs) (AMP, amplification) and whole genome doubling (WGD) (FIG. 6B), and chromosomal arm level SCNAs (gain / loss & LOH) (FIG. 6C) in LU AD predominant subtypes. Recurrent truncal alterations observed in more than 5% of the tumors in the cohort are shown. Asterisks represent the alterations observed in fewer than 10 tumors in both predominantly high- and low / mid-grade predominant tumors. Color scale represents the frequency of the alteration observed within each subtype. FIG. 6D: Across-genome plots showing the frequency of truncal and subclonal SCNAs of low / mid-grade predominant tumors (top) and high-grade predominant tumors (bottom). Within each tumor type, the proportion of patients with gains or amplifications (top) and loss / LOH events (bottom) for each chromosome are described. The black line indicates the total (namely the sum of truncal and subclonal) proportion of tumors with SCNAs;the yellow and grey lines or shades indicate the proportion of tumors with subclonal and truncal gains, respectively. FIG. 6E: The frequency of first and second WGD across LUAD predominant subtypes. The order of WGD clonality for each individual stratified bar graph (from bottom to top) depicted herein is as follows: truncal, subclonal and none. FIG. 6F: Number of genes with differential SCNA between high-grade and low / mid-grade predominant tumors. G2M checkpoint-related genes were not differentially gained in predominantly highgrade tumors (P = 0.20, chi-square goodness of fit test). The order of differential SCNA for each individual stratified bar graph (from top to bottom) depicted herein is as follows: ns, loss and gain. FIGs. 6G-6H: Comparison of stromal TIL scores (FIG. 6G) and PD-L1 expression on cancer cells measured by IHC staining (FIG. 6H) across LUAD predominant subtypes. Each predominant subtype was compared against all other subtype tumors. Centre line, median; box limits, upper and lower quartiles; whiskers, 1.5x interquartile range. P values were corrected for multiple testing according to Benjamini -Hochberg and asterisks indicate q value ranges * q < 0.05, ** q <0.01, *** q < 0.001, **** q < 0.0001. FIG. 61: Adjusted odds ratios for cancer cell PD-L1 positivity (> 1%) estimated by multivariable logistic regression model. Asterisks indicate type II ANOVA P value ranges * P < 0.05, ** P <0.01, *** P < 0.001. Statistically significant covariates are indicated in bold.

[0058] FIGs. 7A-7N show validation of SCNA analysis using orthogonal methods. FIGs. 7A-7B: Comparison of ITH metrics calculated using TRACERx analytical pipeline vs orthogonal methods. Comparison of fraction of the genome subject to loss of heterozygosity (FLOH), weighted genome instability index (wGII), and somatic copy number alteration intratumor heterogeneity (SCNA-ITH) by SCNA profiles generated by the TRACERx pipeline (based on ASCAT with additional multi-sample SCNA estimation approach) against SCNA profiles generated by Sequenza (FIG. 7A) and a comparison of % subclonal tumor mutational burden (TMB) using the TRACERx pipeline (clonality inferred by the modified version of PyClone48) vs % non-ubiquitous TMB (FIG. 7B). FIG. 7C: Correlation of genomic variables calculated using orthogonal methods and the proportion of high-grade patterns within each tumor. Color scale reflects Spearman’s rank correlation coefficient are used, or (rho). Correlation P values were corrected for multiple testing according to Benjamini -Hochberg (BH)and asterisks indicate q value ranges * q < 0.05, ** q <0.01, *** q < 0.001, **** q < 0.0001.FIGs. 7D-7F: Adjusted odds ratios of truncal genomic alterations associated with the predominance of high-grade patterns. Genomic alterations selected by the model simplification are shown when (FIG. 7D) truncal alterations observed in more than 10% of the tumors in the cohort are included in the analysis, or when (FIG. 7E) SCNA profiles generated by Sequenza are used, or when (FIG. 7F) wGII is added to the model shown in FIG. ID. Asterisks indicate type II ANOVA P value ranges * P < 0.05, ** P <0.01, *** P < 0.001. Color represents the type of genomic alteration. Statistically significant alterations are indicated in bold. FIGs. 7G-7H: Mutual exclusivity and cooccurrence of truncal driver gene alterations and chromosome arm somatic copy number alterations when (FIG. 7G) truncal alterations observed in more than 10% of the tumors in the cohort are included in the analysis, or when (FIG. 7H) SCNA profiles generated by Sequenzaare used. Color of the edge represents the relationship (mutual exclusivity vs co-occurrence) and the negative log of the q value (BH method) is represented in blue color scale in predominantly low / mid-grade tumors and red color scale in predominantly high-grade tumors. Relationships with q < 0.1 are shown and asterisks indicate q value ranges * q < 0.05, ** q <0.01 . Covariates in statistically significant relationships are indicated in bold. FIG. 71:. Comparison of wGII between tumors with and without co-occurrence of truncal loss / LOH of chromosome 3p and 3q in predominantly low / mid-grade tumors. P value was calculated using Wilcoxon rank sum test. FIGs. 7J-7K: Comparison of ploidy adjusted mean copy number of chromosomal arm and driver genes between high-grade and low / mid-grade predominant tumors, (FIG. 7 J) using SCNA profiles generated by Sequenza and (FIG. 7K) adding wGII to the regression model. Fixed effect coefficients of the linear mixed effect model with each tumor as a random effect are displayed on the x-axis, and the negative log of the q value (BH method) is displayed on the y-axis. Color represents the sign or the mean ploidy adjusted copy number, stratified with high-grade and low / mid-grade predominance. Data points with q value > 0.05 are colored in grey. Horizontal red dashed line represents q = 0.05. FIGs. 7L-7N: Genomic distance between regions calculated by LOH detected by Sequenza (FIG. 7L, n = 51 tumors) and genomic distance calculated by mutation (FIG. 7M) and LOH (FIG. 7N) only including tumor regions with purity > 0.4 (n = 30 tumors). Each point represents a distance between a pair ofregions in a tumor. Tumors with regions containing both different subtype pair(s) and same subtype pair(s) are included in the analysis. Centre line, median; box limits, upper and lower quartiles; whiskers, 1.5 x interquartile range. P values were calculated using a linear mixed effects model, with each tumor as a random effect.

[0059] FIGs. 8A-8E demonstrate inference of ancestor-like and descendant-like regional pairs in the primary tumor. FIG. 8A: Comparison of tumor mutational burden (TMB) between ancestor-like and descendant-like regions. Each line represents an ancestor-descendant-like regional pair. Each point represents one region and the plotted points were duplicated for regions associated with multiple ancestor-descendant-like pairs within a tumor. To assess the mutational burden shared in the majority of the cancer cells in the region, mutations with estimated cancer cell fraction > 95% were counted (TMB CCF95). Enrichment of higher TMB in descendant-like regions compared with the paired ancestor-like regions was evaluated by permutation test (1000 permutations, randomizing TMB within each tumor, Monte-Carlo procedure). FIG. 8B: Comparison of growth pattern by grades (left) and by the six growth patterns (right) between inferred ancestral-like and descendant-like regions. Tumors with single grades are included in the analysis. Color represents the transition of grade from ancestral-like to descendant-like region. Empirical P value was calculated using a permutation test (1000 permutations, randomizing growth patterns within each tumor, Monte-Carlo procedure). FIG. 8C: Comparison of regional growth pattern grade in ancestor-descendant-like pairs, inferred by various cutoffs of private LOH branch length proportional to the trunk (shared LOH). All combinations of cutoff for ancestor-like and descendant-like inference shown in the figure yielded empirical P value < 0.05 (1000 permutations, Monte-Carlo procedure) when the enrichment of lower-to-higher grade transition (upward transition) was tested. P values were not adjusted for the multiple comparisons shown in this panel. FIG. 8D: Comparison of regional pattern grade in ancestor-descendant-like pairs, inferred by LOH profile generated by Sequenza. FIG. 8E: Comparison of regional pattern grade in ancestor-descendant-like pairs, inferred by both LOH profile and mutational profile (CCF > 95%).

[0060] FIGs. 9A-9I show characterization of purely (homogenously) solid tumors. FIG. 9A: Comparison of G2M checkpoint gene expression in solid-pattern regions within purely solid tumor and mixed pattern tumor as defined by both diagnostic and regional growth pattern assessment. Centre line, median; box limits, upper and lower quartiles; whiskers, 1.5x interquartile range. P values were calculated using a linear mixed effects model, with each tumor as a random effect. FIGs. 9B-9C: Proportion of tumor which are purely solid, mixed pattern with solid component, and without any solid component, compared (FIG. 9B) between tumor with and without truncal gain of chromosome arm 3q and (FIG. 9C) across the tumor stratified by truncal SMARCA4 mutation and / or loss of heterozygosity (LOH). FIGs. 9D-9E. Comparison of the frequency of truncal copy number gain of (FIG. 9D) chromosome arm 3q and (FIG. 9E) arm or focal 3q (3q21.3-3q29) between mixed pattern tumor with solid component and purely solid tumor. FIG. 9F: Comparison of the frequency of truncal SMARCA4 mutation and LOH between mixed pattern tumor with solid component and purely solid tumor. FIGs. 9G-9H: Comparison of the frequency of copy number gain of (FIG. 9G) chromosome arm 3q and (FIG. 9H) arm or focal 3q (3q21.3-3q29) between mixed pattern tumor with solid component and purely solid tumor using somatic copy number alteration (SCNA) profiles generated by Sequenza. FIG. 9D: The order of chr 3q gain for each individual stratified bar graph (from top to bottom) depicted herein is as follows: no gain, subclonal and truncal. FIG. 9E: The order of arm / focal 3q gain for each individual stratified bar graph (from top to bottom) depicted herein is as follows: no gain, subclonal and truncal. FIG. 9F: The order of SMARC4 mut / LOH for each individual stratified bar graph (from top to bottom) depicted herein is as follows: no alt, subclonal and truncal. FIG. 9G: The order of chr 3q gain for each individual stratified bar graph (from top to bottom) depicted herein is as follows: no gain, subclonal and clonal. FIG. 9H: The order of arm / focal 3q gain for each individual stratified bar graph (from top to bottom) depicted herein is as follows: no gain, subclonal and clonal. FIG. 91: Comparison of the frequency of SMARCA4 mutation and LOH between mixed pattern tumor with solid component and purely solid tumor using SCNA profiles generated by Sequenza. FIG. 91: The order of SMARC4 mut / LOH for each individual stratified bar graph (from top to bottom) depicted herein is as follows: no alt, subclonal and clonal.

[0061] FIGs. 10A-10G show analysis of morphology and genomics in metastasis samples.FIG. 10A: Schematic of primary and secondary lung tumors in CRUK0296. Phylogenetic analysis confirmed the contralateral lung lesion to be a metastasis from the resected primary tumor three years ago. Tumor spread through air spaces (STAS) was positive in the primary tumor. FIG. 10B: Phylogenetic tree of a case having lung metastasis with pure lepidic appearance (CRUK0296). Driver mutations are shown in the figure at the concordant mutational cluster. Regional growth pattern is indicated in brackets; n.a., not available. FIG. 10C: Representative hematoxylin and eosin (H&E) slide of a primary tumor of CRUK0296 showing tumor border (arrowheads) and STAS (arrow). FIG. 10D: Representative H&E slide of metastasis tumor in the contralateral lung of CRUK0296, which showed a pure lepidic pattern. FIG. 10E: Characteristics of five patients having lung recurrence samples sequenced and one patient having an intrapulmonary metastasis resected and sequenced at the time of primary surgery. All six patients showed positive STAS in the primary tumors and phylogenetic analysis revealed late metastatic divergence. FIG. 10F: Proportion of the timing of seeding clone divergence across predominant subtypes of primary tumors. The order of timing for each individual stratified bar graph (from top to bottom) depicted herein is as follows: early, and late. FIG. 10G: Frequency of late or early divergence of the metastatic clone compared between tumors with and without STAS. The order of timing for each individual stratified bar graph (from top to bottom) depicted herein is as follows: early, and late.

[0062] FIGs. 11A-11I show characterization of tumors with STAS and pre-operative ctDNA shedding. FIG. 11 A: Overview of the TRACERx421 LU AD cohort, ordered by the positivity of STAS, pre-operative ctDNA detection, and the site of the relapse (n = 223). Patients with synchronous primary lung cancers were excluded. Colloid and fetal adenocarcinomas are included (predominant sutbype = Other). Each column represents each patient. IMA, invasive mucinous adenocarcinoma; LVI, lymphovascular invasion; PL, pleural invasion. Tumors that did not relapse before death or the development of a new primary cancer are treated as no recurrence (No rec). FIG. 11B: Kaplan-Meier curve of disease-free survival, comparing STAS present vs absent. Numbers at risk are described at the bottom. For the patients with multiple tumors, only patients having LU AD as the most advanced tumor were included in the analysis.Hazard ratio (HR) adjusted for age, stage, pack-years, surgery type, and adjuvant therapy is shown. FIG. 11C: STAS positivity across predominant subtypes of the primary tumor. The order of STAS presence for each individual stratified bar graph (from top to bottom) depicted herein is as follows: absent, and present. FIG. 11D: Histopathological features associated with STAS positivity (left) and pre-operative ctDNA detection (right). Negative log of the q values (Benjamini -Hochberg method) in univariable logistic regression analyses are presented. Vertical dotted lines represent q = 0.05, and variables with q < 0.05 are presented in points with colors which represent the direction of the correlation. FIG. HE: Pre-operative ctDNA positivity across predominant subtypes of the primary tumor. FIG. 11F: Kaplan-Meier curve of disease- free survival, comparing patients with predominantly high-grade tumors vs low / mid-grade tumors. Numbers at risk are described at the bottom. Hazard ratio (HR) adjusted for age, stage, pack-years, surgery type, and adjuvant therapy is shown. FIG. 11G: Frequency of the relapse site (intra- and / or extra-thoracic) across predominant subtypes (left) and grades of the predominant subtype (right) of primary tumor. Tumors that did not relapse before death or the development of a new primary cancer are treated as no recurrence (No rec). FIG. 11H: Relapsesite specific (subdistribution) hazard ratio for predominantly high-grade tumors compared with low / mid-grade tumors in all LUADs, adjusted for age, stage, pack-years, surgery type, and adjuvant therapy (n = 185). P < 0.05 are described in red (unadjusted for FDR). FIG. HI: Positivity of necrosis across predominant subtypes of the primary tumor.

[0063] FIGs. 12A-12C: show genomic and transcriptomic analyses of STAS in LU AD. FIG. 12A: Frequency of driver mutations in 10 canonical oncogenic signalling pathways49 in STAS present and absent tumors. P values (Fisher’s exact test) were corrected for multiple testing according to Benjamini-Hochberg (BH) and the asterisk indicates q value range * q < 0.05. FIG. 12B: Comparison of CTNNB1 gene expression (variance stabilisation normalised count) between STAS positive and negative tumors. Centre line, median; box limits, upper and lower quartiles; whiskers, 1.5x interquartile range. P value was calculated using a linear mixed effect model, with each tumor as a random effect. FIG. 12C: Gene set enrichment analysis of Hallmark gene sets between STAS positive and negative tumors. Normalized enrichment scoreis displayed on the x-axis and indicates the enrichment for a given gene set. None of the gene sets showed q < 0.25 (BH method).

[0064] FIGs. 13A-13I show impact of STAS, pre-operative ctDNA positivity, and necrosis on sites and risk of recurrence. FIGs. 13A-13B: Frequency of the relapse site (intra- and / or extra-thoracic), stratified by the positivity of STAS and pre-operative ctDNA detection. Preoperative ctDNA data were based on (FIG. 13A) the assay previously reported by Abbosh et. al9 (TxlOO cohort) and (FIG. 13B) the assay reported in our companion manuscripts 0 (Tx421 cohort), including 7 patients who underwent both assays in each cohort. Tumors that did not relapse before death or the development of a new primary cancer are treated as no recurrence (No rec). FIG. 13C: Positivity of STAS and pre-operative ctDNA detection are incorporated with other tumor and clinical characteristics in a multivariable Cox proportional hazards model (disease-free survival). Hazard ratios (HRs) of each variable with 95% confidence intervals (Cis) are shown on the horizontal axis. FIGs. 13D-13E: Kaplan-Meier curves of disease-free survival, split by the positivity of STAS and pre-operative ctDNA detection in (FIG. 13D) stage I patients and (FIG. 13E) stage II & III patients. HRs were adjusted for age, stage, pack-years, and adjuvant therapy. Surgery type was also added as a covariate for stage I patients but not for stage II & III patients, because only 1 patient underwent sublobar resection in stage II & III patients. The number of patients in each group for every time point is indicated below the time point. FIG. 13F: Frequency of the relapse site (intra- and / or extra-thoracic), stratified by the presence of STAS and necrosis in all LUADs. The order of the relapse site for each individual stratified bar graph (from top to bottom) depicted herein is as follows: No rec, intra-thoracic, intra & extra, and extra-thoracic. FIG. 13G: Relapse-site specific (sub distribution) HR for positivity of necrosis in all LUADs, adjusted for age, stage, pack-years, surgery type, and adjuvant therapy (n=211). P < 0.05 are shown in red (unadjusted for FDR). FIG. 13H: Positivity of STAS and necrosis are incorporated with other tumor and clinical characteristics in a multivariable Cox proportional hazards model for disease-free survival. HRs of each variable with 95% Cis are shown. FIG. 131: Kaplan-Meier curve of disease-free survival, split by the positivity of STAS and the presence of necrosis. HRs were adjusted for age, stage, pack-years,surgery type, and adjuvant therapy. The number of patients in each group for every time point is indicated below the time point.

[0065] FIGs. 14A-14C show external validation of the impact of STAS and necrosis on disease-free survival. FIG. 14A: Summary of patient demographics and clinical characteristics of the Memorial Sloan Kettering Cancer Center cohort (n = 712). FIG. 14B: Positivity of STAS and necrosis are incorporated with other tumor and clinical characteristics in a multivariable Cox proportional hazards model of disease-free survival. FIG. 14C: Kaplan-Meier curve of disease- free survival, split by the positivity of STAS and the presence of necrosis (n = 712). Hazard ratios were adjusted for age, stage, pack-years, surgery type, and adjuvant therapy. The number of patients in each group for every time point is indicated below the time point.

[0066] FIG. 15A is a block diagram depicting an embodiment of a network environment comprising a client device in communication with server device.

[0067] FIG. 15B is a block diagram depicting a cloud computing environment comprising client device in communication with cloud service providers.

[0068] FIGs. 15C and 15D are block diagrams depicting embodiments of computing devices useful in connection with the methods and systems described herein.

[0069] FIG. 16 depicts a system that includes a computing device and a sample processing system according to various potential embodiments.DETAILED DESCRIPTION

[0070] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology. It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0071] Lung adenocarcinomas (LUADs) display a broad histological spectrum from low- grade lepidic tumors through to mid-grade acinar and papillary and high-grade solid, cribriform and micropapillary tumors. Little is known as to how morphology reflects tumor evolutionary history and disease progression. Accordingly, approaches to enhance the overall benefit of lobectomy vs. limited resection in lung cancer patients will be contingent on improved methods for predicting the risk of recurrence and metastases. The present disclosure demonstrates that ctDNA and the aerogenous spread of tumor cells in the peritumoral lung parenchyma beyond the edge of the tumor (a.k.a., “spread through air spaces” or STAS) are useful biomarkers for accurately predicting the risk of recurrence and metastases in lung cancer patients. The presence of solid / cribriform pattern was associated with pre-operative ctDNA detection and increased risk of extra-thoracic recurrence, whilst micropapillary pattern was associated with STAS positivity and intra-thoracic recurrence. As disclosed in the Examples herein, patients who are positive for both STAS and pre-operative ctDNA have significantly poor outcomes. As a combined measure, both STAS positivity and ctDNA detection thus have the potential to predict outcome at resection.Definitions

[0072] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. For example, reference to “a cell” includes a combination of two or more cells, and the like. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, analytical chemistry and nucleic acid chemistry and hybridization described below are those well-known and commonly employed in the art.

[0073] As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1%, 5%, or 10% in either direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context (except where such number would be less than 0% or exceed 100% of a possible value).

[0074] The term “adapter” refers to a short, chemically synthesized, nucleic acid sequence which can be used to ligate to the end of a nucleic acid sequence in order to facilitate attachment to another molecule. The adapter can be single-stranded or double-stranded. An adapter can incorporate a short (typically less than 50 base pairs) sequence useful for PCR amplification or sequencing.

[0075] As used herein, the “administration” of an agent or drug to a subject includes any route of introducing or delivering to a subject a compound to perform its intended function. Administration can be carried out by any suitable route, including but not limited to, orally, intranasally, parenterally (intravenously, intramuscularly, intraperitoneally, or subcutaneously), rectally, intrathecally, intratumorally or topically. Administration includes self-administration and the administration by another.

[0076] As used herein, an “alteration” of a gene or gene product (e.g., a marker gene or gene product) refers to the presence of a mutation or mutations within the gene or gene product, e.g., a mutation, which affects the quantity or activity of the gene or gene product, as compared to the normal or wild-type gene. The genetic alteration can result in changes in the quantity, structure, and / or activity of the gene or gene product in a cancer tissue or cancer cell, as compared to its quantity, structure, and / or activity, in a normal or healthy tissue or cell (e.g., a control). For example, an alteration which is predictive of recurrence and metastases in lung cancer can have an altered nucleotide sequence (e.g., a mutation), amino acid sequence, chromosomal translocation, intra-chromosomal inversion, copy number, expression level, protein level, protein activity, in a cancer tissue or cancer cell, as compared to a normal, healthy tissue or cell.Exemplary mutations include, but are not limited to, point mutations (e.g., silent, missense, or nonsense), deletions, insertions, inversions, linking mutations, duplications, translocations, inter- and intra-chromosomal rearrangements. Mutations can be present in the coding or non-coding region of the gene.

[0077] The terms “cancer” or “tumor” are used interchangeably and refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristicmorphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be a non-tumorigenic cancer cell. As used herein, the term “cancer” includes premalignant, as well as malignant cancers. In some embodiments, the cancer is lung cancer.

[0078] As used herein, a "control" is an alternative sample used in an experiment for comparison purpose. A control can be "positive" or "negative." For example, where the purpose of the experiment is to determine a correlation of the efficacy of a therapeutic agent for the treatment for a particular type of disease, a positive control (a compound or composition known to exhibit the desired therapeutic effect) and a negative control (a subject or a sample that does not receive the therapy or receives a placebo) are typically employed.

[0079] As used herein, a “deletion” refers to a mutation (or a genetic alteration) in which part of a DNA sequence at a chromosome location is absent or lost compared to that observed in a reference genome. A deletion may occur within a gene or may encompass one or more genes. A “homozygous deletion” refers to the loss of both alleles of a gene within a genome. A homozygous deletion may comprise a partial or complete loss of each copy (maternal and paternal) of the gene sequence.

[0080] “Detecting” as used herein refers to determining the presence of a mutation or alteration in a nucleic acid of interest in a sample. Detection does not require the method to provide 100% sensitivity. Analysis of nucleic acid markers can be performed using techniques known in the art including, but not limited to, sequence analysis, and electrophoretic analysis. Non-limiting examples of sequence analysis include Maxam-Gilbert sequencing, Sanger sequencing, capillary array DNA sequencing, thermal cycle sequencing (Sears etal., Biotechniques, 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al., Methods Mol. Cell Biol, 3:39-42 (1992)), sequencing with mass spectrometry such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et l., Nat. Biotechnol, 16:381-384 (1998)), and sequencing by hybridization. Chee et al., Science, 274:610- 614 (1996); Drmanac et al., Science, 260: 1649-1652 (1993); Drmanac et al., Nat. Biotechnol, 16:54-58 (1998). Non-limiting examples of electrophoretic analysis include slab gelelectrophoresis such as agarose or polyacrylamide gel electrophoresis, capillary electrophoresis, and denaturing gradient gel electrophoresis. Additionally, next generation sequencing methods can be performed using commercially available kits and instruments from companies such as the Life Technologies / Ion Torrent PGM or Proton, the Illumina HiSEQ or MiSEQ, and the Roche / 454 next generation sequencing system.

[0081] As used herein, the term “effective amount” refers to a quantity sufficient to achieve a desired therapeutic and / or prophylactic effect, e.g., an amount which results in the prevention of, or a decrease in a disease or condition described herein or one or more signs or symptoms associated with a disease or condition described herein. In the context of therapeutic or prophylactic applications, the amount of a composition administered to the subject will vary depending on the composition, the degree, type, and severity of the disease and on the characteristics of the individual, such as general health, age, sex, body weight and tolerance to drugs. The skilled artisan will be able to determine appropriate dosages depending on these and other factors. The compositions can also be administered in combination with one or more additional therapeutic compounds. In the methods described herein, the therapeutic compositions may be administered to a subject having one or more signs or symptoms of a disease or condition described herein. As used herein, a "therapeutically effective amount" of a composition refers to composition levels in which the physiological effects of a disease or condition are ameliorated or eliminated. A therapeutically effective amount can be given in one or more administrations.

[0082] As used herein, “expression” includes one or more of the following: transcription of the gene into precursor mRNA; splicing and other processing of the precursor mRNA to produce mature mRNA; mRNA stability; translation of the mature mRNA into protein (including codon usage and tRNA availability); and glycosylation and / or other modifications of the translation product, if required for proper expression and function.

[0083] “Gene” as used herein refers to a DNA sequence that comprises regulatory and coding sequences necessary for the production of an RNA, which may have a non-coding function (e.g., a ribosomal or transfer RNA) or which may include a polypeptide or a polypeptideprecursor. The RNA or polypeptide may be encoded by a full length coding sequence or by any portion of the coding sequence so long as the desired activity or function is retained. Although a sequence of the nucleic acids may be shown in the form of DNA, a person of ordinary skill in the art recognizes that the corresponding RNA sequence will have a similar sequence with the thymine being replaced by uracil, i.e., "T" is replaced with "U."

[0084] ‘Next-generation sequencing or NGS” as used herein, refers to any sequencing method that determines the nucleotide sequence of either individual nucleic acid molecules (e.g., in single molecule sequencing) or clonally expanded proxies for individual nucleic acid molecules in a high throughput parallel fashion (e.g., greater than 103, 104, 105or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of the nucleic acid species in the library can be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment.Next generation sequencing methods are known in the art, and are described, e.g., in Metzker, M. Nature Biotechnology Reviews 11 : 31 -46 (2010).

[0085] As used herein, a “sample” refers to a substance that is being assayed for the presence of a mutation in a nucleic acid of interest. Processing methods to release or otherwise make available a nucleic acid for detection are well known in the art and may include steps of nucleic acid manipulation. A biological sample may be a body fluid or a tissue sample. In some cases, a biological sample may consist of or comprise blood, plasma, sera, urine, feces, epidermal sample, vaginal sample, skin sample, cheek swab, sperm, amniotic fluid, cultured cells, bone marrow sample, tumor biopsies, aspirate and / or chorionic villi, cultured cells, and the like.Fresh, fixed or frozen tissues may also be used. In one embodiment, the sample is preserved as a frozen sample or as formaldehyde- or paraformaldehyde-fixed paraffin-embedded (FFPE) tissue preparation. For example, the sample can be embedded in a matrix, e.g., an FFPE block or a frozen sample. Whole blood samples of about 0.5 to 5 ml collected with EDTA, ACD or heparin as anti-coagulant are suitable.

[0086] As used herein, “spread through air spaces” or “STAS” refers to the aerogenous spread of tumor cells in the peritumoral lung parenchyma beyond the edge of the lung tumor.

[0087] As used herein, the terms “subject”, “patient”, or “individual” can be an individual organism, a vertebrate, a mammal, or a human. In some embodiments, the subject, patient or individual is a human.

[0088] “Treating” or “treatment” as used herein covers the treatment of a disease or disorder described herein, in a subject, such as a human, and includes: (i) inhibiting a disease or disorder, i.e., arresting its development; (ii) relieving a disease or disorder, i.e., causing regression of the disorder; (iii) slowing progression of the disorder; and / or (iv) inhibiting, relieving, or slowing progression of one or more symptoms of the disease or disorder. In some embodiments, treatment means that the symptoms associated with the disease are, e.g., alleviated, reduced, cured, or placed in a state of remission.

[0089] The terms “variant allele fraction,” “VAF,” “mutant allele fraction” or “MAF” refer to fractions of a mutant allele over the total number of mutant (alternate allele) plus wild-type alleles (reference allele). ctDNA VAF represents %ctDNA alteration reported as percentage and computed as the number of mutated DNA molecules divided by the total number (mutated plus wild-type) of DNA fragments at that allele. Most of the cell-free DNA is wild-type (germline); therefore, the median VAF of somatic alterations is <0.5%.

[0090] It is also to be appreciated that the various modes of treatment of disorders as described herein are intended to mean “substantial,” which includes total but also less than total treatment, and wherein some biologically or medically relevant result is achieved. The treatment may be a continuous prolonged treatment for a chronic disease or a single, or few time administrations for the treatment of an acute condition.Methods for Detecting Polynucleotides Associated with Elevated Risk of Recurrence and Metastases

[0091] Polynucleotides associated with elevated risk of recurrence and metastases may be detected by a variety of methods known in the art. Non-limiting examples of detection methods are described below. The detection assays in the methods of the present technology may include purified or isolated DNA (genomic or cDNA), RNA or protein or the detection step may beperformed directly from a biological sample without the need for further DNA, RNA or protein puri fi cati on / i sol ati on .Nucleic Acid Amplification and / or Detection

[0092] Polynucleotides associated with elevated risk of recurrence and metastases can be detected by the use of nucleic acid amplification techniques that are well known in the art. The starting material may be genomic DNA, cDNA, RNA, ctDNA, cfDNA, or mRNA. Nucleic acid amplification can be linear or exponential. Specific variants or mutations may be detected by the use of amplification methods with the aid of oligonucleotide primers or probes designed to interact with or hybridize to a particular target sequence in a specific manner, thus amplifying only the target variant.

[0093] Non-limiting examples of nucleic acid amplification techniques include polymerase chain reaction (PCR), real-time quantitative PCR (qPCR), digital PCR (dPCR), reverse transcriptase polymerase chain reaction (RT-PCR), nested PCR, ligase chain reaction (see Abravaya, K. et al., Nucleic Acids Res. (1995), 23:675-682), branched DNA signal amplification (see Urdea, M. S. etal., AIDS (1993), 7(suppl 2): S 11- S14), amplifiable RNA reporters, Q-beta replication, transcription-based amplification, boomerang DNA amplification, strand displacement activation, cycling probe technology, isothermal nucleic acid sequence based amplification (NASBA) (see Kievits, T. et al., J Virological Methods (1991), 35:273-286), Invader Technology, next-generation sequencing technology or other sequence replication assays or signal amplification assays.

[0094] Primers'. Oligonucleotide primers for use in amplification methods can be designed according to general guidance well known in the art as described herein, as well as with specific requirements as described herein for each step of the particular methods described. In some embodiments, oligonucleotide primers for cDNA synthesis and PCR are 10 to 100 nucleotides in length, preferably between about 15 and about 60 nucleotides in length, more preferably 25 and about 50 nucleotides in length, and most preferably between about 25 and about 40 nucleotides in length.

[0095] Tm of a polynucleotide affects its hybridization to another polynucleotide (e.g., the annealing of an oligonucleotide primer to a template polynucleotide). In certain embodiments of the disclosed methods, the oligonucleotide primer used in various steps selectively hybridizes to a target template or polynucleotides derived from the target template (i.e first and second strand cDNAs and amplified products). Typically, selective hybridization occurs when two polynucleotide sequences are substantially complementary (at least about 65% complementary over a stretch of at least 14 to 25 nucleotides, preferably at least about 75%, more preferably at least about 90% complementary). See Kanehisa, M., Polynucleotides Res. (1984), 12:203, incorporated herein by reference. As a result, it is expected that a certain degree of mismatch at the priming site is tolerated. Such mismatch may be small, such as a mono-, di- or tri-nucleotide. In certain embodiments, 100% complementarity exists.

[0096] Probes. Probes are capable of hybridizing to at least a portion of the nucleic acid of interest or a reference nucleic acid ( e., wild-type sequence). Probes may be an oligonucleotide, artificial chromosome, fragmented artificial chromosome, genomic nucleic acid, fragmented genomic nucleic acid, RNA, recombinant nucleic acid, fragmented recombinant nucleic acid, peptide nucleic acid (PNA), locked nucleic acid, oligomer of cyclic heterocycles, or conjugates of nucleic acid. Probes may be used for detecting and / or capturing / purifying a nucleic acid of interest.

[0097] Typically, probes can be about 10 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 50 nucleotides, about 60 nucleotides, about 75 nucleotides, or about 100 nucleotides long. However, longer probes are possible. Longer probes can be about 200 nucleotides, about 300 nucleotides, about 400 nucleotides, about 500 nucleotides, about 750 nucleotides, about 1,000 nucleotides, about 1,500 nucleotides, about 2,000 nucleotides, about 2,500 nucleotides, about 3,000 nucleotides, about 3,500 nucleotides, about 4,000 nucleotides, about 5,000 nucleotides, about 7,500 nucleotides, or about 10,000 nucleotides long.

[0098] Probes may also include a detectable label or a plurality of detectable labels. The detectable label associated with the probe can generate a detectable signal directly. Additionally,the detectable label associated with the probe can be detected indirectly using a reagent, wherein the reagent includes a detectable label, and binds to the label associated with the probe.

[0099] In some embodiments, detectably labeled probes can be used in hybridization assays including, but not limited to Northern blots, Southern blots, microarray, dot or slot blots, and in situ hybridization assays such as fluorescent in situ hybridization (FISH) to detect a target nucleic acid sequence within a biological sample. Certain embodiments may employ hybridization methods for measuring expression of a polynucleotide gene product, such as mRNA. Methods for conducting polynucleotide hybridization assays have been well developed in the art. Hybridization assay procedures and conditions will vary depending on the application and are selected in accordance with the general binding methods known including those referred to in: Maniatis et al. Molecular Cloning: A Laboratory Manual (2nd Ed. Cold Spring Harbor, N.Y., 1989); Berger and Kimmel Methods in Enzymology, Vol. 152, Guide to Molecular Cloning Techniques (Academic Press, Inc., San Diego, Calif, 1987); Young and Davis, PNAS. 80: 1194 (1983).

[0100] Detectably labeled probes can also be used to monitor the amplification of a target nucleic acid sequence. In some embodiments, detectably labeled probes present in an amplification reaction are suitable for monitoring the amount of amplicon(s) produced as a function of time. Examples of such probes include, but are not limited to, the 5'- exonuclease assay (TAQMAN® probes described herein (see also U.S. Pat. No. 5,538,848) various stem-loop molecular beacons (see for example, U.S. Pat. Nos. 6,103,476 and 5,925,517 and Tyagi and Kramer, 1996, Nature Biotechnology 14:303- 308), stemless or linear beacons (see, e.g., WO 99 / 21881), PNA Molecular Beacons™ (see, e.g., U.S. Pat. Nos. 6,355,421 and 6,593,091), linear PNA beacons (see, for example, Kubista et al., 2001, SPIE 4264:53-58), non-FRET probes (see, for example, U.S. Pat. No. 6,150,097), Sunrise® / Amplifluor™ probes (U.S. Pat. No. 6,548,250), stem-loop and duplex Scorpion probes (Solinas et al., 2001, Nucleic Acids Research 29:E96 and U.S. Pat. No. 6,589,743), bulge loop probes (U.S. Pat. No. 6,590,091), pseudo knot probes (U.S. Pat. No. 6,589,250), cyclicons (U.S. Pat. No. 6,383,752), MGB Eclipse™ probe (Epoch Biosciences), hairpin probes (U.S. Pat. No. 6,596,490), peptide nucleic acid (PNA) light-upprobes, self-assembled nanoparticle probes, and ferrocene-modified probes described, for example, in U.S. Pat. No. 6,485,901 ; Mhlanga etal., 2001, Methods 25:463-471 ; Whitcombe et al., 1999, Nature Biotechnology. 17:804-807; Isacsson et al., 2000, Molecular Cell Probes. 14:321-328; Svanvik et al., 2000, Anal Biochem. 281 :26-35; Wolffs et al., 2001 , Biotechniques 766:769-771 ; Tsourkas et al., 2002, Nucleic Acids Research. 30:4208-4215; Riccelli et l., 2002, Nucleic Acids Research 30:4088-4093; Zhang etal., 2002 Shanghai. 34:329-332; Maxwell et al., 2002, J. Am. Chem. Soc. 124:9606-9612; Broude et al., 2002, Trends Biotechnol. 20:249- 56; Huang et al., 2002, Chem. Res. Toxicol. 15: 118- 126; and Yu et al., 2001, J. Am. Chem. Soc 14: 11155-11161.

[0101] In some embodiments, the detectable label is a fluorophore. Suitable fluorescent moi eties include but are not limited to the following fluorophores working individually or in combination: 4-acetamido-4'-isothiocyanatostilbene- 2,2'disulfonic acid; acridine and derivatives: acridine, acridine isothiocyanate; Alexa Fluors: Alexa Fluor® 350, Alexa Fluor® 488, Alexa Fluor® 546, Alexa Fluor® 555, Alexa Fluor® 568, Alexa Fluor® 594, Alexa Fluor® 647 (Molecular Probes); 5-(2- aminoethyl)aminonaphthalene-l -sulfonic acid (EDANS); 4- amino-N-[3- vinylsulfonyl)phenyl]naphthalimide-3,5 disulfonate (Lucifer Yellow VS); N-(4- anilino-1- naphthyl)maleimide; anthranilamide; Black Hole Quencher™ (BHQ™) dyes (biosearch Technologies); BODIPY dyes: BODIPY® R-6G, BOPIPY® 530 / 550, BODIPY® FL; Brilliant Yellow; coumarin and derivatives: coumarin, 7-amino-4-methylcoumarin (AMC, Coumarin 120),7-amino-4-trifluoromethylcouluarin (Coumarin 151); Cy2®, Cy3®, Cy3.5®, Cy5®, Cy5.5®; cyanosine; 4',6-diaminidino-2-phenylindole (DAPI); 5', 5"-dibromopyrogallol- sulfonephthalein (Bromopyrogallol Red); 7-diethylamino-3-(4'-isothiocyanatophenyl)-4- methylcoumarin; diethylenetriamine pentaacetate; 4,4'-diisothiocyanatodihydro-stilbene-2,2'- disulfonic acid; 4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid; 5- [dimethylamino]naphthalene-l -sulfonyl chloride (DNS, dansyl chloride); 4-(4'- dimethylaminophenylazo)benzoic acid (DABCYL); 4-dimethylaminophenylazophenyl-4'- isothiocyanate (DABITC); Eclipse™ (Epoch Biosciences Inc.); eosin and derivatives: eosin, eosin isothiocyanate; erythrosin and derivatives: erythrosin B, erythrosin isothiocyanate; ethidium; fluorescein and derivatives: 5-carboxyfluorescein (FAM), 5-(4,6-dichlorotriazin-2-yl)amino fluorescein (DTAF), 2',7'-dimethoxy-4'5'-dichloro-6-carboxyfluorescein (JOE), fluorescein, fluorescein isothiocyanate (FITC), hexachloro-6-carboxyfluorescein (HEX), QFITC (XRITC), tetrachlorofluorescem (TET); fiuorescamine; IR144; IR1446; lanthamide phosphors; Malachite Green isothiocyanate; 4-methylumbelliferone; ortho cresolphthalein; nitrotyrosine; pararosaniline; Phenol Red; B-phycoerythrin, R-phycoerythrin; allophycocyanin; o- phthal dialdehyde; Oregon Green®; propidium iodide; pyrene and derivatives: pyrene, pyrene butyrate, succinimidyl 1 -pyrene butyrate; QSY® 7; QSY® 9; QSY® 21; QSY® 35 (Molecular Probes); Reactive Red 4 (Cibacron®Brilliant Red 3B-A); rhodamine and derivatives: 6-carboxy- X-rhodamine (ROX), 6-carboxyrhodamine (R6G), lissamine rhodamine B sulfonyl chloride, rhodamine (Rhod), rhodamine B, rhodamine 123, rhodamine green, rhodamine X isothiocyanate, riboflavin, rosolic acid, sulforhodamine B, sulforhodamine 101, sulfonyl chloride derivative of sulforhodamine 101 (Texas Red); terbium chelate derivatives; N,N,N',N'-tetramethyl-6- carboxyrhodamine (TAMRA); tetramethyl rhodamine; tetramethyl rhodamine isothiocyanate (TRITC); and VIC®. Detector probes can also comprise sulfonate derivatives of fluorescenin dyes with S03 instead of the carboxylate group, phosphoramidite forms of fluorescein, phosphorami dite forms of CY 5 (commercially available for example from Amersham).

[0102] Detectably labeled probes can also include quenchers, including without limitation black hole quenchers (Biosearch), Iowa Black (IDT), QSY quencher (Molecular Probes), and Dabsyl and Dabcel sulfonate / carboxylate Quenchers (Epoch).

[0103] Detectably labeled probes can also include two probes, wherein for example a fluorophore is on one probe, and a quencher is on the other probe, wherein hybridization of the two probes together on a target quenches the signal, or wherein hybridization on the target alters the signal signature via a change in fluorescence.

[0104] In some embodiments, interchelating labels such as ethidium bromide, SYBR® Green I (Molecular Probes), and PicoGreen® (Molecular Probes) are used, thereby allowing visualization in real-time, or at the end point, of an amplification product in the absence of a detector probe. In some embodiments, real-time visualization may involve the use of both an intercalating detector probe and a sequence-based detector probe. In some embodiments, thedetector probe is at least partially quenched when not hybridized to a complementary sequence in the amplification reaction, and is at least partially unquenched when hybridized to a complementary sequence in the amplification reaction.

[0105] In some embodiments, the amount of probe that gives a fluorescent signal in response to an excited light typically relates to the amount of nucleic acid produced in the amplification reaction. Thus, in some embodiments, the amount of fluorescent signal is related to the amount of product created in the amplification reaction. In such embodiments, one can therefore measure the amount of amplification product by measuring the intensity of the fluorescent signal from the fluorescent indicator.

[0106] Primers or probes may be designed to selectively hybridize to any portion of a nucleic acid sequence encoding a polypeptide that is over-expressed in lung cancer patients. Exemplary nucleic acid sequences include, but are not limited to, AURKB, BUB IB, CENPE, NCAPG, TROAP, FAM83D, KIF4A, NDC80, POLQ, TPX2, KIF23, AURKA, TACC3, CKAP2L and PTTG1. Primers or probes can be designed so that they hybridize under stringent conditions to mutant nucleotide sequences of a polypeptide that is over-expressed in lung cancer patients, but not to the respective wild-type nucleotide sequences. Primers or probes can also be prepared that are complementary and specific for the wild-type nucleotide sequences of a polypeptide that is over-expressed in lung cancer patients, but not to any of the corresponding mutant nucleotide sequences. In some embodiments, the mutant nucleotide sequences of a polypeptide that is overexpressed in lung cancer patients may be a frameshift mutation, a missense mutation, a deletion, an insertion, a nonsense mutation, an inversion, a translocation, a duplication, or a CNV that results in the altered gene expression and / or activity.

[0107] In some embodiments, detection can occur through any of a variety of mobility dependent analytical techniques based on the differential rates of migration between different nucleic acid sequences. Exemplary mobility-dependent analysis techniques include electrophoresis, chromatography, mass spectroscopy, sedimentation, gradient centrifugation, field-flow fractionation, multi-stage extraction techniques, and the like. In some embodiments, mobility probes can be hybridized to amplification products, and the identity of the target nucleicacid sequence determined via a mobility dependent analysis technique of the eluted mobility probes, as described in Published PCT Applications WO04 / 46344 and WOOl / 92579. In some embodiments, detection can be achieved by various microarrays and related software such as the Applied Biosystems Array System with the Applied Biosystems 1700 Chemiluminescent Microarray Analyzer and other commercially available array systems available from Affymetrix, Agilent, Illumina, and Amersham Biosciences, among others (see also Gerry et al., J. Mol. Biol. 292:251-62, 1999; De Bellis et al., Minerva Biotec 14:247-52, 2002; and Stears et al., Nat. Med. 9: 14045, including supplements, 2003).

[0108] It is also understood that detection can comprise reporter groups that are incorporated into the reaction products, either as part of labeled primers or due to the incorporation of labeled dNTPs during an amplification, or attached to reaction products, for example but not limited to, via hybridization tag complements comprising reporter groups or via linker arms that are integral or attached to reaction products. In some embodiments, unlabeled reaction products may be detected using mass spectrometry.NGS Platforms

[0109] In some embodiments, high throughput, massively parallel sequencing employs sequencing-by-synthesis with reversible dye terminators. In other embodiments, sequencing is performed via sequencing-by-ligation. In yet other embodiments, sequencing is single molecule sequencing. Examples of Next Generation Sequencing techniques include, but are not limited to pyrosequencing, Reversible dye-terminator sequencing, SOLiD sequencing, Ion semiconductor sequencing, Helioscope single molecule sequencing etc.

[0110] The Ion Torrent™ (Life Technologies, Carlsbad, CA) amplicon sequencing system employs a flow-based approach that detects pH changes caused by the release of hydrogen ions during incorporation of unmodified nucleotides in DNA replication. For use with this system, a sequencing library is initially produced by generating DNA fragments flanked by sequencing adapters. In some embodiments, these fragments can be clonally amplified on particles by emulsion PCR. The particles with the amplified template are then placed in a silicon semiconductor sequencing chip. During replication, the chip is flooded with one nucleotide afteranother, and if a nucleotide complements the DNA molecule in a particular microwell of the chip, then it will be incorporated. A proton is naturally released when a nucleotide is incorporated by the polymerase in the DNA molecule, resulting in a detectable local change of pH. The pH of the solution then changes in that well and is detected by the ion sensor. If homopolymer repeats are present in the template sequence, multiple nucleotides will be incorporated in a single cycle. This leads to a corresponding number of released hydrogens and a proportionally higher electronic signal.

[0111] The 454TM GS FLX ™ sequencing system (Roche, Germany), employs a light-based detection methodology in a large-scale parallel pyrosequencing system. Pyrosequencing uses DNA polymerization, adding one nucleotide species at a time and detecting and quantifying the number of nucleotides added to a given location through the light emitted by the release of attached pyrophosphates. For use with the 454™ system, adapter-ligated DNA fragments are fixed to small DNA-capture beads in a water-in-oil emulsion and amplified by PCR (emulsion PCR). Each DNA-bound bead is placed into a well on a picotiter plate and sequencing reagents are delivered across the wells of the plate. The four DNA nucleotides are added sequentially in a fixed order across the picotiter plate device during a sequencing run. During the nucleotide flow, millions of copies of DNA bound to each of the beads are sequenced in parallel. When a nucleotide complementary to the template strand is added to a well, the nucleotide is incorporated onto the existing DNA strand, generating a light signal that is recorded by a CCD camera in the instrument.

[0112] Sequencing technology based on reversible dye-terminators: DNA molecules are first attached to primers on a slide and amplified so that local clonal colonies are formed. Four types of reversible terminator bases (RT-bases) are added, and non-incorporated nucleotides are washed away. Unlike pyrosequencing, the DNA can only be extended one nucleotide at a time. A camera takes images of the fluorescently labeled nucleotides, then the dye along with the terminal 3' blocker is chemically removed from the DNA, allowing the next cycle.

[0113] Helicos's single-molecule sequencing uses DNA fragments with added polyA tail adapters, which are attached to the flow cell surface. At each cycle, DNA polymerase and asingle species of fluorescently labeled nucleotide are added, resulting in template-dependent extension of the surface-immobilized primer-template duplexes. The reads are performed by the Helioscope sequencer. After acquisition of images tiling the full array, chemical cleavage and release of the fluorescent label permits the subsequent cycle of extension and imaging.

[0114] Sequencing by synthesis (SBS), like the "old style" dye-termination electrophoretic sequencing, relies on incorporation of nucleotides by a DNA polymerase to determine the base sequence. A DNA library with affixed adapters is denatured into single strands and grafted to a flow cell, followed by bridge amplification to form a high-density array of spots onto a glass chip. Reversible terminator methods use reversible versions of dye-terminators, adding one nucleotide at a time, detecting fluorescence at each position by repeated removal of the blocking group to allow polymerization of another nucleotide. The signal of nucleotide incorporation can vary with fluorescently labeled nucleotides, phosphate-driven light reactions and hydrogen ion sensing having all been used. Examples of SBS platforms include Illumina GA and HiSeq 2000. The MiSeq® personal sequencing system (Illumina, Inc.) also employs sequencing by synthesis with reversible terminator chemistry.

[0115] In contrast to the sequencing by synthesis method, the sequencing by ligation method uses a DNA ligase to determine the target sequence. This sequencing method relies on enzymatic ligation of oligonucleotides that are adjacent through local complementarity on a template DNA strand. This technology employs a partition of all possible oligonucleotides of a fixed length, labeled according to the sequenced position. Oligonucleotides are annealed and ligated and the preferential ligation by DNA ligase for matching sequences results in a dinucleotide encoded color space signal at that position (through the release of a fluorescently labeled probe that corresponds to a known nucleotide at a known position along the oligo). This method is primarily used by Life Technologies’ SOLiD™ sequencers. Before sequencing, the DNA is amplified by emulsion PCR. The resulting beads, each containing only copies of the same DNA molecule, are deposited on a solid planar substrate.

[0116] SMRT™ sequencing is based on the sequencing by synthesis approach. The DNA is synthesized in zero-mode wave-guides (ZMWs)-small well-like containers with the capturingtools located at the bottom of the well. The sequencing is performed with use of unmodified polymerase (attached to the ZMW bottom) and fluorescently labeled nucleotides flowing freely in the solution. The wells are constructed in a way that only the fluorescence occurring at the bottom of the well is detected. The fluorescent label is detached from the nucleotide at its incorporation into the DNA strand, leaving an unmodified DNA strand.Methods for Predicting the Risk of Recurrence and Metastases Using ctDNA as a Biomarker

[0117] In one aspect, the present disclosure provides a method for preventing risk of recurrence and / or metastases in a lung cancer patient in need thereof comprising (a) detecting ctDNA levels in a biological sample obtained from the lung cancer patient, wherein the ctDNA levels are detected at a variant allele fraction (VAF) detection limit of at least 0.1 %-0.5%; and (b) performing a lobectomy on the lung cancer patient, wherein the lung cancer patient comprises at least one lung tumor and exhibits an aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor.

[0118] In another aspect, the present disclosure provides a method for preventing risk of recurrence and / or metastases in a lung cancer patient in need thereof comprising performing a lobectomy on the lung cancer patient, wherein the lung cancer patient comprises detectable preoperative ctDNA levels, wherein the ctDNA levels are detected at a variant allele fraction (VAF) detection limit of at least 0.1%-0.5%; and wherein the lung cancer patient comprises at least one lung tumor and exhibits an aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor.

[0119] Additionally or alternatively, in some embodiments of the methods disclosed herein, the ctDNA levels are detected at a VAF detection limit of from about 0.1% to about 0.5%, from about 0.5% to about 2%, from about 2% to about 10% or from about 10% to about 99%. In certain embodiments, the ctDNA levels are detected at a VAF detection limit of about 0.1%, about 0.2%, about 0.3%, about 0.4%, about 0.5%, about 0.6%, about 0.7%, about 0.8%, about 0.9%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%,about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99%.

[0120] In some embodiments, the biological sample has a cfDNA concentration ranging from about 3 pg / pL to 5.5 ng / pL. In some embodiments, the biological sample has a cfDNA concentration of about 3 pg / pL, about 4 pg / pL, about 5 pg / pL, about 6 pg / pL, about 7 pg / pL, about 8 pg / pL, about 9 pg / pL, about 10 pg / pL, about 15 pg / pL, about 20 pg / pL, about 25 pg / pL, about 30 pg / pL, about 35 pg / pL, about 40 pg / pL, about 45 pg / pL, about 50 pg / pL, about 55 pg / pL, about 60 pg / pL, about 65 pg / pL, about 70 pg / pL, about 75 pg / pL, about 80 pg / pL, about 85 pg / pL, about 90 pg / pL, about 100 pg / pL, about 125 pg / pL, about 150 pg / pL, about 175 pg / pL, about 200 pg / pL, about 225 pg / pL, about 250 pg / pL, about 275 pg / pL, about 300 pg / pL, about 325 pg / pL, about 350 pg / pL, about 375 pg / pL, about 400 pg / pL, about 425 pg / pL, about 450 pg / pL, about 475 pg / pL, about 500 pg / pL, about 525 pg / pL, about 550 pg / pL, about 575 pg / pL, about 600 pg / pL, about 625 pg / pL, about 650 pg / pL, about 675 pg / pL, about 700 pg / pL, about 725 pg / pL, about 750 pg / pL, about 775 pg / pL, about 800 pg / pL, about 825 pg / pL, about 850 pg / pL, about 875 pg / pL, about 900 pg / pL, about 925 pg / pL, about 950 pg / pL, about 975 pg / pL, about 1 ng / pL, about 1.25 ng / pL, about 1.5 ng / pL, about 1.75 ng / pL, about 2 ng / pL, about 2.25 ng / pL, about 2.5 ng / pL, about 2.75 ng / pL, about 3 ng / pL, about 3.25 ng / pL, about 3.5 ng / pL, about 3.75 ng / pL, about 4 ng / pL, about 4.25 ng / pL, about 4.5 ng / pL, about 4.75 ng / pL, about 5 ng / pL, about 5.25 ng / pL, or about 5.5 ng / pL.

[0121] In yet another aspect, the present disclosure provides a method for preventing risk of recurrence and / or metastases in a lung cancer patient in need thereof comprising performing a limited resection or segmentectomy on the lung cancer patient, wherein the lung cancer patient comprises at least one lung tumor and exhibits an aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor; and wherein the lung cancer patient does not comprise any detectable pre-operative ctDNA levels.

[0122] In any of the preceding embodiments of the methods disclosed herein, the lung cancer is lung adenocarcinoma. The lung adenocarcinoma may comprise a histologic subtype selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. Additionally or alternatively, in some embodiments, the cancer is Stage 1, Stage 2, Stage 3, or Stage 4 cancer.

[0123] Additionally or alternatively, in certain embodiments, the lung cancer comprises a truncal loss or loss of heterozygosity (LOH) of 3p, a truncal loss or loss of heterozygosity (LOH) of 3q, a truncal gain of Iq, and / or a truncal gain of 8q. In other embodiments of the methods described herein, the lung cancer comprises a TP53 mutation, a KRAS mutation, a SMARCA4 mutation, a CCNE1 amplification, and / or a truncal loss or loss of heterozygosity (LOH) of chromosome 21q.

[0124] In any and all embodiments of the methods disclosed herein, the aerogenous spread of tumor cells in the lung parenchyma are colocalized with M2 macrophages and T regulatory cells. Additionally or alternatively, in some embodiments, the aerogenous spread of tumor cells in the peritumoral lung parenchyma beyond the edge of the at least one lung tumor is identified via histopathologic analysis of lung tissue sections obtained from the lung cancer patient. The lung tissue sections may be frozen sections or permanent sections. In any of the above embodiments of the methods disclosed herein, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0125] In any and all embodiments of the methods disclosed herein, the biological sample is whole blood, serum or plasma.Systems, Devices, and Methods for Predicting the Risk of Recurrence and Metastases in Lung Cancer Patients

[0126] Aspects of the operating environment as well as associated system components (e.g., hardware elements) in connection with various embodiments of the methods and systems described herein will now be discussed. Referring to FIG. 15A, an embodiment of a network environment is depicted. In brief overview, the network environment includes one or more clients 102a-102n (also generally referred to as local machine(s) 102, client(s) 102, client node(s) 102, client machine(s) 102, client computer(s) 102, client device(s) 102, endpoint(s) 102, or endpoint node(s) 102) in communication with one or more servers 106a-106n (also generally referred to as server(s) 106, node 106, or remote machine(s) 106) via one or more networks 104. In some embodiments, a client 102 has the capacity to function as both a client node seeking access to resources provided by a server and as a server providing access to hosted resources for other clients 102a-102n.

[0127] Although FIG. 15A shows a network 104 between the clients 102 and the servers 106, the clients 102 and the servers 106 may be on the same network 104. In some embodiments, there are multiple networks 104 between the clients 102 and the servers 106. In one of these embodiments, a network 104’ (not shown) may be a private network and a network 104 may be a public network. In another of these embodiments, a network 104 may be a private network and a network 104’ a public network. In still another of these embodiments, networks 104 and 104’ may both be private networks.

[0128] The network 104 may be connected via wired or wireless links. Wired links may include Digital Subscriber Line (DSL), coaxial cable lines, or optical fiber lines. The wireless links may include BLUETOOTH, Wi-Fi, Worldwide Interoperability for Microwave Access (WiMAX), an infrared channel or satellite band. The wireless links may also include any cellular network standards used to communicate among mobile devices, including standards that qualify as 1G, 2G, 3G, 4G, or 5G. The network standards may qualify as one or more generation of mobile telecommunication standards by fulfilling a specification or standards such as the specifications maintained by International Telecommunication Union. The 3G standards, for example, may correspond to the International Mobile Telecommunications-2000 (IMT-2000)specification, and the 4G standards may correspond to the International Mobile Telecommunications Advanced (IMT- Advanced) specification. Examples of cellular network standards include AMPS, GSM, GPRS, UMTS, LTE, LTE Advanced, Mobile WiMAX, and WiMAX-Advanced. Cellular network standards may use various channel access methods e.g. FDMA, TDMA, CDMA, or SDMA. In some embodiments, different types of data may be transmitted via different links and standards. In other embodiments, the same types of data may be transmitted via different links and standards.

[0129] The network 104 may be any type and / or form of network. The geographical scope of the network 104 may vary widely and the network 104 can be a body area network (BAN), a personal area network (PAN), a local-area network (LAN), e.g. Intranet, a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The topology of the network 104 may be of any form and may include, e.g., any of the following: point-to-point, bus, star, ring, mesh, or tree. The network 104 may be an overlay network which is virtual and sits on top of one or more layers of other networks 104’. The network 104 may be of any such network topology as known to those ordinarily skilled in the art capable of supporting the operations described herein. The network 104 may utilize different techniques and layers or stacks of protocols, including, e.g., the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, or the SDH (Synchronous Digital Hierarchy) protocol. The TCP / IP internet protocol suite may include application layer, transport layer, internet layer (including, e.g., IPv6), or the link layer. The network 104 may be a type of a broadcast network, a telecommunications network, a data communication network, or a computer network.

[0130] In some embodiments, the system may include multiple, logically-grouped servers 106. In one of these embodiments, the logical group of servers may be referred to as a server farm 38 or a machine farm 38. In another of these embodiments, the servers 106 may be geographically dispersed. In other embodiments, a machine farm 38 may be administered as a single entity. In still other embodiments, the machine farm 38 includes a plurality of machine farms 38. The servers 106 within each machine farm 38 can be heterogeneous - one or more ofthe servers 106 or machines 106 can operate according to one type of operating system platform (e.g., WINDOWS NT, manufactured by Microsoft Corp, of Redmond, Washington), while one or more of the other servers 106 can operate on according to another type of operating system platform (e.g., Unix, Linux, or Mac OS X).

[0131] In one embodiment, servers 106 in the machine farm 38 may be stored in high- density rack systems, along with associated storage systems, and located in an enterprise data center. In this embodiment, consolidating the servers 106 in this way may improve system manageability, data security, the physical security of the system, and system performance by locating servers 106 and high performance storage systems on localized high performance networks. Centralizing the servers 106 and storage systems and coupling them with advanced system management tools allows more efficient use of server resources.

[0132] The servers 106 of each machine farm 38 do not need to be physically proximate to another server 106 in the same machine farm 38. Thus, the group of servers 106 logically grouped as a machine farm 38 may be interconnected using a wide-area network (WAN) connection or a metropolitan-area network (MAN) connection. For example, a machine farm 38 may include servers 106 physically located in different continents or different regions of a continent, country, state, city, campus, or room. Data transmission speeds between servers 106 in the machine farm 38 can be increased if the servers 106 are connected using a local-area network (LAN) connection or some form of direct connection. Additionally, a heterogeneous machine farm 38 may include one or more servers 106 operating according to a type of operating system, while one or more other servers 106 execute one or more types of hypervisors rather than operating systems. In these embodiments, hypervisors may be used to emulate virtual hardware, partition physical hardware, virtualize physical hardware, and execute virtual machines that provide access to computing environments, allowing multiple operating systems to run concurrently on a host computer. Native hypervisors may run directly on the host computer. Hypervisors may include VMware ESX / ESXi, manufactured by VMWare, Inc., of Palo Alto, California; the Xen hypervisor, an open source product whose development is overseen by Citrix Systems, Inc.; the HYPER-V hypervisors provided by Microsoft or others. Hosted hypervisorsmay run within an operating system on a second software level. Examples of hosted hypervisors may include VMware Workstation and VIRTU ALBOX.

[0133] Management of the machine farm 38 may be de-centralized. For example, one or more servers 106 may comprise components, subsystems and modules to support one or more management services for the machine farm 38. In one of these embodiments, one or more servers 106 provide functionality for management of dynamic data, including techniques for handling failover, data replication, and increasing the robustness of the machine farm 38. Each server 106 may communicate with a persistent store and, in some embodiments, with a dynamic store.

[0134] Server 106 may be a file server, application server, web server, proxy server, appliance, network appliance, gateway, gateway server, virtualization server, deployment server, SSL VPN server, or firewall. In one embodiment, the server 106 may be referred to as a remote machine or a node. In another embodiment, a plurality of nodes 290 may be in the path between any two communicating servers.

[0135] Referring to FIG. 15B, a cloud computing environment is depicted. A cloud computing environment may provide client 102 with one or more resources provided by a network environment. The cloud computing environment may include one or more clients 102a- 102n, in communication with the cloud 108 over one or more networks 104. Clients 102 may include, e.g., thick clients, thin clients, and zero clients. A thick client may provide at least some functionality even when disconnected from the cloud 108 or servers 106. A thin client or a zero client may depend on the connection to the cloud 108 or server 106 to provide functionality. A zero client may depend on the cloud 108 or other networks 104 or servers 106 to retrieve operating system data for the client device. The cloud 108 may include back end platforms, e.g., servers 106, storage, server farms or data centers.

[0136] The cloud 108 may be public, private, or hybrid. Public clouds may include public servers 106 that are maintained by third parties to the clients 102 or the owners of the clients. The servers 106 may be located off-site in remote geographical locations as disclosed above or otherwise. Public clouds may be connected to the servers 106 over a public network. Privateclouds may include private servers 106 that are physically maintained by clients 102 or owners of clients. Private clouds may be connected to the servers 106 over a private network 104. Hybrid clouds 108 may include both the private and public networks 104 and servers 106.

[0137] The cloud 108 may also include a cloud based delivery, e.g. Software as a Service (SaaS) 110, Platform as a Service (PaaS) 112, and Infrastructure as a Service (laaS) 114. laaS may refer to a user renting the use of infrastructure resources that are needed during a specified time period. laaS providers may offer storage, networking, servers or virtualization resources from large pools, allowing the users to quickly scale up by accessing more resources as needed. Examples of laaS can include infrastructure and services (e.g., EG-32) provided by OVH HOSTING of Montreal, Quebec, Canada, AMAZON WEB SERVICES provided by Amazon.com, Inc., of Seattle, Washington, RACKSPACE CLOUD provided by Rackspace US, Inc., of San Antonio, Texas, Google Compute Engine provided by Google Inc. of Mountain View, California, or RIGHTSCALE provided by RightScale, Inc., of Santa Barbara, California. PaaS providers may offer functionality provided by laaS, including, e.g., storage, networking, servers or virtualization, as well as additional resources such as, e.g., the operating system, middleware, or runtime resources. Examples of PaaS include WINDOWS AZURE provided by Microsoft Corporation of Redmond, Washington, Google App Engine provided by Google Inc., and HEROKU provided by Heroku, Inc. of San Francisco, California. SaaS providers may offer the resources that PaaS provides, including storage, networking, servers, virtualization, operating system, middleware, or runtime resources. In some embodiments, SaaS providers may offer additional resources including, e.g., data and application resources. Examples of SaaS include GOOGLE APPS provided by Google Inc., SALESFORCE provided by Salesforce.com Inc. of San Francisco, California, or OFFICE 365 provided by Microsoft Corporation. Examples of SaaS may also include data storage providers, e.g. DROPBOX provided by Dropbox, Inc. of San Francisco, California, Microsoft SKYDRIVE provided by Microsoft Corporation, Google Drive provided by Google Inc., or Apple ICLOUD provided by Apple Inc. of Cupertino, California.

[0138] Clients 102 may access laaS resources with one or more laaS standards, including, e.g., Amazon Elastic Compute Cloud (EC2), Open Cloud Computing Interface (OCCI), CloudInfrastructure Management Interface (CIMI), or OpenStack standards. Some laaS standards may allow clients access to resources over HTTP, and may use Representational State Transfer (REST) protocol or Simple Object Access Protocol (SOAP). Clients 102 may access PaaS resources with different PaaS interfaces. Some PaaS interfaces use HTTP packages, standard Java APIs, JavaMail API, Java Data Objects (JDO), Java Persistence API (JPA), Python APIs, web integration APIs for different programming languages including, e.g., Rack for Ruby, WSGI for Python, or PSGI for Perl, or other APIs that may be built on REST, HTTP, XML, or other protocols. Clients 102 may access SaaS resources through the use of web-based user interfaces, provided by a web browser (e.g. GOOGLE CHROME, Microsoft INTERNET EXPLORER, or Mozilla Firefox provided by Mozilla Foundation of Mountain View, California). Clients 102 may also access SaaS resources through smartphone or tablet applications, including, e.g., Salesforce Sales Cloud, or Google Drive app. Clients 102 may also access SaaS resources through the client operating system, including, e.g., Windows file system for DROPBOX.

[0139] In some embodiments, access to laaS, PaaS, or SaaS resources may be authenticated. For example, a server or authentication server may authenticate a user via security certificates, HTTPS, or API keys. API keys may include various encryption standards such as, e.g., Advanced Encryption Standard (AES). Data resources may be sent over Transport Layer Security (TLS) or Secure Sockets Layer (SSL).

[0140] The client 102 and server 106 may be deployed as and / or executed on any type and form of computing device, e.g. a computer, network device or appliance capable of communicating on any type and form of network and performing the operations described herein. FIGs. 15C and 15D depict block diagrams of a computing device 100 useful for practicing an embodiment of the client 102 or a server 106. As shown in FIGs. 15C and 15D, each computing device 100 includes a central processing unit 121, and a main memory unit 122. As shown in FIG. 15C, a computing device 100 may include a storage device 128, an installation device 116, a network interface 118, an I / O controller 123, display devices 124a- 124n, a keyboard 126 and a pointing device 127, e.g. a mouse. The storage device 128 may include, without limitation, an operating system, software, and a software of a genomic dataprocessing system 120. As shown in FIG. 15D, each computing device 100 may also include additional optional elements, e.g. a memory port 103, a bridge 170, one or more input / output devices 130a- 13 On (generally referred to using reference numeral 130), and a cache memory 140 in communication with the central processing unit 121.

[0141] The central processing unit 121 is any logic circuitry that responds to and processes instructions fetched from the main memory unit 122. In many embodiments, the central processing unit 121 is provided by a microprocessor unit, e.g. : those manufactured by Intel Corporation of Mountain View, California; those manufactured by Motorola Corporation of Schaumburg, Illinois; the ARM processor and TEGRA system on a chip (SoC) manufactured by Nvidia of Santa Clara, California; the POWER7 processor, those manufactured by International Business Machines of White Plains, New York; or those manufactured by Advanced Micro Devices of Sunnyvale, California. The computing device 100 may be based on any of these processors, or any other processor capable of operating as described herein. The central processing unit 121 may utilize instruction level parallelism, thread level parallelism, different levels of cache, and multi-core processors. A multi-core processor may include two or more processing units on a single computing component. Examples of multi-core processors include the AMD PHENOM IIX2, INTEL CORE i5 and INTEL CORE i7.

[0142] Main memory unit or memory device 122 may include one or more memory chips capable of storing data and allowing any storage location to be directly accessed by the microprocessor 121. Main memory unit or device 122 may be volatile and faster than storage 128 memory. Main memory units or devices 122 may be Dynamic random access memory (DRAM) or any variants, including static random access memory (SRAM), Burst SRAM or SynchBurst SRAM (BSRAM), Fast Page Mode DRAM (FPM DRAM), Enhanced DRAM (EDRAM), Extended Data Output RAM (EDO RAM), Extended Data Output DRAM (EDO DRAM), Burst Extended Data Output DRAM (BEDO DRAM), Single Data Rate Synchronous DRAM (SDR SDRAM), Double Data Rate SDRAM (DDR SDRAM), Direct Rambus DRAM (DRDRAM), or Extreme Data Rate DRAM (XDR DRAM). In some embodiments, the main memory 122 or the storage 128 may be non-volatile; e.g., non-volatile read access memory(NVRAM), flash memory non-volatile static RAM (nvSRAM), Ferroelectric RAM (FeRAM), Magnetoresistive RAM (MRAM), Phase-change memory (PRAM), conductive-bridging RAM (CBRAM), Silicon-Oxide-Nitride-Oxide-Silicon (SONOS), Resistive RAM (RRAM), Racetrack, Nano-RAM (NRAM), or Millipede memory. The main memory 122 may be based on any of the above described memory chips, or any other available memory chips capable of operating as described herein. In the embodiment shown in FIG. 15C, the processor 121 communicates with main memory 122 via a system bus 150 (described in more detail below). FIG. 15D depicts an embodiment of a computing device 100 in which the processor communicates directly with main memory 122 via a memory port 103. For example, in FIG. 15D the main memory 122 may be DRDRAM.

[0143] FIG. 15D depicts an embodiment in which the main processor 121 communicates directly with cache memory 140 via a secondary bus, sometimes referred to as a backside bus. In other embodiments, the main processor 121 communicates with cache memory 140 using the system bus 150. Cache memory 140 typically has a faster response time than main memory 122 and is typically provided by SRAM, BSRAM, or EDRAM. In the embodiment shown in FIG. 15D, the processor 121 communicates with various I / O devices 130 via a local system bus 150. Various buses may be used to connect the central processing unit 121 to any of the I / O devices 130, including a PCI bus, a PCI-X bus, or a PCI-Express bus, or a NuBus. For embodiments in which the I / O device is a video display 124, the processor 121 may use an Advanced Graphics Port (AGP) to communicate with the display 124 or the I / O controller 123 for the display 124. FIG. 15D depicts an embodiment of a computer 100 in which the main processor 121 communicates directly with I / O device 130b or other processors 121’ via HYPERTRANSPORT, RAPIDIO, or INFINIBAND communications technology. FIG. 15D also depicts an embodiment in which local busses and direct communication are mixed: the processor 121 communicates with I / O device 130a using a local interconnect bus while communicating with I / O device 130b directly.

[0144] A wide variety of I / O devices 130a-130n may be present in the computing device 100. Input devices may include keyboards, mice, trackpads, trackballs, touchpads, touch mice,multi-touch touchpads and touch mice, microphones, multi-array microphones, drawing tablets, cameras, single-lens reflex camera (SLR), digital SLR (DSLR), CMOS sensors, accelerometers, infrared optical sensors, pressure sensors, magnetometer sensors, angular rate sensors, depth sensors, proximity sensors, ambient light sensors, gyroscopic sensors, or other sensors. Output devices may include video displays, graphical displays, speakers, headphones, inkjet printers, laser printers, and 3D printers.

[0145] Devices 130a-130n may include a combination of multiple input or output devices, including, e.g., Microsoft KINECT, Nintendo Wiimote for the WII, Nintendo WII U GAMEPAD, or Apple IPHONE. Some devices 130a- 13 On allow gesture recognition inputs through combining some of the inputs and outputs. Some devices 130a- 13 On provides for facial recognition which may be utilized as an input for different purposes including authentication and other commands. Some devices 130a-130n provides for voice recognition and inputs, including, e.g., Microsoft KINECT, SIR! for IPHONE by Apple, Google Now or Google Voice Search.

[0146] Additional devices 130a-130n have both input and output capabilities, including, e.g., haptic feedback devices, touchscreen displays, or multi-touch displays. Touchscreen, multitouch displays, touchpads, touch mice, or other touch sensing devices may use different technologies to sense touch, including, e.g., capacitive, surface capacitive, projected capacitive touch (PCT), in-cell capacitive, resistive, infrared, waveguide, dispersive signal touch (DST), incell optical, surface acoustic wave (SAW), bending wave touch (BWT), or force-based sensing technologies. Some multi-touch devices may allow two or more contact points with the surface, allowing advanced functionality including, e.g., pinch, spread, rotate, scroll, or other gestures. Some touchscreen devices, including, e.g., Microsoft PIXELSENSE or Multi-Touch Collaboration Wall, may have larger surfaces, such as on a table-top or on a wall, and may also interact with other electronic devices. Some I / O devices 130a-130n, display devices 124a-124n or group of devices may be augment reality devices. The I / O devices may be controlled by an I / O controller 123 as shown in FIG. 15C. The I / O controller may control one or more I / O devices, such as, e.g., a keyboard 126 and a pointing device 127, e.g., a mouse or optical pen. Furthermore, an I / O device may also provide storage and / or an installation medium 116 for thecomputing device 100. In still other embodiments, the computing device 100 may provide USB connections (not shown) to receive handheld USB storage devices. In further embodiments, an I / O device 130 may be a bridge between the system bus 150 and an external communication bus, e.g. a USB bus, a SCSI bus, a FireWire bus, an Ethernet bus, a Gigabit Ethernet bus, a Fibre Channel bus, or a Thunderbolt bus.

[0147] In some embodiments, display devices 124a-124n may be connected to I / O controller 123. Display devices may include, e.g., liquid crystal displays (LCD), thin fdm transistor LCD (TFT-LCD), blue phase LCD, electronic papers (e-ink) displays, flexile displays, light emitting diode displays (LED), digital light processing (DLP) displays, liquid crystal on silicon (LCOS) displays, organic light-emitting diode (OLED) displays, active-matrix organic light-emitting diode (AMOLED) displays, liquid crystal laser displays, time-multiplexed optical shutter (TMOS) displays, or 3D displays. Examples of 3D displays may use, e.g. stereoscopy, polarization fdters, active shutters, or autostereoscopy. Display devices 124a-124n may also be a head-mounted display (HMD). In some embodiments, display devices 124a- 124n or the corresponding I / O controllers 123 may be controlled through or have hardware support for OPENGL or DIRECTX API or other graphics libraries.

[0148] In some embodiments, the computing device 100 may include or connect to multiple display devices 124a-124n, which each may be of the same or different type and / or form. As such, any of the I / O devices 130a-130n and / or the I / O controller 123 may include any type and / or form of suitable hardware, software, or combination of hardware and software to support, enable or provide for the connection and use of multiple display devices 124a-124n by the computing device 100. For example, the computing device 100 may include any type and / or form of video adapter, video card, driver, and / or library to interface, communicate, connect or otherwise use the display devices 124a-124n. In one embodiment, a video adapter may include multiple connectors to interface to multiple display devices 124a-124n. In other embodiments, the computing device 100 may include multiple video adapters, with each video adapter connected to one or more of the display devices 124a-124n. In some embodiments, any portion of the operating system of the computing device 100 may be configured for using multipledisplays 124a-124n. In other embodiments, one or more of the display devices 124a-124n may be provided by one or more other computing devices 100a or 100b connected to the computing device 100, via the network 104. In some embodiments software may be designed and constructed to use another computer’s display device as a second display device 124a for the computing device 100. For example, in one embodiment, an Apple iPad may connect to a computing device 100 and use the display of the device 100 as an additional display screen that may be used as an extended desktop. One ordinarily skilled in the art will recognize and appreciate the various ways and embodiments that a computing device 100 may be configured to have multiple display devices 124a-124n.

[0149] Referring again to FIG. 15C, the computing device 100 may comprise a storage device 128 (e.g. one or more hard disk drives or redundant arrays of independent disks) for storing an operating system or other related software, and for storing application software programs such as any program related to the software for the genomic data processing system 120. Examples of storage device 128 include, e.g., hard disk drive (HDD); optical drive including CD drive, DVD drive, or BLU-RAY drive; solid-state drive (SSD); USB flash drive; or any other device suitable for storing data. Some storage devices may include multiple volatile and non-volatile memories, including, e.g., solid state hybrid drives that combine hard disks with solid state cache. Some storage device 128 may be non-volatile, mutable, or read-only. Some storage device 128 may be internal and connect to the computing device 100 via a bus 150.Some storage devices 128 may be external and connect to the computing device 100 via an I / O device 130 that provides an external bus. Some storage device 128 may connect to the computing device 100 via the network interface 118 over a network 104, including, e.g., the Remote Disk for MACBOOK AIR by Apple. Some client devices 100 may not require a nonvolatile storage device 128 and may be thin clients or zero clients 102. Some storage device 128 may also be used as an installation device 116, and may be suitable for installing software and programs. Additionally, the operating system and the software can be run from a bootable medium, for example, a bootable CD, e.g. KNOPPIX, a bootable CD for GNU / Linux that is available as a GNU / Linux distribution from knoppix.net.

[0150] Client device 100 may also install software or application from an application distribution platform. Examples of application distribution platforms include the App Store for iOS provided by Apple, Inc., the Mac App Store provided by Apple, Inc., GOOGLE PLAY for Android OS provided by Google Inc., Chrome Webstore for CHROME OS provided by Google Inc., and Amazon Appstore for Android OS and KINDLE FIRE provided by Amazon.com, Inc. An application distribution platform may facilitate installation of software on a client device 102. An application distribution platform may include a repository of applications on a server 106 or a cloud 108, which the clients 102a-102n may access over a network 104. An application distribution platform may include application developed and provided by various developers. A user of a client device 102 may select, purchase and / or download an application via the application distribution platform.

[0151] Furthermore, the computing device 100 may include a network interface 118 to interface to the network 104 through a variety of connections including, but not limited to, standard telephone lines LAN or WAN links (e.g., 802.11, Tl, T3, Gigabit Ethernet, Infiniband), broadband connections (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet-over- SONET, ADSL, VDSL, BPON, GPON, fiber optical including FiOS), wireless connections, or some combination of any or all of the above. Connections can be established using a variety of communication protocols (e.g., TCP / IP, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), IEEE 802.1 la / b / g / n / ac CDMA, GSM, WiMax and direct asynchronous connections). In one embodiment, the computing device 100 communicates with other computing devices 100’ via any type and / or form of gateway or tunneling protocol e.g. Secure Socket Layer (SSL) or Transport Layer Security (TLS), or the Citrix Gateway Protocol manufactured by Citrix Systems, Inc. of Ft. Lauderdale, Florida. The network interface 118 may comprise a built-in network adapter, network interface card, PCMCIA network card, EXPRESSCARD network card, card bus network adapter, wireless network adapter, USB network adapter, modem or any other device suitable for interfacing the computing device 100 to any type of network capable of communication and performing the operations described herein.

[0152] A computing device 100 of the sort depicted in FIGs. 15B and 15C may operate under the control of an operating system, which controls scheduling of tasks and access to system resources. The computing device 100 can be running any operating system such as any of the versions of the MICROSOFT WINDOWS operating systems, the different releases of the Unix and Linux operating systems, any version of the MAC OS for Macintosh computers, any embedded operating system, any real-time operating system, any open source operating system, any proprietary operating system, any operating systems for mobile computing devices, or any other operating system capable of running on the computing device and performing the operations described herein. Typical operating systems include, but are not limited to: WINDOWS 2000, WINDOWS Server 2022, WINDOWS CE, WINDOWS Phone, WINDOWS XP, WINDOWS VISTA, and WINDOWS 7, WINDOWS RT, WINDOWS 8, and WINDOWS 10, all of which are manufactured by Microsoft Corporation of Redmond, Washington; MAC OS and iOS, manufactured by Apple, Inc. of Cupertino, California; and Linux, a freely-available operating system, e.g. Linux Mint distribution (“distro”) or Ubuntu, distributed by Canonical Ltd. of London, United Kingdom; or Unix or other Unix-like derivative operating systems; and Android, designed by Google, of Mountain View, California, among others. Some operating systems, including, e.g., the CHROME OS by Google, may be used on zero clients or thin clients, including, e.g., CHROMEBOOKS.

[0153] The computer system 100 can be any workstation, telephone, desktop computer, laptop or notebook computer, netbook, ULTRABOOK, tablet, server, handheld computer, mobile telephone, smartphone or other portable telecommunications device, media playing device, a gaming system, mobile computing device, or any other type and / or form of computing, telecommunications or media device that is capable of communication. The computer system 100 has sufficient processor power and memory capacity to perform the operations described herein. The computer system 100 can be of any suitable size, such as a standard desktop computer or a Raspberry Pi 4 manufactured by Raspberry Pi Foundation, of Cambridge, United Kingdom. In some embodiments, the computing device 100 may have different processors, operating systems, and input devices consistent with the device. The Samsung GALAXYsmartphones, e.g., operate under the control of Android operating system developed by Google, Inc. GALAXY smartphones receive input via a touch interface.

[0154] In some embodiments, the computing device 100 is a gaming system. For example, the computer system 100 may comprise a PLAYSTATION 3, or PERSONAL PLAYSTATION PORTABLE (PSP), or a PLAYSTATION VITA device manufactured by the Sony Corporation of Tokyo, Japan, a NINTENDO DS, NINTENDO 3DS, NINTENDO WII, or a NINTENDO WII U device manufactured by Nintendo Co., Ltd., of Kyoto, Japan, an XBOX 360 device manufactured by the Microsoft Corporation of Redmond, Washington.

[0155] In some embodiments, the computing device 100 is a digital audio player such as the Apple IPOD, IPOD Touch, and IPOD NANO lines of devices, manufactured by Apple Computer of Cupertino, California. Some digital audio players may have other functionality, including, e.g., a gaming system or any functionality made available by an application from a digital application distribution platform. For example, the IPOD Touch may access the Apple App Store. In some embodiments, the computing device 100 is a portable media player or digital audio player supporting fde formats including, but not limited to, MP3, WAV, M4A / AAC, WMA Protected AAC, AIFF, Audible audiobook, Apple Lossless audio file formats and .mov, ,m4v, and .mp4 MPEG-4 (H.264 / MPEG-4 AVC) video file formats.

[0156] In some embodiments, the computing device 100 is a tablet e.g. the IPAD line of devices by Apple; GALAXY TAB family of devices by Samsung; or KINDLE FIRE, by Amazon.com, Inc. of Seattle, Washington. In other embodiments, the computing device 100 is an eBook reader, e.g. the KINDLE family of devices by Amazon.com, or NOOK family of devices by Barnes & Noble, Inc. of New York City, New York.

[0157] In some embodiments, the communications device 102 includes a combination of devices, e.g. a smartphone combined with a digital audio player or portable media player. For example, one of these embodiments is a smartphone, e.g. the IPHONE family of smartphones manufactured by Apple, Inc.; a Samsung GALAXY family of smartphones manufactured by Samsung, Inc.; or a Motorola DROID family of smartphones. In yet another embodiment, the communications device 102 is a laptop or desktop computer equipped with a web browser and amicrophone and speaker system, e.g. a telephony headset. In these embodiments, the communications devices 102 are web-enabled and can receive and initiate phone calls. In some embodiments, a laptop or desktop computer is also equipped with a webcam or other video capture device that enables video chat and video call.

[0158] In some embodiments, the status of one or more machines 102, 106 in the network 104 are monitored, generally as part of network management. In one of these embodiments, the status of a machine may include an identification of load information (e.g., the number of processes on the machine, CPU and memory utilization), of port information (e.g., the number of available communication ports and the port addresses), or of session status (e.g., the duration and type of processes, and whether a process is active or idle). In another of these embodiments, this information may be identified by a plurality of metrics, and the plurality of metrics can be applied at least in part towards decisions in load distribution, network traffic management, and network failure recovery as well as any aspects of operations of the present solution described herein. Aspects of the operating environments and components described above will become apparent in the context of the systems and methods disclosed herein.

[0159] Referring to FIG. 16, in various embodiments, a system 2400 may include a computing device 2410 (or multiple computing devices, co-located or remote to each other), a sample processing system 2480, and an electronic health record (EHR) system 2490. In various embodiments, computing device 2410 (or components thereof) may be integrated with the sample processing system 2480 (or components thereof) and / or EHR system 2490 (or components thereof). In various embodiments, the sample processing system 2480 may include, may be, or may employ, in situ hybridization, PCR, Next-generation sequencing, Northern blotting, microarray, dot or slot blots, FISH, Western blotting, ELISA, colorimetric dye binding assays, complete blood count (CBC) panels, FACs, electrophoresis, chromatography, and / or mass spectroscopy on such biological sample as blood, plasma, serum, and / or tissue and / or Whole-body MRI and PET-CT scans of a subject. For example, in certain embodiments, the sample processing system 2490 may be or may include a Next-generation sequencer. In various embodiments, the EHR system 2490 may include, may be, or may employ, various computingdevices that include health records of patients and study subjects (including devices of hospitals, clinics, healthcare practitioners, etc.), obtained from various sources, such as entries by healthcare practitioners, sample processing system 2480, university and hospital systems, government agency systems, etc.

[0160] In various embodiments, the computing device 2410 (or multiple computing devices) may be used to control, and receive signals acquired via, components of sample processing system 2480. The computing device 2410 may include one or more processors and one or more volatile and non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated. The computing device 2410 may include a control unit 2415 that in certain embodiments may be configured to exchange control signals with sample processing system 2480, allowing the computing device 2410 to be used to control, for example, processing of samples and / or scans and / or delivery of data generated and / or acquired through processing of samples and / or scans.

[0161] In various embodiments, computing device 2410 may include a data acquisition unit 2420 that may be configured to exchange control signals, or otherwise communicate, with sample processing system 2480 (or components thereof) and / or EHR system 2490, allowing the computing device 2410 to be used to control the capture of physiological data and / or signals via sensors of the sample processing system 2480, retrieve data or signals (e.g., from sample processing system 2480, EHR system 2490, and / or memory devices where data is stored), and direct transfer of data or signals (e.g., to sample processing system 2490 as feedback thereto, to EHR system 2490, to memory for storage, and / or to other systems or devices).

[0162] In various embodiment, a data analyzer 2425 may direct analysis of the data and signals, and output analysis results. Data analyzer 2425 may be used, for example, to transform raw data captured or obtained via sample processing system 2480 and / or EHR system 2490, and may employ pre-processing procedures involved in generating a training dataset. For example, in some implementations, data may be generated as a multi-dimensional array or vector with values representing, and to prevent the machine learning system from overemphasizing certain readings, values may be normalized to a predetermined range (e.g. 0-1, 0-100, or any other suchrange). The normalization may comprise linear rescaling, or may be a more complex function. In some implementations, dimension reduction may be performed to reduce large and sparse arrays or vectors. In some implementations, feature recognition may be performed to select a subset of features for further analysis, such as principal component analysis.

[0163] In various embodiments, a machine learning system 2430 may be used to implement various machine learning functionality discussed herein. Machine learning system 2430 may include a training engine 2435 configured to train predictive models using, for example, data obtained from or via data acquisition unit 2420 and / or processed data obtained from or via data analyzer 2425. The training engine 2435 may, for example, generate or obtain training datasets from or via data analyzer 2425 and may perform validation of datasets. The training engine 2435 may comprise a feature analyzer used to evaluate features by, for example, quantifying the impact of each feature on the developed model. Such a feature analyzer may, for example, uncover clinically important features that were globally predictive of the outcome, and may determine, for example, contributions of all features, or the top features (e.g., the top 2, top 5, top 10, top 15, top 20, top 25, top 30, etc.) on individual predictions. Features may be selected based on a threshold, such a percent contribution to predicting a medical condition, such as 0.5%, 1%, 2%, 5%, 10%, etc. A testing and application engine 2440 may be configured to test and apply models trained via training engine 2435 to, for example, study subject and / or patient data from data acquisition unit 2420 and / or data analyzer 2425.

[0164] In various embodiments, a transceiver 2445 allows the computing device 2410 to exchange readings, control commands, and / or other data with sample processing system 2480 (or components thereof) and / or EHR system 2490 (or components thereof). The transceiver 2445 may additionally or alternatively include a network interface permitting the computing device 2410 to communicate with other remote devices and systems via, for example, a telecommunications network such as the internet. One or more user interfaces 2450 allow the computing device 2410 to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, etc.) and provide outputs (e.g., via a touchscreen or other display screen, audio speakers, haptic devices, etc. ). A display screen may be employed, for example, to provide real time ornear real time waveforms or other readings or measurements obtained via sensors being used to capture physiological data from subjects and patients. The computing device 2410 may additionally include one or more databases 2455 (stored in, e.g., one or more computer-readable non-volatile memory devices) for storing, for example, data and analyses obtained from or via data acquisition unit 2420, data analyzer 2425, machine learning system 2430 (e.g., training engine 2435 and / or testing and application engine 2440), sample processing system 2480, and / or EHR system 2490. In some implementations, database 2455 (or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote and in communication with computing device 2410, sample processing system 2480 (or components thereof), and / or EHR system 2490.

[0165] In one aspect, the present disclosure provides a method of training a machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, comprising: (a) receiving data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; (b) generating a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, wherein applying the machine learning method comprises: (i) applying a machine learning technique to the training dataset; (ii) performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the classifier; and (iii) determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In some embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0166] Additionally or alternatively, in some embodiments, the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4. Examples of surgery type include, but are not limited to, pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection. The adjuvant therapy may comprise chemotherapy or immunotherapy. In other embodiments, the lung cancer patient has not received adjuvant therapy.

[0167] In any of the preceding embodiments of the methods disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of clustercell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0168] In any of the preceding embodiments of the methods disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0169] Additionally or alternatively, in some embodiments, the methods of the present technology further comprise applying the classifier to data on a lung cancer patient to generate apredictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In certain embodiments, the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0170] Additionally or alternatively, in certain embodiments, the methods of the present technology further comprise performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments, the methods described herein further comprise performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

[0171] In one aspect, the present disclosure provides a method of estimating risk of recurrence and / or metastases in a lung cancer patient using a machine learning classifier, the method comprising: (a) receiving patient data corresponding to a plurality of features for the lung cancer patient; (b) applying the machine learning classifier to the patient data to generate a predictor; and (c) determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the machine learning classifier is trained by: (A) receiving cohort data on a cohort of lung adenocarcinoma (LUAD) subjects, each LUAD subject in the cohort comprising at least one lung tumor; (B) generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (C) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein themachine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0172] In some embodiments, the methods of the present technology further comprise performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In some embodiments, the methods disclosed herein further comprise administering an effective amount of an adjuvant therapy to the lung cancer patient, optionally wherein the adjuvant therapy comprises chemotherapy or immunotherapy. In other embodiments, the methods of the present technology further comprise performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold. The predictor may comprise a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0173] Additionally or alternatively, in some embodiments, the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

[0174] In any of the preceding embodiments of the methods disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of clustercell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0175] In any of the above embodiments of the methods disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regressiontechnique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0176] Additionally or alternatively, in some embodiments of the methods described herein, one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA, and / or one or more of the plurality of features for each subject in the cohort are determined by assaying blood and / or sequencing tumor DNA.

[0177] In another aspect, the present disclosure provides a machine learning system for training a machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, the system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients; wherein applying the machine learning method comprises: (a) applying a machine learning technique to the training dataset; (b) performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and (c) determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves forthe training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0178] In any of the preceding embodiments of the machine learning systems disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0179] Additionally or alternatively, in some embodiments, the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4. Examples of surgery type include, but are not limited to, pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection. The adjuvant therapy may comprise chemotherapy or immunotherapy. In other embodiments, the lung cancer patient has not received adjuvant therapy.

[0180] In any of the above embodiments of the machine learning systems disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more ofcluster-cell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0181] Additionally or alternatively, in some embodiments of the machine learning systems of the present technology, the instructions further cause the processor to apply the classifier to data on a lung cancer patient to generate a predictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In certain embodiments, the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0182] Additionally or alternatively, in certain embodiments of the machine learning systems disclosed herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the machine learning systems disclosed herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

[0183] In yet another aspect, the present disclosure provides a computing system for estimating risk of recurrence and / or metastases in lung cancer patients, the computing system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: (a) receive patient data corresponding to a plurality of features for the lung cancer patient; (b) apply a machine learning classifier to the patient data to generate a predictor; and (c) determine whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the classifier is trained by: (A) receiving cohort data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; (B) generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (C)applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0184] In any of the above embodiments of the computing systems disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0185] Additionally or alternatively, in some embodiments of the computing system described herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the computing system described herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrenceand / or metastases based on the predictor and the operating-point threshold. The predictor may comprise a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0186] Additionally or alternatively, in some embodiments, the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

[0187] In any of the preceding embodiments of the computing systems disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0188] In any and all embodiments of the computing systems described herein, one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA.

[0189] In one aspect, the present disclosure provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a machine learning system, configure the machine learning system to train a machine learning classifier to estimate risk of recurrence and / or metastases in lung cancer patients, the instructions configured to cause the processor to: (a) receive data on a cohort of lung adenocarcinoma (LUAD) subjects, each LUAD subject in the cohort comprising at least one lung tumor; (b) generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (c) apply a machine learning method to the training dataset to develop the machine learning classifier for estimatingrisk of recurrence and / or metastases in lung cancer patients; wherein applying the machine learning method comprises: (A) applying a machine learning technique to the training dataset; (B) performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and (C) determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0190] In any of the above embodiments of the computer-readable storage medium disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0191] Additionally or alternatively, in some embodiments, the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4. Examples of surgery type include, but are not limited to, pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection. The adjuvant therapy may comprisechemotherapy or immunotherapy. In other embodiments, the lung cancer patient has not received adjuvant therapy.

[0192] In any of the above embodiments of the computer-readable storage medium disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non- circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0193] Additionally or alternatively, in some embodiments of the computer-readable storage medium of the present technology, the instructions further cause the processor to apply the classifier to data on a lung cancer patient to generate a predictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In certain embodiments, the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0194] Additionally or alternatively, in some embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

[0195] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a computing system, configure the computing system to estimate risk of recurrence and / or metastases in lung cancer patients, the instructions configured to cause the processor to: (a) receive patient data corresponding to a plurality of features for the lung cancer patient; (b) apply a machine learningclassifier to the patient data to generate a predictor; and (c) determine whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the classifier is trained by: (A) receiving cohort data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; (B) generating a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and (C) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset, wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients. In certain embodiments, the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

[0196] In any of the above embodiments of the computer-readable storage medium disclosed herein, the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique. In certain embodiments, the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model. In other embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, the machine learning technique models survival outcomes with competing risks. In some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.'ll

[0197] Additionally or alternatively, in some embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold. In other embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold. The predictor may comprise a cumulative incidence function (CIF) for lung recurrence and / or metastases.

[0198] Additionally or alternatively, in some embodiments, the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis. The histological subtype of lung adenocarcinoma may be selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern. In some embodiments, the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

[0199] In any of the foregoing embodiments of the computer-readable storage medium disclosed herein, the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects. The lung tissue sections may be frozen sections or permanent sections. Additionally or alternatively, in certain embodiments, STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non- circumferential STAS. In some embodiments, the status of STAS is determined by a thoracic pathologist.

[0200] In any of the preceding embodiments of the computer-readable storage medium disclosed herein, one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA.EXAMPLES

[0201] The present technology is further illustrated by the following Examples, which should not be construed as limiting in any way.Example 1: Experimental methods

[0202] The TRACERx 421 cohort, sample collection, and DNA / RNA sequencing

[0203] All patients were enrolled into the TRACERx study (https: / / clinicaltrials.gov / ct2 / show / NCT01888601, approved by an independent research ethics committee, 13 / LO / 1546), with multi-region sampling of tumors, DNA and RNA extraction, whole exome sequencing (WES) and RNA sequencing as previously described1>2. The first 421 patients, which constitute the first half of the originally scheduled full cohort of the clinical trial, are included in the TRACERx 421 cohort. Inclusion / exclusion criteria of the clinical study, clinical data acquisition, and bioinformatic processing of DNA and RNA sequencing data, including fusion gene analysis and pre-operative ctDNA analysis, were performed as described in our companion manuscripts3 6.

[0204] Histopathological assessment

[0205] (1) Central histopathological review

[0206] The diagnostic slides from all lung adenocarcinoma (LU AD) cases in the cohort were requested from the local pathology departments, scanned using a Hamamatsu Nanozoomer S210 slide scanner at 40x scanning magnification and retained within a central digital histology archive. In a small number of cases, full diagnostic slides were not available and therefore pathology review was conducted using a combination of a single representative diagnostic slide and regional TRACERx tissue samples. Full diagnostic slides were used for central pathology review to confirm tumor subtype and to generate adenocarcinoma growth pattern fractions following review by central study pathologists (FIG. 5K). Tumors were categorised into six growth patterns - the five patterns currently defined in the WHO classification7, with the addition of a cribriform pattern which has elsewhere been included as part of the acinar growth pattern subtype (FIGs. 5A-5H). As per standard clinical diagnostic practice and current guidance, we labelled each tumor using the predominant histological subtype according to the proportions of each growth pattern. Invasive mucinous adenocarcinoma (including mixed invasive mucinous and nonmucinous adenocarcinoma) was labelled as a separate entity, in line with the WHO classification, though pattern proportions of the six growth patterns were still ascribed. For28 / 242 patients whose full diagnostic slides were not available, pattern proportion was based upon local histopathology reporting, provided this matched broadly with the appearances of any available material at central review. Any discrepancy in tumor type between the clinical pathology report and central review was subject to additional expert review for a final definitive diagnosis. The presence of STAS was defined as previously described8,9.

[0207] (2) Nuclear grade and mitotic index

[0208] The nuclear grade was determined based upon nuclear size in the highest grade area of the tumor, as previously described10. The mitotic index was determined from diagnostic material from the area of the tumor with the highest mitotic activity. The count was performed over a 2.4mm2area of tumor on scanned slides, equal to the area of 10 high-power microscopic fields used in previous LU AD grading studies11.

[0209] (3) Regional growth pattern annotation

[0210] Where regional fresh tissue samples were sufficient, these samples were split between fresh frozen tissue for DNA and RNA extraction, and formalin-fixed paraffin-embedded (FFPE) tissue to allow histological assessment of the sequenced regions. This regional histology was used to generate regional growth pattern data in lung adenocarcinoma and regional stromal TIL scores for the entire cohort.

[0211] (4) Definition of tumor growth pattern homogeneity

[0212] A tumor was defined as morphologically homogeneous when the tumor met both of the following criteria: 1) predominant subtype was dominating 90% or more of the tumor area in diagnostic slides and 2) predominant subtype was dominating 90% or more of the annotated regional patterns.

[0213] (5) Growth pattern annotation in the metastatic sample

[0214] In metastatic disease, where possible, the metastatic tumor was sampled for sequencing analysis. If sufficient tissue was available for histological assessment, growth pattern was also characterised in lung adenocarcinoma samples. Biopsies that were too limited forgrowth pattern annotation or were too dissociated, such as some aspirated samples, were excluded from the analysis.

[0215] (6) PD-L1 and Ki-67 immunohistochemical staining

[0216] PD-L1 (22C3 clone) and Ki-67 (MIB-1 clone) immunohistochemistry were performed on a single representative diagnostic FFPE tissue block from the resection specimen using a Link48 Autostainer (Agilent) for PD-L1 and a Bond- III Autostainer (Leica Biosystems) for Ki-67, according to the manufacturer’s instructions. Fractions of positively stained tumor cells (cancer cells) were scored manually by a pathologist in line with clinical guidelines.

[0217] (7) Pathology TIL estimates

[0218] TILs were estimated from haematoxylin and eosin (H&E) stained slides using established international guidelines, developed by the International Immuno-Oncology Biomarker Working Group, as described in the previous reports2,12’13. In brief, the relative proportion of stromal area to tumor area was determined from the pathology slide of a given tumor region. TILs were reported for the stromal compartment (= percent stromal TILs). The denominator used to determine the percent stromal TILs was the area of stromal tissue. Therefore percent stromal TILs equalled the area occupied by mononuclear inflammatory cells over the total intratumoral stromal area rather than the fraction of total stromal nuclei that represent mononuclear inflammatory cell nuclei. This method has been demonstrated to be reproducible among trained pathologists14. The International Immuno-Oncology Biomarker Working Group has developed a freely available training tool to train pathologists for optimal TIL assessment on H&E slides (www.tilsincancer.org).

[0219] TRACERx analytical pipeline and orthogonal method for SCNA profiling and clonality inference

[0220] (1) SCNA profiling

[0221] In the TRACERx WES pipeline, copy number segmentation data were produced using ASCAT15and then a multi-sample somatic copy number alteration (SCNA) estimation approach3,16was applied in which single nucleotide polymorphisms (SNPs) are phased ontopaternal and maternal alleles using samples with allelic imbalance and detectable B-allele- frequency (BAF) separation. This approach allowed copy number aberrations present in one region to be tested for in other regions, and enabled us to more accurately profde copy number states in low purity tumor regions. To determine genome- wide copy number gain and loss events, copy number data for each sample was divided by the sample mean ploidy, and log2 transformed. Amplification, gain, and loss thresholds were defined as log2 (4 / 2), log2(2.5 / 2), and log2(1.5 / 2), respectively. LOH was defined as a floating point copy number of the minor allele of <0.5. As discussed previously1, our pipeline may be under-calling homozygous deletions, due to the nature of the relatively low resolution of the WES data (restricted to exonic regions) making it difficult to call very focal events. Therefore, we excluded homozygous deletions from the analysis.

[0222] (2) Clonality inference

[0223] In the TRACERx WES pipeline, mutations were classified as truncal or subclonal using a modified version of PyClone(v0.13.1)17. Several additional steps were then carried out during phylogenetic tree building to avoid overcalling subclonality and inappropriately increasing the reported degree of tumor heterogeneity1 3. For copy number amplification, if at least one region showed an amplified mean copy number, we called the amplification truncal if all other regions showed a copy number gain of ploidy +1 copy. The gene would be called subclonally amplified if at least one region showed no gain, or if a truncal event overlapped with SCNAs that disrupt the same genomic region but affect different parental alleles within separate tumor subclones, what we called a mirrored subclonal allelic imbalance in a previous publication1. Truncal / subclonal chromosomal arm gain and loss were defined on a per tumor basis, by requiring at least one region to show at least 98% gain or loss of the arm. Truncal arm gain or loss was then called if the same chromosomal arm showed at least 75% gain or loss across all remaining regions. Subclonal arm gain or loss was called if at least one region showed less than 75% gain or loss of the chromosomal arm or if a chromosomal arm was subject to mirrored subclonal allelic imbalance1.

[0224] (3) Orthogonal methods

[0225] As an orthogonal method for SCNA profiling, we used the default output of Sequenza18for each sample. As a default output, Sequenza returns integer copy numbers for major and minor alleles. LOH was defined as integer copy number = 0 for the minor allele and >0 for the major allele. As an alternative method to call clonality for mutations and SCNAs, we called any ubiquitous events truncal (e.g. amplification observed in all regions per tumor), and any non-ubiquitous events subclonal (e.g. one region with amplification but all other regions were gains).

[0226] Enumeration of chromosomal complexity, clonal architecture, and intra-tumor heterogeneity

[0227] (1) Weighted genome instability index (wGII score)

[0228] The wGII score was calculated as the proportion of the genome with aberrant copy number relative to the median ploidy, weighted on a per chromosome basis19.

[0229] (2) Fraction of genome subject to loss of heterozygosity (FLOH score)

[0230] The FLOH score was defined as the proportion of the genome subject to loss of heterozygosity.

[0231] (3) % subclonal SCNA (= SCNA-ITH)

[0232] The percentage of the genome subject to subclonal SCNA events was divided by the percentage of the genome involved in any SCNA event in each tumor1’3.

[0233] (4) % subclonal TMB (= Mut-ITH).

[0234] The number of mutations estimated to be subclonal was divided by the total number of mutations classified as either truncal or subclonal after phylogenetic tree building in each tumor1 3.

[0235] (5) Subclonal TMB (clonal in > 1 region) (‘illusion of clonality’-type subclonal mutation burden)

[0236] Mutational clusters used for phylogenetic tree building were determined to be subclonal or clonal within every region by testing whether cancer cell fraction (CCF, fraction ofcancer cells harbouring the cluster of mutations) was significantly lower than 1. Subcl onal TMB of regionally clonal mutations in > 1 region but not in all regions (i.e. subclonal at tumor level), previously described as ‘illusion of clonality’1,3mutations because they may appear as clonal when only one region is sampled per tumor1, was calculated for each tumor.

[0237] (6) Subclonal diversity within the tumor region

[0238] First, cancer cell clone proportions, namely what percentage of the cells in that region come from each clone, were calculated using the cancer cell fraction of each mutational cluster that was used for phylogenetic tree construction. Then, the Shannon diversity index of the clone proportions was calculated to give the subclonal diversity of each region. Minimum subclonal diversity within the tumor region was used to represent the subclonal diversity per tumor3.

[0239] (7) Recent subclonal expansion score

[0240] A recent subclonal expansion score per tumor, reflecting the size of the largest recent subclonal expansion within each tumor region, was calculated as follows3. First, for each tumor phylogenetic tree, the terminal nodes on the tree (i.e. leaf nodes) were identified. Then for each tumor region, the maximum CCF of any of these leaf nodes was identified. Lastly, as a tumor level metric, the subclonal expansion score was calculated by taking the maximum across the regional scores, therefore describing the maximal size of the most recent subclone expansion in each tumor.

[0241] Determinants of predominant subtype

[0242] (1) Truncal genomic alterations associated with predominantly high-grade or low / mid-grade tumors

[0243] To investigate the truncal genomic alterations associated with predominantly highgrade pattern tumors, we first compiled recurrent truncal events observed in more than 5% of the tumors and in at least 10 tumors in either predominantly high-or low / mid-grade tumors in the TRACERx 421 LU AD cohort. The compiled list was composed of 13 truncal driver alterations (7 mutations, 6 amplifications) and 31 truncal chromosomal arm SCNAs (8 gains and 23 loss / LOHs) (FIGs. 6A-6C). Logistic regression analysis was performed to construct a model todistinguish between tumors with predominantly high-grade and low / mid-grade. Specifically, we constructed an initial model with the presence / absence of these 44 truncal events as explanatory variables. Stepwise model simplification was performed using the stepAIC function (MASS (7.3-54) R package). The final model was composed of 11 truncal genomic alterations, of which 7 were determined to be significantly independent variables (FIG. ID). The results were consistent with the results when only truncal alterations observed in at least 10% of the cohort were included in the analysis, when an orthogonal tool (Sequenza18) for SCNA profiling was applied, or when wGII was added to the final model as a covariate to control for general genomic instability.

[0244] (2) Differential copy number analysis of driver genes and chromosomal arms between predominantly high-grade and low / mid-grade tumors

[0245] To capture SCNAs with a frequency in the cohort of lower than 5%, we also compared the ploidy-adjusted copy number of driver genes and chromosomal arms between predominantly high-grade and low / mid-grade tumors. To account for multi-region input from a single tumor, a linear mixed effect model was applied (response variable = ploidy adjusted copy number of each SCNA, fixed effect = predominant subtype (high-vs low / mid-grade), random effect = tumor ID), using the nlme (3.1-153) R package. P values were adjusted for multiple comparisons using the Benj ami ni -Hochberg (BH) method. The results were consistent with the results when Sequenza18was applied for SCNA profiling, or when wGII was added to the linear mixed effect model as a covariate.

[0246] (3) Mutual exclusivity and co-occurrence of truncal genomic alterations associated with predominantly high-grade or low / mid-grade tumors

[0247] To determine significantly mutually exclusive and co-occurring (me-co) relationships between recurrent truncal genomic alterations observed in more than 5% of the tumors in the TRACERx 421 LU AD cohort, the Rediscover (0.3.0) R package, which applies statistical analysis based on the Poisson-Binomial distribution to take into account the alteration rate of genes and samples, was implemented20. The truncal events observed in at least 10 tumors in either predominantly high- or low / mid-grade tumors were included in the analysis. A getMutexfunction was applied, with a binary matrix of the presence / absence of truncal driver gene mutations and a binary matrix of truncal SCNA alterations (driver gene amplifications, chromosome arm gains, and arm loss / LOH) provided as input data, for the low / mid-grade predominant and high-grade predominant tumors separately. A getMutexAB function was also run, with binary matrices of the presence / absence of truncal driver gene mutations and truncal SCNA alterations provided as input data. The outputs from getMutex and getMutexAB functions were combined, and the probabilities of mutual exclusivity and co-occurrence were adjusted for multiple comparisons using the BH method. To focus on the mutually exclusive or co-occurring truncal events specific to predominantly high-grade or low / mid-grade tumors, the events with unadjusted P value < 0.05 in both predominantly high-grade and low / mid-grade tumors were filtered out. Analyses were conducted using R 4.0.0 (R Foundation for Statistical Computing, Vienna, Austria). The same analyses were conducted including only truncal alterations observed in at least 10% of the cohort or using Sequenza18as an orthogonal tool for SCNA profiling.

[0248] GISTIC2.0 peak identification for tumor regions with solid pattern

[0249] GISTIC2.021takes as input a copy number profile across the genome from one sample per tumor. To investigate genomic regions of recurrent gains and amplifications associated with solid pattern, the copy number profiles from all solid-pattern regions from the same tumor were uniformly segmented by taking minimum consistent segmentations and the single sample copy number profile for each tumor was constructed by selecting the minimum ploidy-corrected total copy number per segment across the genome. By taking the minimum ploidy-corrected total copy number per segment, a significant peak (r / < 0.05) in chromosome 3q (chr3: 131091386-191871390, 3q21.3-3q29) was inferred as a truncal focal amp / gain peak associated with the presence of the solid-pattern regions. To investigate the presence of the gain involving the focal 3q segment in each tumor region, mean ploidy-adjusted copy number (CN) of the focal 3q segment was calculated and log2 transformed. When all tumor regions harboured mean ploidy-adjusted CN > log2(2.5 / 2), the tumor was classified as having a truncal gain of the focal 3q segment (including chromosome 3q arm gain), and the tumor was classified as having subclonal gain when not all but at least one region harboured a gain of the segment.

[0250] Identification of factors independently associated with tumor cell PD-L1 protein expression

[0251] To investigate whether the proportion of solid pattern per tumor is independently associated with tumor cell (cancer cell) PD-L1 protein expression (0% vs >1%), we performed logistic regression analysis (response variable = PD-L1 protein expression) to account for potential confounders including the amount of TILs (pathology TIL scores), truncal and subclonal neoantigen burden, and the presence of any loss of heterozygosity of human leukocyte antigen (HLA LOH) inferred using LOHHLA23. Variables with P<0.2 at univariate analysis were included for the multivariable analysis: pathology TIL scores, truncal and subclonal neoantigen burden (loglO), and solid pattern %.

[0252] Genomic distance

[0253] The genomic distance using mutations was calculated as previously described2. In brief, all detected mutations (SNVs and indels) present in any region of a tumor were turned into a binary matrix (1 : mutation present, 0: mutation absent), in which the rows were the mutations and the columns were tumor regions. The pairwise Euclidean distance between any two tumor regions within each tumor was calculated. The genomic distance using LOH was calculated similarly. In brief, the presence of cytoband level LOH was first assigned when the copy-number status of the largest genomic segment overlapping each cytoband was LOH. Then, all copynumber states per cytoband in any region from a tumor were turned into a binary matrix (1 : LOH present, 0: LOH absent), in which the rows were genomic segments (cytoband) and the columns were tumor regions, and the pairwise Euclidean distance between any two tumor regions was calculated. Mixed effects models are regression analyses for repeated data measures (e.g. where one patient provides multiple outcomes), and a linear mixed effect model was applied for the comparison of genomic distances between regional pairs with same vs different growth patterns to account for multiple regional pairs from a single tumor (response variable = genomic distance, fixed effect = growth pattern (different vs same), random effect = tumor ID), using the nlme (3.1-153) R package.

[0254] Sequential evolution from lower-grade pattern to higher-grade pattern

[0255] (1) Application of grade pattern scoring

[0256] To compare the regional growth patterns between specific groups of regions (for example, seeding vs non-seeding regions in the primary tumor, or seeding regions vs metastasis samples), growth pattern was transformed into integer scores as follows: lepidic, 0 (low-grade pattern); papillary and acinar, 1 (mid-grade pattern); and cribriform, micropapillary, and solid, 3 (high-grade pattern) and then the mean of regional pattern grade scores calculated for each regional group within each tumor.

[0257] (2) Ancestor-descendant-like relation inference

[0258] Within the primary tumors, regional pairs having ancestor-descendant-like relations were investigated under the assumption that although a tumor region has not directly evolved from another region, there might be a regional pair in which one region is harbouring a common ancestral-like clone and another region is harbouring a descendant-like clone of the common ancestral clone. Assignments of ancestor-like region and descendant-like region were made as follows:

[0259] First, primary tumor regions with purity <0.15 were removed, in case LOH calling in low purity regions was not robust;

[0260] Next, a LOH tree was built for all regional pairs by counting shared (=trunk) LOH and private (=branch) LOH per cytoband;

[0261] Based on the hypothesis that the majority of LU AD tumors have one or more truncal arm-level LOH events, and that regional pairs without any shared arm-level LOH may be due to technical errors (e.g. inappropriate ploidy / purity solutions), only LOH trees with at least one shared arm-level LOH were retained for further analysis;

[0262] For the remaining regional pairs and LOH trees,

[0263] An ancestor-like region, namely, a region that harbours a clone similar to a recent common ancestral clone, was defined as a region with a private LOH branch length less than X% of the trunk length, where X=2 is used in the main text and figure.

[0264] A descendant-like region was defined as a region with a private LOH branch length of more than Y% of the trunk length and more than one arm-level private LOH, where Y=10 is used in the main text and figure.

[0265] To test if the ancestor-descendant-like relations inferred by the LOH profile conflict with mutational profiles, dominant mutations with cancer cell fraction (CCF) >95% (namely, mutations shared among >95% cancer cells within each region) were compared between paired descendant-like regions and ancestor-like regions.

[0266] Ancestor-descendant-like regional pairs were also inferred using both LOH and mutational profiles together, applying the same method for dominant mutations (CCF >95%). Namely, a mutation tree (CCF >95%) was built for all regional pairs by counting shared mutations (=trunk) and private mutations (=branch), and the ancestor-like region was defined as a region with private mutational branch length less than 2% of the trunk length, and the descendant-like region was defined as a region with a private mutational branch length of more than 10% of the trunk length. In the combined method of LOH and mutational profiles, the regional pairs called as having ancestor-descendant-like relations by both methods were inferred as ancestor-descendant-like regional pairs. Additionally, as an orthogonal method for calling SCNA profiles, we applied Sequenza18, and otherwise, the same definition for inferring ancestor- descendant-like relations.

[0267] (3) Permutation test

[0268] To test the enrichment of higher grade patterns in the descendant-like regions and the enrichment of higher mutational burden (CCF>95%) in the descendant-like regions (namely, lower-to-higher (upward) transition), permutation tests were applied to obtain empirical P values using the Monte-Carlo procedure22. Firstly, the number of ancestor-descendant-like regional pairs with the upward transition was counted (= observed frequency). Next, for each permutation, regional growth patterns and regional mutational burdens were randomised within each tumor and the frequency of upward transition was compared against the observed frequency. Finally, the empirical P value was calculated by Equation 1 :

[0269] P=(r+l) / (n+l) [Equation 1]

[0270] where r is the number of permutations that produced a higher frequency of upward transition compared with the observed frequency and n is the number of permutations.

[0271] Histological factors associated with tumor cell spread through air spaces (STAS) and pre-operative circulating tumor DNA (ctDNA) positivity

[0272] To elucidate the biological difference between STAS positivity and pre-operative ctDNA positivity, a univariable logistic regression model was applied. For the response variable, either STAS positivity or preoperative ctDNA positivity was used. For the explanatory variable, each of the following histological variables was included in the model: the presence of each growth pattern, mitotic index, nuclear grade, Ki-67 fraction of tumor cells, type of tumor (IMA or not), presence of necrosis, lymphovascular invasion, visceral pleural invasion, and pathological tumor size. P values of ANOVA in each univariable model were adjusted for multiple comparisons by the BH method.

[0273] Gene set enrichment analyses

[0274] The median value of the variant stabilising transformed count of each gene was calculated per tumor, and gene set enrichment analyses (GSEA) were performed using the fgseaLabel function (fgsea (1. 12.0) R package) to compare differentially expressed gene sets between high-grade predominant tumors and low / mid-grade predominant tumors, and between STAS positive and STAS negative tumors using the Molecular Signatures Database (MSigDB) hallmark gene sets24. Multiple comparisons were adjusted by the BH method and gene sets with q < 0.25 were determined to be significantly enriched as described elsewhere25.

[0275] Enrichment analysis of G2M checkpoint gene SCNA

[0276] Mean ploidy-adjusted copy number of each gene was calculated per tumor region and was compared between predominantly high-grade tumors and low / mid-grade tumors using a linear mixed effect model by setting the ploidy adjusted copy number of each gene as a response variable, predominance of growth pattern (high vs low / mid) as a fixed effect variable, and tumor ID as a random effect using the nlme (3.1-153) R package. ANOVA P values were adjusted for multiple comparisons by the BH method and the genes with q < 0.05 were determined to havesignificantly variant copy numbers between predominantly high- and low / mid-grade tumors. Among the 13728 genes tested, 1032 were significantly gained and 749 were lost in predominantly high-grade tumors. Over-representation of SCNA of G2M checkpoint genes24was tested by chi-square goodness of fit test.

[0277] Survival analysis (TRACERx cohort)

[0278] Disease-free survival (DFS) was defined as the period from the date of registration to the time of radiological confirmation of the recurrence of the primary tumor registered for the TRACERx or the time of death by any cause. During the follow-up, three patients (CRUK0512, CRUK0428, and CRUK0511) developed new primary cancer and subsequent recurrence from either the first primary lung cancer or the new primary cancer diagnosed during the follow-up. These cases were censored at the time of the diagnosis of new primary cancer for DFS analysis, due to the uncertainty of the origin of the third tumor. As for the patients who harboured synchronous multiple primary lung cancers, when associating genomic / pathologic data from the tumors with patient level clinical information, we used only data from the tumor of the highest pathological TMN stage. Hazard ratios and P values were calculated using the coxph function (survival (3.3.1) R package), through multivariable Cox regression analyses, adjusted for age, stage, pack-years, surgery type, and adjuvant therapy. Kaplan-Meier plots were generated using ggsurvplot function (survminer (0.4.9) R package). Intra-thoracic relapse was defined as any relapse found within the thoracic cavity including mediastinum and parietal pleura but not ribs. Extra-thoracic relapse was defined as any relapse found outside the thoracic cavity, including ribs and axillary lymph nodes. To estimate the relapse-site specific (subdistribution) hazard ratio, the Fine-Gray competing risk regression model was applied using cmprsk (2.2-11) R package and adjusted for age, stage, pack-years, surgery type, and adjuvant therapy. In this analysis, relapses at the specific site (intra-thoracic-only or extra-thoracic with or without intra-thoracic) are counted as events, and relapses at other sites or non-lung cancer deaths are treated as competing events. Relapsed cases with uncertain sites and / or uncertain origins were excluded from relapse site-specific analysis.

[0279] Survival analysis (Memorial Sloan Kettering Cancer Center cohort)

[0280] (1) Patient selection

[0281] From the Memorial Sloan Kettering Cancer Center Thoracic Surgery Service prospectively maintained database, 712 patients who had undergone surgical resection for p- stage (TNM staging, version 7) IB to IIIA primary lung adenocarcinoma from January 2006 to December 2014 were identified. Exclusion criteria were: receipt of induction therapy, non- invasive histology, and R2 resection classification.

[0282] (2) Statistical analysis

[0283] The distribution of patient clinicopathologic characteristics was summarised as frequencies and percentages or medians, 25th, and 75thpercentiles for categorical and continuous factors, respectively. Disease-free survival (DFS) was calculated from the date of surgery to the date of any recurrence or death, whichever occurs first. Patients were otherwise censored on the date of the last follow-up. DFS was estimated using the Kaplan-Meier method and compared between groups using the log-rank test. The relationships between factors of interest (STAS status, necrosis status, and the combination of the two) and DFS were quantified using multivariable Cox proportional hazards models, controlling for age at surgery, pathologic stage (1, 2, 3), pack-years, surgery type (sublobar vs lobar or greater), and any adjuvant therapy. All statistical tests were two-sided and P<0.05 was considered statistically significant. Analyses were conducted using Stata 15.1 (StataCorp, College Station, TX) and R 4.1.2 (R Foundation for Statistical Computing, Vienna, Austria).

[0284] Statistical analysis

[0285] All statistical tests were performed in R (version 3.6.1 unless specified otherwise). Tests involving correlations were performed using cor. test with Spearman’s method. Tests involving comparisons of distributions were performed using Wilcoxon rank sum test (or Wilcoxon signed-rank test for paired analysis) or linear mixed effect regression analysis as stated, using wilcox.test or Ime (nlme (3. 1-153) R package) functions, respectively. P values were adjusted for multiple comparisons using the BH method unless otherwise stated, and reported as q values. All statistical tests were two-sided unless specified otherwise, and the numbers of data points included are plotted and / or annotated in the corresponding figures orfigure legends. Plotting was done by using ggplot2 (3.3.5) and ComplexHeatmap (2.2.0) R packages.Example 2: Clonal evolutionary characteristics of LU AD growth patterns

[0286] To elucidate the characteristics of each growth pattern in LU AD, the clinical, pathological, and genomic features across the different predominant subtypes at the whole tumour level were analysed (FIG. 1A, FIG. 5L, data not shown).

[0287] The proportion of high-grade patterns (solid, cribriform, and micropapillary) within a tumour, based on sectional area, was assessed in the context of variables associated with subclonal architecture and genomic instability (FIG. IB, Methods). In TRACERx, an increasing proportion of high-grade patterns at the whole tumour level was associated with higher truncal, but not subclonal, TMB (truncal TMB, Spearman’s rho = 0.25, q value (FDR adjusted P value) = 0.0012, subclonal TMB, rho = 0.11, q = 0.12). Features of chromosomal complexity and instability, including the mean weighted genome instability index (wGII), mean fraction of the genome subject to loss of heterozygosity (FLOH) calculated across multiple regions within a tumour, and the percentage of subclonal SCNAs calculated as the fraction of the aberrant genome which is heterogeneous across tumour regions, were significantly associated with the proportion of high-grade patterns within a tumour (wGII, rho = 0.15, q = 0.039; FLOH, rho = 0.35, q = 2.9 x 10’6, % subclonal SCNA, rho = 0.32, q = 1.5 x 10’3, Methods) (FIG. IB). Similar results were observed when using an orthogonal tool (Sequenza)12for SCNA analysis (FIGs. 7A-7C, Methods).

[0288] In our companion manuscripts, we show that metastasizing subclones tend to be spread across tumour regions, and a recent subclonal expansion score, defined as the largest cancer cell fraction (CCF) of any subclone terminal to the phylogenetic tree in any tumour region, is associated with shorter disease-free survival10,11. In LU AD, measures capturing the presence of a recent subclonal sweep and associated mutational homogeneity within individual tumour regions were significantly associated with an increasing proportion of high-grade patterns (FIG. IB). These included the number of subclonal mutations which are clonal in at least one region (rho = 0.22, q = 0.0035) and the recent subclonal expansion score10(rho =0.16, q = 0.033, Methods). Conversely, the subclonal diversity index, a metric reflective of the number of coexisting subclones in a region and the absence of large recent clonal expansions, was lower in those tumours with a higher proportion of high-grade patterns (minimum subclonal diversity per tumour, rho = -0.23, q = 0.0030, Methods). Ki-67 fraction was also significantly associated with an increasing proportion of high-grade patterns (rho = 0.51, q = 1.6 x 10'11), consistent with previous reports7 13.

[0289] Taken together, these data suggest that while different regions from a high-grade tumour are frequently genomically distinct and can harbour distinct SCNAs, individual regions tend to be highly proliferative and clonally pure, potentially reflective of large intra-regional recent subclonal expansions.

[0290] When the proportion of each individual growth pattern was compared against these genomic features, the presence of highly proliferative and recent subclonal expansion was most strongly associated with the proportion of the solid-pattern component within a tumour (Ki-67 fraction, rho = 0.51, q = 8.1 x 10'11, recent subclonal expansion score, rho = 0.19, q = 0.026) (FIG. 1C). Intriguingly, although micropapillary pattern is regarded as a high-grade pattern that is associated with poor prognosis1,5, an increasing proportion of micropapillary-pattern component was associated with increasing subclonal diversity and a lack of evidence for clear subclonal expansions (minimum subclonal diversity per tumour, rho = 0.22, q = 0.0091, recent subclonal expansion score, rho = -0.18, q = 0.028), suggesting distinct biology and clonal evolutionary characteristics between high-grade solid and micropapillary growth patterns.Example 3: Determinants of inter-tumoural growth pattern heterogeneity

[0291] To explore the evolutionary determinants of growth patterns in LU AD, the presence of truncal genomic alterations was correlated with the predominant pattern within a tumour, assuming that specific early genomic events may influence the subsequent growth pattern. In total, 13 truncal driver alterations (7 mutations, 6 amplifications) and 31 chromosomal arm-level truncal SCNAs (8 gains and 23 losses / LOHs) observed in at least 5% of the cohort and at least 10 tumours in either predominantly high-or low / mid-grade tumours were included in a logistic regression analysis (FIGs. 6A-6C, Methods). Previous studies have reported an increasedfrequency of 7P556 14’15and KRAS16mutations in solid predominant tumours. In the TRACERx cohort, in addition to truncal TP53 and KRAS driver mutations, truncal SMARCA4 mutation, truncal CCNE1 amplification, and truncal loss / LOH of chromosome 21q were associated with predominantly high-grade pattern tumours, while truncal gains of Iq and 8q were associated with predominantly low / mid-grade tumours (FIG. ID). Similar results were observed when only including truncal alterations observed in at least 10% of the cohort, applying an orthogonal tool (Sequenza)12for SCNA profiling, or after adding genomic instability as a covariate in the regression model (FIGs. 7D-7F, Methods).

[0292] Co-occurrence of truncal loss / LOH of 3p and 3q was observed in predominantly low / mid-grade tumours but not in predominantly high-grade tumours, suggesting the whole loss of chromosome 3 may be an early evolutionary event specifically in predominantly low / mid- grade tumours (FIG. IE, FIGs. 7G-7H, Methods). Co-occurrence of truncal 3p and 3q loss / LOH within predominantly low / mid-grade tumours was not associated with higher wGII, suggesting that co-occurrence of the loss of these chromosome arms is not a reflection of genomic instability (P = 0.71, Wilcoxon rank sum test) (FIG. 71). These results may indicate that a cooccurrence of 3p and 3q losses, possibly reflecting the loss of one allele of chromosome 3 as one event, is a distinct evolutionary route to predominantly low / mid-grade tumours.

[0293] To capture SCNAs with <5% frequency in the cohort, which would not have been included in the regression analysis, an expanded analysis of the driver alterations and chromosomal arm-level copy number alterations varying between predominantly high- and low / mid-grade tumours was carried out (Methods). The characteristics most strongly associated with predominantly high-grade tumours versus predominantly low / mid-grade tumours were gains of chromosome arms 3q (whole arm 3q, encompassing the SOX2 and TERC genes) and 12p (whole arm 12p, encompassing the KRAS gene, consistent with a previous report6), both of which were often seen as loss / LOH in predominantly low / mid-grade tumours (FIG. IF, FIGs. 6C-6D). Similar results were observed when applying an orthogonal tool (Sequenza)12for SCNA profiling, or after adding genomic instability as a covariate in the regression model (FIGs. 7J- 7K, Methods).

[0294] In 59% of lepidic predominant invasive adenocarcinomas, a truncal whole genome doubling (WGD) event was detected (FIGs. 6B-6E). By contrast, in the preinvasive lesions, atypical adenomatous hyperplasia (AAH) and adenocarcinoma in situ (AIS), analysed in the TRACERx pre-invasive cohort, <10% demonstrated evidence of WGD17. These data highlight the potential biological differences between alveolar wall surface growth (lepidic pattern) in the setting of pre-invasive disease versus a partly invasive-pattern adenocarcinoma. This suggests that WGD may occur prior to malignant transition from pre-invasive to invasive disease, as previously shown in oesophageal adenocarcinoma18.

[0295] At the transcriptome level, consistent with our observation that increasing Ki-67 fraction is associated with the proportion of high-grade patterns observed at the whole tumour level (FIG. IB), gene-set enrichment analyses revealed the greatest differential expression between high-grade and low / mid-grade predominant tumours involved genes related to cell cycle and cell proliferation (FIG. 1G, Methods), including upregulated G2M checkpoint, mTORCl signalling and PI3K-AKT-mTOR signalling. Notably, predominantly high-grade tumours did not necessarily harbour increased copy number of cell cycle genes compared with predominantly low / mid-grade tumours (significantly increased copy number in high-grade predominant tumours compared with low / mid-grade predominant tumours, G2M checkpoint genes, 6 / 182; all genes, 749 / 13728; P = 0.20, chi-square goodness of fit test) (FIG. 6F, Methods), suggesting that the overexpression of cell cycle genes in predominantly high-grade tumours may not be directly driven by gains of cell cycle genes.

[0296] Although there was no association observed between stromal TIL infiltration and the predominant growth pattern (FIG. 6G), cancer cell PD-L1 expression, assessed by immunohistochemistry, was significantly higher in solid predominant tumours than all other histological subtypes (q = 4.8 x 10'7, Wilcoxon rank sum test, FIG. 6H), as previously reported14. The proportion of solid pattern remained significantly associated with cancer cell expression of PD-L1 after adjustment for potential confounders, including stromal TILs and neoantigen burden (OR= 1.23 [95%CI 1.11-1.38] per 10% increase, P = 3.3 x 10'5, ANOVA, Methods) (FIG. 61), suggesting that overexpression of PD-L1 in solid predominanttumours may be driven by cancer cell intrinsic characteristics, such as AKT-mTOR pathway• • 1 Q activation .Example 4: Morphological intra-tumour heterogeneity reflects genomic intra-tumour heterogeneity

[0297] The data presented thus far have focussed on growth pattern characterisation at the whole tumour level. However, the multi-region sampling and sequencing in TRACERx allows for the analysis of intra-tumour growth pattern heterogeneity in the context of the genomic and transcriptomic landscape. A total of 200 tumours had regional histology data available, amounting to 603 pathological regions used in this analysis (FIG. 1A, FIG. 5J , data not shown).

[0298] To determine whether the variation in morphology within a tumour reflected the degree of subclonal alterations, genomic distances based on mutations (SNVs and indels) and LOH were calculated between regions of the same tumour and explored in relation to the different regional growth patterns (FIG. 2A, Methods). Consistent with findings from the whole tumour analysis that FLOH, but not subclonal TMB, was associated with the proportion of highgrade patterns, the genomic distance between regions with different growth patterns was significantly greater than the genomic distance between regions with the same growth pattern when calculated using LOH (P = 0.0073, linear mixed effect model, ANOVA), but not when using mutations (P = 0.14) (FIG. 2B, FIGs. 7L-7N). This may reflect the fact that LOH is irreversible. By contrast, truncal mutations may be subject to copy number loss, rendering them subclonal and potentially less reliable as a marker of evolutionary divergence.

[0299] To assess whether different morphological patterns reflect an evolutionary trajectory from low-to high-grade pattern, we utilised the irreversible nature of LOH to determine ancestordescendant-like relationships in tumour regions and the context of their respective morphological grades (FIG. 2C, Methods). This was based on the assumption that, although a tumour region is not directly evolved from another region, in some regional pairs one region will harbour a common ancestral-like clone while the other region may harbour a descendant-like clone of the common ancestor. We identified 151 regional pairs with ancestor-descendant-like relationships within 54 tumours using LOH profiles. Comparison of the mutational profiles of these ancestor-descendant-like pairs revealed that descendant-like regions harboured additional mutations at a high cancer cell fraction (CCF > 95%) compared to their ancestor-like counterpart regions (P = 0.007, permutation test, Methods) (FIG. 8A). Within the mixed pattern grade tumours, descendant-like regions exhibited a significantly higher grade compared with their respective ancestral-like regions (P = 0.002, permutation test, Methods) (FIG. 2D, FIG. 8B). This association remained significant when various cut-offs were applied to infer ancestordescendant-like regional pairs (FIG. 8C), or when using an orthogonal tool (Sequenza)12for LOH profiling (FIG. 8D). This association also remained significant when ancestor-descendantlike regional pairs were inferred using a combination of LOH and mutational profiles (FIG.8E, Methods). These data suggest that heterogeneity in morphological patterns may be partially underpinned by genomic alterations. Notably, LUADs did not always follow an evolutionary route towards higher grade patterns, suggesting that, although very rare (7 / 151 ancestordescendant-like regional pairs) (FIG. 8B), tumours may transition to lower grade patterns during sub clonal evolution.Example 5: Distinct evolutionary trajectory of purely solid tumours

[0300] Tumours lacking any adenocarcinoma architectural differentiation, namely purely solid pattern tumours, are typically classified as adenocarcinoma on the basis of immunohistochemical expression of TTF-1 alone. Molecular characteristics of purely solid tumours are poorly understood, and the evolutionary trajectory of purely solid tumours, especially whether these have evolved from mixed pattern adenocarcinomas which include a solid component, remains unclear. To address this, we explored if solid-pattern regions in purely solid tumours and solid-pattern regions in mixed pattern tumours harboured similar genomic and transcriptomic features, and if purely solid tumours were genomically distinct from mixed pattern tumours with a solid component.

[0301] Solid regions in purely solid tumours showed overexpression of G2M checkpoint genes compared with solid regions in other mixed growth pattern tumours (P = 0.0033, linear mixed effect model, ANOVA) (FIG. 9A). At the whole tumour level, the presence of truncal 3q gain (purely solid vs other tumours, OR = 6.2 [95%CI 0.91-32.3], P = 0.031, Fisher’s exact test),especially the truncal focal gain of 3q21.3-3q29 which GISTIC2.0 analysis20detected as a significant peak in tumour regions with solid pattern (purely solid vs other tumours, OR = 10.2 [95%CI 3.0-34.1], P = 7.1 x 10'5, Fisher’s exact test, Methods), and truncal SMARCA4 mutations and / or LOH (purely solid vs other tumours, OR = 6.2 [95%CI 1.7-33.9], P = 0.0015, Fisher’s exact test) were associated with increased likelihood of purely solid tumours (FIGs. 2E- 2F, FIGs. 9B-9I) These results suggest that purely solid tumours have genomic alterations and an evolutionary trajectory distinct from mixed growth pattern tumours with a solid component.Example 6: Evolution of LU AD growth patterns from primary tumour to metastasis

[0302] To determine the relationship between growth patterns in available matched primary and metastatic tumours, 121 metastatic samples from 65 patients were studied specifically with reference to the phylogeny of metastatic clones (FIG. 3A). Metastatic tumours consisted of lymph nodes (LNs) (n=83) and intrapulmonary metastasis (n=2) removed at the time of primary surgery, and sites of disease relapse sampled using diagnostic biopsies including surgical resection (n=36), and were subjected to centralised pathological review and subtyping (proportion of the samples pathologically reviewed for growth pattern, 107 / 121, 88%). The majority of metastatic samples displayed high-grade patterns (82 / 107, 77%) (FIGs. 3A-3B). Each primary tumour region was classified into metastasis seeding or metastasis non-seeding regions based on combined phylogenetic analysis11. In brief, seeding regions were defined as the primary tumour regions harbouring metastasis seeding clones, and the seeding clones were defined as the most recent shared clone between the primary tumour and metastasis.

[0303] Notably, in certain tumours multiple seeding clones were identified, and in certain cases these were spread across all tumour regions. Using a numerical score assigned to high, mid, and low-grade patterns, the mean grade for each seeding region was calculated and compared with the mean grade for each matched metastatic tumour region and primary tumour non-seeding regions (FIG. 3C). While no significant difference was observed using this numerical score to compare seeding and non-seeding regions in the primary tumour (P = 0.096, Wilcoxon signed-rank test) (FIG. 3D), the metastatic regions were typically the same or higher grade compared with their matched seeding regions (P = 1.6 x 10’4, Wilcoxon signed-rank test)(FIG. 3E). One such example was a papillary predominant primary adenocarcinoma (CRUK0543, FIGs. 3F-3G) in which two out of three metastatic lymph nodes (LN#7 and LN#8), deriving from different phylogenetic branches, showed high-grade (cribriform) pattern, while the majority (4 / 5) of the seeding regions showed mid-grade (papillary) pattern, suggestive of parallel evolution of growth patterns during the metastatic cascade.

[0304] A particular case of interest involved a diagnosis of primary lepidic predominant (non-mucinous) adenocarcinoma (CRUK0296) which was found to have a pure lepidic intrapulmonary metastasis during follow-up. After primary resection, the patient developed a subsequent tumour in the contralateral lung, which was surgically resected, and a diagnosis of a pure lepidic pattern LU AD was made. This was deemed a second primary lung cancer due to its 100% lepidic pattern, however, through TRACERx genomic profiling both tumours were found to be clonally related indicating that this was an intrapulmonary metastasis (FIGs. 10A-10D). Notably, pure lepidic intrapulmonary metastasis were not identified in a previous study of 23 intrapulmonary LU AD metastases, although both the primary and metastasis showed part-lepidic patterns in 14 (61%) of the cases21. Furthermore, the primary tumour demonstrated evidence of ‘spread through air spaces’ (STAS) (FIG. 10C), and the presence of a confirmed metastatic lesion that is purely lepidic, and therefore without stromal invasion, supports the hypothesis that free-floating cells are capable of seeding distant tumours through the airways by aerogenous spread. There were five additional patients in whom lung metastases were sequenced upon primary surgery sampling and / or during follow-up, and in which the primary tumour demonstrated positive STAS. Phylogenetic analysis revealed late divergence in all cases (FIG. 10E), which was defined by the divergence of metastatic clone occurred after the last complete clonal sweep within the primary tumour11. Although the predominant primary tumour subtype was unrelated to the timing of metastatic divergence (FIG. 10F), STAS positivity was significantly associated with late divergence (P = 0.019, Fisher’s exact test) (FIG. 10G), suggesting that the ability to metastasize through the airway may be a late event during LU AD evolution, or that tumours acquiring the ability to metastasize through the airway early in their evolution may be rare in our current surgical cohort. Overall, these findings prompted a moredetailed analysis of the relationship between histological pattern, STAS, and patient outcome, including the site of relapse (FIG. 11 A).Example 7: Impact of tumour morphology upon site and risk of recurrence

[0305] The morphological feature of STAS is defined as free-floating tumour cells, or tumour cell clusters, in air spaces beyond the boundary of the tumour, and is known to be associated with intra-thoracic recurrence in limited (sublobar) resections in stage I LU AD3,22,23. Others have reported an association with STAS and poor prognosis in more advanced stage LU AD24,25as well as in non-LUAD histologies26,27. In a multivariable analysis of the TRACERx 421 LU AD cohort, disease-free survival (DFS) of STAS positive cases was shorter than STAS negative cases (HR=2.2 [95%CI 1.4-3.6], adjusted for age, stage, pack-years, surgery type, and adjuvant therapy) (FIG. 11B).

[0306] STAS positivity was associated with the presence of high-grade pattern in each tumour (r / = 0.0096, univariate logistic regression, ANOVA, FDR adjusted; Methods), and was associated with micropapillary patterns (q = 0.0096) (FIGs. 11C-11D), as described in other cohorts3,24. Immunohistochemical nuclear beta-catenin positivity and an epithelial to mesenchymal transition (EMT) phenotype has previously been shown to be associated with STAS28. Driver mutations in the Wnt pathway were enriched in STAS positive tumours, (q = 0.033, Fisher’s exact test, FDR adjusted) (FIG. 12A), and the bulk tumour transcriptomic profiles showed higher CTNNB1 gene expression (P = 0.0076, linear mixed effect model, ANOVA) (FIG. 12B). However, we did not observe enrichment of EMT pathway or Wnt-beta- catenin signalling gene expression modules in STAS positive tumours (FIG. 12C), potentially due to the difficulty of capturing phenotypic differences related to STAS using bulk transcriptomic data.

[0307] The presence of pre-operative ctDNA is known to be associated with increased risk of relapse in LUAD29. In our companion manuscript, we show the presence of preoperative ctDNA is particularly associated with extra-thoracic recurrence30, which may reflect the increased risk of hematogenous metastatic dissemination. In a subset of the LUAD cohort excluding the patients with synchronous primary lung cancers (136 / 242 patients), pre-operative ctDNA detection fromtwo assays (53 patients with an assay previously reported by our group in Abbosh et al9, and 90 patients with an assay reported in our companion manuscript30, including 7 patients analysed in both assays), and STAS status were integrated to compare the biological features of these two prognostic indicators in relation to the risk and site of metastasis (FIG. 4A). Patients with both STAS positivity and pre-operative ctDNA detection had primary tumours enriched for predominantly high-grade patterns (P = 7.0 x 10'5, Fisher’s exact test) (FIG. 4B).

[0308] Detection of pre-operative ctDNA was associated with the presence of high-grade patterns (q = 5.4 x 10'4, univariate logistic regression, ANOVA, FDR adjusted; Methods), in particular solid (q = 1.0 x 1 O'6) and cribriform (q = 0.008) patterns, and a lack of lepidic (q = 2.7x10'5) and acinar patterns (q = 0.0011) (FIGs. 11D-11E), consistent with previously reported radiological characteristics in ctDNA shedding tumours29. As described in an earlier TRACERx cohort, histological evidence of necrosis (q = 2.2 x 1 O'13), tumour size (q = 2.8 x 1 O'3), Ki-67 fraction (q = 9.9 x 10'7), mitotic index (q = 1.1 x 10'4), degree of nuclear pleomorphism (nuclear grade, q = 2.3 x 10'4), and the presence of pleural invasion (q = 0.0014) were associated with pre-operative ctDNA detection9, but not with STAS positivity (FIG. 11D). Whilst predominantly high-grade tumours were associated with shorter DFS than predominantly low / mid-grade tumours (HR = 1.7 [95%CI 1.1 -2.6], multivariable Cox regression) (FIG. HF), predominance of high-grade pattern was not significantly associated with relapse site (FIGs. 11G-11H). In contrast, the presence of micropapillary pattern was associated with intra-thoracic-only recurrence (subdistribution HR = 2.3 [95%CI 1.1-4.6], multivariable Fine-Gray regression) and the presence of solid and / or cribriform patterns was associated with extra-thoracic recurrence (subdistribution HR = 3.2 [95%CI 1.1 -9.4]) (FIG. 4C, FIG. HI), consistent with findings in stage I LUADs reported previously31,32. Similarly, STAS positivity was associated with increased risk of intra-thoracic-only recurrence (subdistribution HR = 3.0 [95%CI 1.0-9.1], multivariable Fine-Gray regression), but not extra-thoracic recurrence (subdistribution HR = 2.0 [95%CI 0.7-5.4]). Although it is worth noting that our cohort may be underpowered to detect the risk of extra-thoracic recurrence in STAS positive tumours, which has been reported previously in a larger cohort24,25. Pre-operative ctDNA detection was associated with extra-thoracic recurrence (subdistribution HR = 4.6 [95%CI 1.5-13.8], multivariable Fine-Gray regression) butnot intra-thoracic-only recurrence (subdistribution HR = 1.2 [95%CI 0.3-3.9]) (FIG. 4D), as reported in our companion manuscript30. Of note, STAS was detected in 17 out of 21 patients (81%) who had disease relapse despite having undetectable pre-operative ctDNA, which was significantly higher than STAS detection in tumours with undetectable pre-operative ctDNA and no subsequent relapse (37 / 70, 52.8%) (P = 0.024, Fisher’s exact test) (FIG. 4E, FIGs. 13A- 13B)

[0309] Patients with both STAS positivity and pre-operative ctDNA detection had an increased risk of disease relapse compared to patients in whom neither were detected (HR = 8.1 [95%CI 3.2-20.6], multivariable Cox regression) (FIG. 4F). Both STAS positivity and preoperative ctDNA detection were independent predictors of prognosis in a multivariable analysis that included age, stage, pack-years, surgery type, and adjuvant therapy (STAS, HR = 3.4 [95%CI 1.8-6.4]; pre-operative ctDNA, HR = 2.4 [95%CI 1.3-4.2], multivariable Cox regression) (FIG. 13C). These results suggest that whilst STAS positivity and pre-operative ctDNA detection are both associated with disease recurrence, the underlying biology of the metastatic process in tumours with each of these characteristics is distinct. Furthermore, the combination of STAS positivity and pre-operative ctDNA detection has the potential to identify patients with an increased risk of relapse during follow-up, independent of TNM staging, and is therefore of potential clinical utility (FIGs. 13D-13E).

[0310] Finally, since histological evidence of necrosis was more significantly associated with pre-operative ctDNA detection than any other histological feature (FIG. 11D), we tested whether necrosis could be used as a proxy for pre-operative ctDNA detection. The presence of necrosis was associated with solid and cribriform predominant tumours (solid / cribriform vs others, P = 3.7 x 10"13, Fisher’s exact test) (FIG. Ill) and an increased risk of extra-thoracic recurrence (subdistribution HR = 2.9 [95%CI 1.5-5.6], multivariable Fine-Gray regression), but not intra- thoracic-only recurrence (subdistribution HR = 1.6 [95%CI 0.7-3.6]) (FIGs. 13F-13G). As a combined measure, these two histological features remained significant independent predictors of outcome in a multivariable analysis (STAS, HR = 2.4 [95%CI 1.5-3.9]; necrosis, HR = 2.1 [95%CI 1.3-3.2], multivariable Cox regression) (FIG. 13H). The combination of STAS andnecrosis demonstrated that patients with tumours positive for both had an increased risk of disease relapse (HR = 5.8 [95%CI 3.0-11.4] versus patients negative for both, multivariable Cox regression) (FIG. 131). A similar result was observed in a larger independent external cohort of surgically resected stage IB-IIIA LUADs (n = 712, HR = 2.0 [95%CI 1.4-2.8] versus patients negative for both, multivariable Cox regression) (FIGs. 14A-14C). The combination of STAS and necrosis may therefore have clinical value in predicting metastatic risk in patients in the absence of pre-operative ctDNA sampling and analysis.Example 8: ctDNA / STAS biomarkers accurately predicts risk of recurrence and / or metastases in lung cancer patients

[0311] Machine learning model details. Multiple machine learning models such as Random survival forest (RSF; Ishwaran et al The Annals of Applied Statistics 2008), Extreme Gradient Boosting (XGB) etc., to predict recurrence / metastases risk are implemented. Models are implemented in python using the sksurv library. Input variables include STAS classification and as well as liquid biopsy-related parameters (i.e. logiocfDNA concentration and max VAF as continuous variables. The STAS classification input can be based on pathologist assessments or through machine learning predictions. Models are trained and validated using 5-fold cross validation. Performance will be assessed using C-index, Fl score, recall, precision, specificity, sensitivity, and area under the receiver operating characteristic curve (AUC). among others.

[0312] Without wishing to be bound by theory, it is believed that the inclusion of data from ctDNA sequencing assays in multivariable machine learning models including cell-free (cf)DNA concentrations, and other STAS would improve prediction of recurrence and metastases. The ability of ctDNA / STAS for early intervention using nonrandomized, real-world evidence is assessed. It is anticipated that the models with achieve high sensitivity and specificity metrics and will reduce the frequency of clinical errors.EQUIVALENTS

[0313] The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can bemade without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0314] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0315] As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 cells refers to groups having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to groups having 1, 2, 3, 4, or 5 cells, and so forth.

[0316] All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification.

Claims

CLAIMS1. A method of training a machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, comprising: a. receiving data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; b. generating a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peri turn oral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and c. applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset; wherein the classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients, optionally wherein the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

2. The method of claim 1, wherein the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis.

3. The method of claim 1 or 2, wherein the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects.

4. The method of claim 3, wherein the lung tissue sections are frozen sections or permanent sections.

5. The method of any of claims 1-4, wherein the STAS comprises one or more of clustercell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS.

6. The method of any of claims 1-5, wherein the status of STAS is determined by a thoracic pathologist.

7. The method of any one of claims 1-6, wherein the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique.

8. The method of any one of claims 1-7, wherein the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model.

9. The method of any one of claims 1-8, wherein the machine learning classifier is an ensemble learning random forest classifier.

10. The method of any one of claims 1-9, wherein the machine learning technique models survival outcomes with competing risks.

11. The method of any one of claims 1-10, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

12. The method of any one of claims 1-11, further comprising applying the classifier to data on a lung cancer patient to generate a predictor, and determining whether the lung cancerpatient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

13. The method of claim 12, wherein the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

14. The method of claim 12 or 13, further comprising performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

15. The method of claim 12 or 13, further comprising performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

16. The method of any one of claims 2-15, wherein the histological subtype of lung adenocarcinoma is selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern.

17. The method of any one of claims 2-16, wherein the adjuvant therapy comprises chemotherapy or immunotherapy.

18. The method of any one of claims 2-17, wherein the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

19. The method of any one of claims 2-18, wherein the surgery type is pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection.

20. A method of estimating risk of recurrence and / or metastases in a lung cancer patient using a machine learning classifier, the method comprising: a. receiving patient data corresponding to a plurality of features for the lung cancer patient; b. applying the machine learning classifier to the patient data to generate a predictor; andc. determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the machine learning classifier is trained by: i. receiving cohort data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; ii. generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and iii. applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients.

21. The method of claim 20, further comprising performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

22. The method of claim 20, further comprising performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

23. The method of any one of claims 21-22, wherein the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

24. The method of any one of claims 20-23, wherein the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis.

25. The method of any one of claims 20-24, wherein the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects.

26. The method of claim 25, wherein the lung tissue sections are frozen sections or permanent sections.

27. The method of any of claims 20-26, wherein the STAS comprises one or more of clustercell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS.

28. The method of any of claims 20-27, wherein the status of STAS is determined by a thoracic pathologist.

29. The method of any of claims 24-28, wherein the histological subtype of lung adenocarcinoma is selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern.

30. The method of any one of claims 24-29, wherein the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

31. The method of any one of claims 21-30, further comprising administering an effective amount of an adjuvant therapy to the lung cancer patient, optionally wherein the adjuvant therapy comprises chemotherapy or immunotherapy.

32. The method of any one of claims 20-30, wherein the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

33. The method of any one of claims 20-32, wherein the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique.

34. The method of any one of claims 20-33, wherein the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model.

35. The method of any one of claims 20-34, wherein the machine learning classifier is an ensemble learning random forest classifier.

36. The method of any one of claims 20-35, wherein the machine learning technique models survival outcomes with competing risks.

37. The method of any one of claims 20-36, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

38. The method of any one of claims 20-37, wherein one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA39. The method of any one of claims 20-38, wherein one or more of the plurality of features for each subject in the cohort are determined by assaying blood and / or sequencing tumor DNA.I l l40. A machine learning system for training a machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients, the system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients; wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients.

41. The machine learning system of claim 40, wherein the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, alinear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique.

42. The machine learning system of any one of claims 40-41, wherein the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model.

43. The machine learning system of any one of claims 40-42, wherein the machine learning classifier is an ensemble learning random forest classifier.

44. The machine learning system of any one of claims 40-43, wherein the machine learning technique models survival outcomes with competing risks.

45. The machine learning system of any one of claims 40-44, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

46. The machine learning system of any one of claims 40-45, wherein the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis.

47. The machine learning system of any one of claims 40-46, wherein the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects.

48. The machine learning system of claim 47, wherein the lung tissue sections are frozen sections or permanent sections.

49. The machine learning system of any of claims 40-48, wherein the STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non- circumferential STAS.

50. The machine learning system of any of claims 40-49, wherein the status of STAS is determined by a thoracic pathologist.

51. The machine learning system of any one of claims 40-50, wherein the instructions further cause the processor to apply the classifier to data on a lung cancer patient to generate a predictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

52. The machine learning system of claim 51, wherein the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

53. The machine learning system of any one of claims 40-51, wherein the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

54. The machine learning system of any one of claims 40-51, wherein the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

55. The machine learning system of any one of claims 46-54, wherein the histological subtype of lung adenocarcinoma is selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern.

56. The machine learning system of any one of claims 46-55, wherein the adjuvant therapy comprises chemotherapy or immunotherapy.

57. The machine learning system of any one of claims 46-56, wherein the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

58. The machine learning system of any one of claims 46-57, wherein the surgery type is pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection.

59. The machine learning system of any one of claims 40-58, wherein the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

60. A computing system for estimating risk of recurrence and / or metastases in lung cancer patients, the computing system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive patient data corresponding to a plurality of features for the lung cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the classifier is trained by: receiving cohort data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients.

61. The computing system of claim 60, wherein the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique.

62. The computing system of claim 60 or 61, wherein the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model.

63. The computing system of any one of claims 60-62, wherein the machine learning classifier is an ensemble learning random forest classifier.

64. The computing system of any one of claims 60-63, wherein the machine learning technique models survival outcomes with competing risks.

65. The computing system of any one of claims 60-64, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

66. The computing system of any one of claims 60-65, wherein the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

67. The computing system of any one of claims 60-65, wherein the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

68. The computing system of any one of claims 60-67, wherein the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

69. The computing system of any one of claims 60-68, wherein the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis.

70. The computing system of any one of claims 60-69, wherein the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects.

71. The computing system of claim 70, wherein the lung tissue sections are frozen sections or permanent sections.

72. The computing system of any of claims 60-71, wherein the STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS.

73. The computing system of any of claims 60-72, wherein the status of STAS is determined by a thoracic pathologist.

74. The computing system of any of claims 60-73, wherein the histological subtype of lung adenocarcinoma is selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern.

75. The computing system of any one of claims 60-74, wherein the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

76. The computing system of any one of claims 60-75, wherein the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

77. The computing system of any one of claims 60-76, wherein one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA.

78. A non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a machine learning system, configure the machine learning system to train a machine learning classifier to estimate risk of recurrence and / or metastases in lung cancer patients, the instructions configured to cause the processor to:receive data on a cohort of lung adenocarcinoma (LU AD) subjects, each LU AD subject in the cohort comprising at least one lung tumor; generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases in lung cancer patients; wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients.

79. The computer-readable storage medium of claim 78, wherein the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique.

80. The computer-readable storage medium of claim 78 or 79, wherein the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model.

81. The computer-readable storage medium of any one of claims 78-80, wherein the machine learning classifier is an ensemble learning random forest classifier.

82. The computer-readable storage medium of any one of claims 78-81, wherein the machine learning technique models survival outcomes with competing risks.

83. The computer-readable storage medium of any one of claims 78-82, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

84. The computer-readable storage medium of any one of claims 78-83, wherein the plurality of features further comprises surgery type, histological subtype of lung adenocarcinoma, adjuvant therapy, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis.

85. The computer-readable storage medium of any one of claims 78-84, wherein the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LU AD subjects.

86. The computer-readable storage medium of claim 85, wherein the lung tissue sections are frozen sections or permanent sections.

87. The computer-readable storage medium of any of claims 78-86, wherein the STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS.

88. The computer-readable storage medium of any of claims 78-87, wherein the status of STAS is determined by a thoracic pathologist.

89. The computer-readable storage medium of any one of claims 78-88, wherein the instructions further cause the processor to apply the machine learning classifier to data ona cancer patient to generate a predictor, and determining whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

90. The computer-readable storage medium of claim 89, wherein the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

91. The computer-readable storage medium of any one of claims 78-90, wherein the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

92. The computer-readable storage medium of any one of claims 78-90, wherein the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

93. The computer-readable storage medium of any one of claims 84-92, wherein the histological subtype of lung adenocarcinoma is selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern.

94. The computer-readable storage medium of any one of claims 84-93, wherein the adjuvant therapy comprises chemotherapy or immunotherapy.

95. The computer-readable storage medium of any one of claims 84-94, wherein the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

96. The computer-readable storage medium of any one of claims 84-95, wherein the surgery type is pneumonectomy, bilobectomy, lobectomy, segmentectomy, or wedge resection.

97. The computer-readable storage medium of any one of claims 78-96, wherein the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

8. A non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a computing system, configure the computing system to estimate risk of recurrence and / or metastases in lung cancer patients, the instructions configured to cause the processor to: receive patient data corresponding to a plurality of features for the lung cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the lung cancer patient is at risk for recurrence and / or metastases based on the predictor and an operating-point threshold, wherein the classifier is trained by: receiving cohort data on a cohort of lung adenocarcinoma (LUAD) subjects, each LUAD subject in the cohort comprising at least one lung tumor; generating a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) cell free DNA concentration, (ii) maximum ctDNA VAF, and (iii) the status of aerogenous spread of tumor cells in peritumoral lung parenchyma beyond an edge of the at least one lung tumor (STAS); and applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of recurrence and / or metastases, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of receiver operating characteristic (ROC) curves for the training dataset;wherein the machine learning classifier is configured to receive the plurality of features for lung cancer patients and generate predictors for risk of recurrence and / or metastases in lung cancer patients.

99. The computer-readable storage medium of claim 98, wherein the machine learning technique is a random forest technique, a decision tree technique, a logistic regression technique, a linear regression technique, a nearest neighbor technique, an artificial neural network technique, or a support vector machine technique.

100. The computer-readable storage medium of claim 98 or 99, wherein the one or more machine learning models are selected from among a supervised model, an unsupervised model or a reinforcement model.

101. The computer-readable storage medium of any one of claims 98-100, wherein the machine learning classifier is an ensemble learning random forest classifier.

102. The computer-readable storage medium of any one of claims 98-101, wherein the machine learning technique models survival outcomes with competing risks.

103. The computer-readable storage medium of any one of claims 98-102, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

104. The computer-readable storage medium of any one of claims 98-103, wherein the instructions further cause the processor to recommend performing a lobectomy on the lung cancer patient predicted to be at high risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

105. The computer-readable storage medium of any one of claims 98-103, wherein the instructions further cause the processor to recommend performing a segmentectomy on the lung cancer patient predicted to be at low risk for recurrence and / or metastases based on the predictor and the operating-point threshold.

106. The computer-readable storage medium of any one of claims 98-105, wherein the predictor comprises a cumulative incidence function (CIF) for lung recurrence and / or metastases.

107. The computer-readable storage medium of any one of claims 98-106, wherein the plurality of features further comprises histological subtype of lung adenocarcinoma, cancer stage, spatial distance of STAS from the at least one lung tumor, and the presence of necrosis.

108. The computer-readable storage medium of any one of claims 98-107, wherein the status of STAS is identified via histopathologic analysis of lung tissue sections obtained from the cohort of the LUAD subjects.

109. The computer-readable storage medium of claim 108, wherein the lung tissue sections are frozen sections or permanent sections.

110. The computer-readable storage medium of any of claims 98-109, wherein the STAS comprises one or more of cluster-cell STAS, single-cell STAS, circumferential STAS, or non-circumferential STAS.

111. The computer-readable storage medium of any of claims 98-110, wherein the status of STAS is determined by a thoracic pathologist.

112. The computer-readable storage medium of any of claims 107-111, wherein the histological subtype of lung adenocarcinoma is selected from among a lepidic pattern, a papillary pattern, an acinar pattern, a cribriform pattern, a micropapillary pattern, and a solid pattern.

113. The computer-readable storage medium of any one of claims 107-1 12, wherein the cancer stage is Stage 1, Stage 2, Stage 3, or Stage 4.

114. The computer-readable storage medium of any one of claims 98-113, wherein the metastases comprise metastasis of one or more organs selected from among adrenal gland, bone, brain, liver, lymph, and pleura.

15. The computer-readable storage medium of any one of claims 98-114, wherein one or more of the plurality of features for the lung cancer patient are determined by assaying blood and / or sequencing tumor DNA.

Citation Information

Patent Citations

  • Method for analyzing cell-free nucleic acid and application thereof

    CN115443341A

  • Data based cancer research and treatment systems and methods

    US20230223121A1

  • Composite metastasis score with weighted coefficients for predicting breast cancer metastasis, and uses thereof

    US8557525B1