Selecting cancer therapies based on levels of intratumoral escherichia in cancer patients
Patent Information
- Application Number
- PCT/US2025/029878
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-05-16
- Publication Date
- 2026-01-29
AI Technical Summary
Existing methods struggle to accurately determine cancer patients' responsiveness to immune checkpoint inhibitor (ICI) therapy due to challenges in detecting and distinguishing intratumoral microbiota, particularly Escherichia, from environmental contaminants like bacteriophages, using formalin-fixed, paraffin-embedded (FFPE) samples.
A computational system and method that filters genome sequence reads to distinguish true Escherichia reads from bacteriophage reads, using a selective and specific filter, to predict patient responsiveness to ICI therapy based on intratumoral Escherichia levels.
The system enables statistically significant and accurate prediction of improved overall survival and progression-free survival in cancer patients treated with ICI therapy, by identifying individuals with high intratumoral Escherichia levels, and suggests alternative therapies for those with low levels.
Smart Images

Figure US2025029878_29012026_PF_FP_ABST
Abstract
Description
[0001] Atty. Dkt. No.: 115872-3221 SELECTING CANCER THERAPIES BASED ON LEVELS OF INTRATUMORAL ESCHERICHIA IN CANCER PATIENTS CROSS REFERENCE TO RELATED APPLICATIONS The application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 649,138, filed May 17, 2024, which is incorporated herein by reference in its entirety. GOVERNMENT SUPPORT This invention was made with government support under CA008748 and CA251591 awarded by the National Institutes of Health. The government has certain rights in the invention. TECHNICAL FIELD Disclosed herein are computational systems and methods of selecting an anti- cancer therapy for a patient in need thereof based on levels of intratumoral Escherichia. BACKGROUND The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology. Consistent with its role in the pathogenesis of a variety of human diseases, the host microbiome is increasingly recognized as a hallmark of cancer and a determinant of therapeutic outcomes. The gut microbiota profiling of patients with cancer has demonstrated bacterial signature associated with favorable and unfavorable response to immune check point inhibitor (ICI) therapy. Specifically, in patients with advanced non-small cell lung cancer (NSCLC), presence of Akkermansia muciniphila was found to be enriched in patients with -1- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 response to ICI confirmed in a prospective, observational trial. While there has been extensive literature and microbiota profiling of the gut microbiota, few studies have evaluated the role of the intratumor microbiota and response to ICI. Difficulties associated with detecting intratumoral microbiota (Nejman, D. et al. Science 368, 973-980 (2020); Heymann, C. J. F. et al. Cancer Letters 522, 63-79 (2021) have hindered our ability to explore associations between intratumoral bacteria and ICI response. Characterization of the lung tissue microbiome has been a challenge, as 16s ribosomal (r)RNA sequencing has broad limitations when performed on clinically available formalin-fixed, paraffin-embedded (FFPE) samples. Accordingly, there is an urgent need for methods for determining whether cancer patients would be responsive to ICI therapy based on intratumoral microbiota. SUMMARY Aspects of the present disclosure are directed to systems and methods of selecting a cancer therapy in a cancer subject. One or more processors may obtain, for a subject at risk of or diagnosed with cancer, a dataset comprising genome sequence reads of a biological sample obtained from the cancer subject. The one or more processors may identify, from the dataset identifying the genome sequence reads, a first portion of the genome sequence reads corresponding to candidate Escherichia reads. The one or more processors may filter genome sequence reads corresponding to a bacteriophage from the first portion of the genome sequence reads to generate a second portion of the genome sequence reads corresponding to true Escherichia reads. The one or more processors may determine, using the second portion of the genome sequence reads, a score indicating a number of true reads for Escherichia. The one or more processors may select, from a plurality of anti-cancer therapies, the anti-cancer therapy for the cancer subject based on the score relative to a fixed threshold or a determined threshold from a control group. The one or more processors may store, using one or more data structures, an association between the cancer subject and the anti-cancer therapy. -2- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 In some embodiments, the one or more processors may generate the determined threshold for the control group based on a plurality of number of reads of Escherichia from a corresponding plurality of non-cancer subjects. In some embodiments, the one or more processors may determine that the score satisfies at least one of the fixed threshold or the determined threshold. In some embodiments, the one or more processors may select the anti- cancer therapy including immune checkpoint inhibitor (ICI) therapy, responsive to determining that the score satisfies at least one of the fixed threshold or the determined threshold. In some embodiments, the one or more processors may determine that the score does not satisfy the threshold. In some embodiments, the one or more processors may select the anti-cancer therapy including a combination therapy of the ICI therapy and chemotherapy, responsive to determining that the score does not satisfy the threshold. In some embodiments, the one or more processors may classify the cancer subject into one of (i) a first group corresponding to a presence of intratumoral Escherichia or (ii) a second corresponding to an absence of the intratumoral Escherichia, based on comparing the number of reads with the threshold. In some embodiments, the one or more processors may select the anti-cancer therapy based on classifying the cancer subject into one of the first group or the second group. In some embodiments, the plurality of anti-cancer therapies may include at least one of (i) an ICI therapy or (ii) a combination therapy of the ICI therapy and chemotherapy. In some embodiments, the ICI therapy may include at least one of a PD-L1 antibody, a PD-1 antibody, or a CTLA-4 antibody. In some embodiments, the ICI therapy may include one or more of pembrolizumab, nivolumab, cemiplimab, atezolizumab, avelumab, durvalumab, ipilimumab, tremelimumab, ticlimumab, JTX-4014, Spartalizumab (PDR001), Camrelizumab (SHR1210), Sintilimab (IBI308), Tislelizumab (BGB-A317), Toripalimab (JS 001), Dostarlimab (TSR-042, WBP-285), INCMGA00012 (MGA012), AMP-224, AMP-514, KN035, CK-301, AUNP12, CA-170, or BMS-986189. In some embodiments, administration of the selected cancer therapy including the ICI therapy to the cancer subject may result in an improved overall survival (OS), relative to an OS of a second cancer subject administered with ICI therapy associated with a second score -3- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 indicating a number of true reads for Escherichia not satisfying the threshold. In some embodiments, administration of the selected cancer therapy including the ICI therapy to the cancer subject results in an improved progression-free survival (PFS), relative to a PFS of a second cancer subject administered with ICI therapy and associated with a second score indicating a number of true reads for Escherichia not satisfying the threshold. In some embodiments, the one or more processors may filter the sequence reads by identifying, from the first portion of the genome sequence reads, the sequence reads attributable to the bacteriophage and the true Escherichia reads. In some embodiments, the bacteriophage may include ΦX (phiX). In some embodiments, the true Escherichia reads are associated with E. coli. In some embodiments, the one or more processors may obtain the dataset comprising the genome sequence reads of the biological sample identified as having the candidate Escherichia reads via E. coli fluorescence in situ hybridization (FISH). In some embodiments, FISH may include detecting E. coli nucleic acids using a probe of SEQ ID NO: 1. In some embodiments, the cancer may include lung cancer. In some embodiments, the lung cancer optionally may include non-small cell lung cancer (NSCLC). The cancer may be one of an early stage, a localized stage, an advanced stage, or a metastasized stage in the cancer subject. In some embodiments, the biological sample may be obtained from at least one of a primary cancer site or a metastatic site associated with the cancer. BRIEF DESCRIPTION OF THE DRAWINGS The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which: FIGs. 1A–1C. Intratumoral Escherichia is associated with improved overall survival in univariable and multivariable analyses in patients treated with single-agent immune checkpoint inhibition (MSK discovery cohort). A. Overall survival in Escherichia- positive vs negative groups in patients treated with single-agent immune checkpoint blockade. -4- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 The patients were identified using a filter to distinguish between Escherichia reads and bacteriophages. B. Multivariable analysis for overall survival in patients treated with single- agent immune checkpoint inhibition. Escherichia-pos; Escherichia-positive group. Escherichia- neg; Escherichia-negative group. ECOG PS; Eastern Cooperative Oncology Group performance status. F; female. M; male. OS; overall survival. mOS; median OS. MSK; Memorial Sloan Kettering Cancer Center. C. An overall survival in Escherichia-positive vs negative groups in patients treated with single-agent immune checkpoint blockade. The filter to distinguish between Escherichia reads and bacteriophages was not used to identify the individuals. As shown, without the use of the filter, the statistical significance drops from p=0.0065 as in A to p=0.1 as in C. FIGs. 2A and 2B. Intratumoral Escherichia is associated with improved overall survival in multivariable analyses in patients treated with single-agent immune checkpoint inhibition (FMI validation cohort). A. Overall survival in Escherichia-positive vs negative groups in patients treated with single-agent immune checkpoint blockade. B. Multivariable analysis for overall survival in patients treated with single-agent immune checkpoint inhibition Escherichia-pos; Escherichia-positive group. Escherichia-neg; Escherichia-negative group. ECOG PS; Eastern Cooperative Oncology Group performance status. F; female. M; male. OS; overall survival. mOS; median OS. FMI; Foundation Medicine. The FMI analyses have been adjusted for delayed entry to mitigate left truncation bias. As such, not all patients are displayed as at risk at Time 0 in the bottom at-risk tables. FIGs. 3A–3D. Intratumoral visualization of Escherichia coli A. Colon adenocarcinoma, resection specimen from sigmoid colon, 13237 Escherichia reads. B. LUAD, section from left lower lobe resection; 0 Escherichia reads. C. LUAD, resection specimen from right middle lobe; 24 Escherichia reads. D. LUAD, biopsy of bladder metastasis, 166 Escherichia reads. Fluorescence (top sections): Blue - DAPI; Red - E.Coli. Hematoxylin and eosing staining shown below. All images captured at same magnification. Representative images shown. -5- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 FIGs. 4A–4D. Bulk RNA sequencing reveals potential associations between intratumoral Escherichia and tumor microenvironment. A. Volcano plot showing differentially expressed genes (DEGs) between Escherichia-positive (pos) and Escherichia- negative (neg) samples. Grey dots represent non-significant value; green dots represent DEG significant by p-value but not by fold-change threshold; blue dots represent DEG significant by fold-change threshold but not by p-value; red dots represent significant DEG by both fold- change and p-value threshold. B. Boxplots of relevant DEGs between Escherichia-positive and Escherichia-negative samples. P-value by Wilcox test. C. Volcano plot showing cell type enrichment between Escherichia-pos and Escherichia-neg samples. D. Pathway enrichment analysis between Escherichia-pos and Escherichia-neg samples. All p-values adjusted (p-adj) for false discovery rate. FIG. 5. Cohort assembly diagram for MSK cohort. NGS; next-generation sequencing. ICI monotherapy means single-agent ICI with either pembrolizumab, nivolumab, atezolizumab, or durvalumab. Chemo-ICI means combination of platinum doublet chemotherapy in combination with ICI. Regimens that were excluded were regimens incorporating dual ICI (ipilimumab plus nivolumab), or exploratory regimens including ICI in combination with TKI or other small molecules. FIG. 6. Overview of pipeline to detect bacteria within tumors using formalin-fixed, paraffin embedded (FFPE) tissue. Overview of the pipeline, described in detail in the Methods. Biopsies from either primary or metastatic tumor samples undergo formalin fixation and paraffin embedding. DNA is extracted and sequenced using the MSK- IMPACT targeted next-generation sequencing platform. Tumor BAM files are then queried for unmapped reads and aligned against the NCBI BLASTN database. FIG. 7. Mitigating contamination. No template control (NTC) analysis (n=2,378) comparing Escherichia bacterial reads across MSK-IMPACT versions 3, 5 and 6. NTCs are samples that only include reagents. Y-axis represents readcount at the genus level for -6- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Escherichia detected within these NTCs. X-axis represents the number of NTCs with Escherichia reads. This analysis was only available within the MSK cohort. FIGs. 8A–8C. Univariable and multivariable analysis for association of progression-free survival in patients treated with monotherapy ICI in the MSK Cohort. A. Univariable progression-free survival analysis. The patients were identified using a filter to distinguish between Escherichia reads and bacteriophages. B. Multivariable progression-free survival analysis. Escherichia-pos; Escherichia-positive group. Escherichia-neg; Escherichia- negative group. ECOG PS; Eastern Cooperative Oncology Group performance status. F; female. M; male. PFS; progression-free survival. mPFS; median PFS. MSK; Memorial Sloan Kettering Cancer Center. C. Univariable progression-free survival analysis. The filter to distinguish between Escherichia reads and bacteriophages was not used to identify the individuals. As shown, without the use of the filter, the statistical significance drops from p=0.0057 as in A to p=0.2 as in C. FIGs. 9A–9D. Intratumoral Escherichia is not associated with progression- free or overall survival in univariable and multivariable analyses in patients treated with platinum-doublet chemotherapy in combination with ICI (MSK discovery cohort). A. Overall survival and B. progression-free survival in Escherichia-positive vs negative groups in patients treated with Chemo-ICI. Multivariable analysis in patients treated with Chemo-IO for C. overall survival and D. progression-free survival. Escherichia-pos; Escherichia-positive group. Escherichia-neg; Escherichia-negative group. ECOG PS; Eastern Cooperative Oncology Group performance status. F; female. M; male. PFS; progression-free survival. OS; overall survival. mPFS; median PFS. mOS; median OS. MSK; Memorial Sloan Kettering Cancer Center. FIG. 10. Cohort assembly for FMI Validation cohort. FIGs. 11A and 11B. Univariable and multivariable analysis for association of progression-free survival in patients treated with monotherapy ICI in the FMI Cohort. -7- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 A. Univariable progression-free survival analysis. B. Multivariable progression-free survival analysis. Escherichia-pos; Escherichia-positive group. Escherichia-neg; Escherichia-negative group. ECOG PS; Eastern Cooperative Oncology Group performance status. F; female. M; male. PFS; progression-free survival. mPFS; median PFS. FMI; Foundation Medicine. FIGs. 12A–12D. Intratumoral Escherichia is not associated with progression-free or overall survival in univariable and multivariable analyses in patients treated with platinum-doublet chemotherapy in combination with ICI (FMI validation cohort). A. Overall survival and B. progression-free survival in Escherichia-positive vs negative groups in patients treated with Chemo-ICI. Multivariable analysis in patients treated with Chemo-IO for C. overall survival and D. progression-free survival. Escherichia-pos; Escherichia-positive group. Escherichia-neg; Escherichia-negative group. ECOG PS; Eastern Cooperative Oncology Group performance status. F; female. M; male. PFS; progression-free survival. OS; overall survival. mPFS; median PFS. mOS; median OS. FMI; Foundation Medicine. FIG. 13. E. coli FISH signal as measured by the E. coli+-cell ratio according to E. coli status on MSK-IMPACT. Readcounts of E. coli as detected by MSK-IMPACT are printed near the datapoint. FIG. 14 Graph of Escherichia-pos and Escherichia-neg groups by overall survival rate over a timeframe of months. Both groups of patients were treated with a combination of immune checkpoint inhibitor (ICI) treatments (e.g., CTLA-4 or anti-PD1 treatment). As shown, while Escherichia-pos group of patents showed improvement relative to the Escherichia-neg group of patients when administered with the combination of ICI treatments, the statistical significance is low (p = 0.37). FIG. 15 depicts a block diagram of a system for selecting cancer therapies for subjects based on Escherichia reads from biological samples, in accordance with an illustrative embodiment. -8- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 FIG. 16 depicts a block diagram of a process for parsing genome sequence data in the system for selecting cancer therapies based on Escherichia reads, in accordance with an illustrative embodiment. FIG. 17 depicts a block diagram of a process for providing outputs based on classifications from genome sequence data in the system for selecting cancer therapies based on Escherichia reads, in accordance with an illustrative embodiment. FIG. 18 depicts a flow diagram of a method of selecting cancer therapies for subjects based on Escherichia reads from biological samples, in accordance with an illustrative embodiment. FIG. 19 depicts a block diagram of a server system and a client computer system, in accordance with one or more implementations. DETAILED DESCRIPTION Following below are more detailed descriptions of various concepts related to, and embodiments of, systems and methods for selecting cancer therapies for subjects based on Escherichia reads from a biological sample. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways, as the disclosed concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes. The impact of the intratumoral microbiome on immune checkpoint inhibitor (ICI) efficacy in patients with cancer (e.g., non-small cell lung cancer (NSCLC)) is unknown. For instance, the gut microbiota profiling of patients with cancer has demonstrated bacterial signature associated with both favorable and unfavorable response to ICI. Furthermore, it can be challenge to detect certain intratumoral microbiota, such as Escherichia (e.g., E. coli), because genomic data correlated with Escherichia can also be associated with other conflicting sequence -9- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 readouts (e.g., bacteriophage reads). To address these issues, the present disclosure provides for systems and methods for selecting anti-cancer therapies for subjects based on Escherichia reads from a biological sample. It is further shown that the use of a filter detailed herein to discriminate between Escherichia and bacteriophages lead to statistically significant and accurate predictions in identifying likely responders to ICI therapy. The present disclosure demonstrates that the presence of intratumoral Escherichia determined from genomic data that are filtered using a highly selective, precise, and specific filter can be used to select individuals for ICI. Individuals with high intratumoral Escherichia levels are shown to have significant improvement in overall survival (OS) and progression -free survival (PFS) when administered with ICI, relative to individuals with low intratumoral Escherichia levels. Alternatively, individuals with low intratumoral Escherichia levels are likely to benefit from anti-cancer therapies beyond ICI. Preclinically, intratumoral Escherichia is associated with a proinflammatory tumor microenvironment and decreased metastases. The determination of whether intratumoral Escherichia was associated with outcome to ICI in patients with NSCLC was investigated. The intratumoral microbiome in 958 patients with advanced NSCLC treated with ICI was examined by querying unmapped next-generation sequencing reads against a bacterial genome database. Putative environmental contaminants were filtered using “no template” controls (n=2,378). The impact of intratumoral Escherichia detection on overall survival (OS) was assessed in univariable and multivariable analyses. Findings were further validated in an external independent cohort of 772 patients. Escherichia fluorescence in situ hybridization (FISH) and transcriptomic profiling were performed. As demonstrated herein, failure to accurately filter out the Escherichia genomic datasets for environmental contaminants (e.g., conflicting bacteriophage sequence reads) obscured the ability to accurately discriminate between ICI responders and ICI non-responders based on the original Escherichia genomic datasets. See FIGs. 1A versus 1C, and FIGs. 8A versus 8C. -10- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 In the discovery cohort, filtered reads mapping to intratumoral Escherichia were associated with significantly longer OS (16 vs 11 months, HR 0.73, 95% CI 0.59-0.92, p=0.0065) in patients treated with single-agent ICI, but not combination chemo-immunotherapy. The association with OS remained statistically significant in multivariable analysis adjusting for prognostic features including PD-L1 expression (p=0.023). Analysis of an external validation cohort confirmed the association with improved OS in univariable and multivariable analyses of patients treated with single-agent ICI, and not in patients treated with chemo-immunotherapy. Escherichia localization within tumor cells was supported by co-registration of FISH staining and serial H&E sections. Transcriptomic analysis correlated Escherichia-positive samples with expression signatures of immune cell infiltration. Reads mapping to potential intratumoral Escherichia was associated with survival to single-agent ICI in two independent cohorts of patients with NSCLC. Definitions Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. For example, reference to “a cell” includes a combination of two or more cells, and the like. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, analytical chemistry and nucleic acid chemistry and hybridization described below are those well-known and commonly employed in the art. As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1%, 5%, or 10% in either direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context (except where such number would be less than 0% or exceed 100% of a possible value). -11- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 As used herein, the “administration” of an agent or drug to a subject includes any route of introducing or delivering to a subject a compound to perform its intended function. Administration can be carried out by any suitable route, including but not limited to, orally, intranasally, parenterally (intravenously, intramuscularly, intraperitoneally, or subcutaneously), rectally, intrathecally, intratumorally or topically. Administration includes self-administration and the administration by another. The term “bacteriophage”, often called a phage, is a virus that infects and replicates within bacterial cells. These viruses are ubiquitous and can be found in various environments where bacteria exist. The terms “cancer” or “tumor” are used interchangeably and refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be a non-tumorigenic cancer cell. As used herein, the term “cancer” includes premalignant, as well as malignant cancers. In some embodiments, the cancer is bladder cancer, breast cancer, colorectal cancer, esophagogastric cancer, gynecological cancer (e.g., uterine cancer, cervical cancer, ovarian cancer), head and neck cancer, hepatobiliary cancer, high-grade glioma, low-grade glioma, lung cancer, melanoma, pancreatic cancer, prostate cancer, renal cancer, or soft tissue sarcoma. In some embodiments, the cancer is lung cancer (e.g., non-small cell lung cancer (NSCLC) such as adenocarcinoma, squamous cell carcinoma, or large cell carcinoma). “Detecting” as used herein refers to determining the presence of a mutation, alteration, or level of a nucleic acid of interest in a sample. Detection does not require the method to provide 100% sensitivity. Analysis of nucleic acid markers can be performed using techniques known in the art including, but not limited to, sequence analysis, and electrophoretic analysis. Non-limiting examples of sequence analysis include Maxam-Gilbert sequencing, Sanger sequencing, capillary array DNA sequencing, thermal cycle sequencing (Sears et al., -12- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Biotechniques, 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al., Methods Mol. Cell Biol, 3:39-42 (1992)), sequencing with mass spectrometry such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et al., Nat. Biotechnol, 16:381-384 (1998)), and sequencing by hybridization. Chee et al., Science, 274:610- 614 (1996); Drmanac et al., Science, 260:1649-1652 (1993); Drmanac et al., Nat. Biotechnol, 16:54-58 (1998). Non-limiting examples of electrophoretic analysis include slab gel electrophoresis such as agarose or polyacrylamide gel electrophoresis, capillary electrophoresis, and denaturing gradient gel electrophoresis. Additionally, next generation sequencing methods can be performed using commercially available kits and instruments from companies such as the Life Technologies / Ion Torrent PGM or Proton, the Illumina HiSEQ or MiSEQ, and the Roche / 454 next generation sequencing system. As used herein, the term “effective amount” refers to a quantity sufficient to achieve a desired therapeutic and / or prophylactic effect, e.g., an amount which results in the prevention of, or a decrease in a disease or condition described herein or one or more signs or symptoms associated with a disease or condition described herein. In the context of therapeutic or prophylactic applications, the amount of a composition administered to the subject will vary depending on the composition, the degree, type, and severity of the disease and on the characteristics of the individual, such as general health, age, sex, body weight and tolerance to drugs. The skilled artisan will be able to determine appropriate dosages depending on these and other factors. The compositions can also be administered in combination with one or more additional therapeutic compounds. In the methods described herein, the therapeutic compositions may be administered to a subject having one or more signs or symptoms of a disease or condition described herein. As used herein, a "therapeutically effective amount" of a composition refers to composition levels in which the physiological effects of a disease or condition are ameliorated or eliminated. A therapeutically effective amount can be given in one or more administrations. The term “Escherichia” refers to a genus of bacteria within the family Enterobacteriacea. The bacteria of the Escherichia genus may be gram-negative and -13- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 facultatively anaerobic. The term “intratumoral Escherichia” refers to Escherichia associated with tumor or cancer, for example, found at a primary cancer site or a metastatic site. The term “Escherichia reads” refer to the genomic sequence reads in a genome sequence dataset that match or correspond to Escherichia. In some embodiments, Escherichia is Escherichia coli (E. coli). “Next-generation sequencing or NGS” as used herein, refers to any sequencing method that determines the nucleotide sequence of either individual nucleic acid molecules (e.g., in single molecule sequencing) or clonally expanded proxies for individual nucleic acid molecules in a high throughput parallel fashion (e.g., greater than 103, 104, 105 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of the nucleic acid species in the library can be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment. Next generation sequencing methods are known in the art, and are described, e.g., in Metzker, M. Nature Biotechnology Reviews 11:31-46 (2010). As used herein, a “sample” refers to a substance that is being assayed for the presence of a nucleic acid of interest. Processing methods to release or otherwise make available a nucleic acid for detection are well known in the art and may include steps of nucleic acid manipulation. A biological sample may be a body fluid or a tissue sample. In some cases, a biological sample may consist of or comprise blood, plasma, sera, urine, feces, epidermal sample, vaginal sample, skin sample, cheek swab, sperm, amniotic fluid, cultured cells, bone marrow sample, tumor biopsies, aspirate and / or chorionic villi, cultured cells, and the like. Fresh, fixed or frozen tissues may also be used. In one embodiment, the sample is preserved as a frozen sample or as formaldehyde- or paraformaldehyde-fixed paraffin-embedded (FFPE) tissue preparation. For example, the sample can be embedded in a matrix, e.g., an FFPE block or a frozen sample. -14- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 As used herein, the terms “subject”, “patient”, or “individual” can be an individual organism, a vertebrate, a mammal, or a human. In some embodiments, the subject, patient or individual is a human. As used herein, the term “therapeutic agent” is intended to mean a compound that, when present in an effective amount, produces a desired therapeutic effect on a subject in need thereof. “Treating” or “treatment” as used herein covers the treatment of a disease or disorder described herein, in a subject, such as a human, and includes: (i) inhibiting a disease or disorder, i.e., arresting its development; (ii) relieving a disease or disorder, i.e., causing regression of the disorder; (iii) slowing progression of the disorder; and / or (iv) inhibiting, relieving, or slowing progression of one or more symptoms of the disease or disorder. In some embodiments, treatment means that the symptoms associated with the disease are, e.g., alleviated, reduced, cured, or placed in a state of remission. It is also to be appreciated that the various modes of treatment of disorders as described herein are intended to mean “substantial,” which includes total but also less than total treatment, and wherein some biologically or medically relevant result is achieved. The treatment may be a continuous prolonged treatment for a chronic disease or a single, or few time administrations for the treatment of an acute condition. NGS Platforms In some embodiments, high throughput, massively parallel sequencing employs sequencing-by-synthesis with reversible dye terminators. In other embodiments, sequencing is performed via sequencing-by-ligation. In yet other embodiments, sequencing is single molecule sequencing. Examples of Next Generation Sequencing techniques include, but are not limited to pyrosequencing, Reversible dye-terminator sequencing, SOLiD sequencing, Ion semiconductor sequencing, Helioscope single molecule sequencing etc. -15- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 The Ion TorrentTM (Life Technologies, Carlsbad, CA) amplicon sequencing system employs a flow-based approach that detects pH changes caused by the release of hydrogen ions during incorporation of unmodified nucleotides in DNA replication. For use with this system, a sequencing library is initially produced by generating DNA fragments flanked by sequencing adapters. In some embodiments, these fragments can be clonally amplified on particles by emulsion PCR. The particles with the amplified template are then placed in a silicon semiconductor sequencing chip. During replication, the chip is flooded with one nucleotide after another, and if a nucleotide complements the DNA molecule in a particular microwell of the chip, then it will be incorporated. A proton is naturally released when a nucleotide is incorporated by the polymerase in the DNA molecule, resulting in a detectable local change of pH. The pH of the solution then changes in that well and is detected by the ion sensor. If homopolymer repeats are present in the template sequence, multiple nucleotides will be incorporated in a single cycle. This leads to a corresponding number of released hydrogens and a proportionally higher electronic signal. The 454TM GS FLX TM sequencing system (Roche, Germany), employs a light-based detection methodology in a large-scale parallel pyrosequencing system. Pyrosequencing uses DNA polymerization, adding one nucleotide species at a time and detecting and quantifying the number of nucleotides added to a given location through the light emitted by the release of attached pyrophosphates. For use with the 454TM system, adapter-ligated DNA fragments are fixed to small DNA-capture beads in a water-in-oil emulsion and amplified by PCR (emulsion PCR). Each DNA-bound bead is placed into a well on a picotiter plate and sequencing reagents are delivered across the wells of the plate. The four DNA nucleotides are added sequentially in a fixed order across the picotiter plate device during a sequencing run. During the nucleotide flow, millions of copies of DNA bound to each of the beads are sequenced in parallel. When a nucleotide complementary to the template strand is added to a well, the nucleotide is incorporated onto the existing DNA strand, generating a light signal that is recorded by a CCD camera in the instrument. -16- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Sequencing technology based on reversible dye-terminators: DNA molecules are first attached to primers on a slide and amplified so that local clonal colonies are formed. Four types of reversible terminator bases (RT-bases) are added, and non-incorporated nucleotides are washed away. Unlike pyrosequencing, the DNA can only be extended one nucleotide at a time. A camera takes images of the fluorescently labeled nucleotides, then the dye along with the terminal 3' blocker is chemically removed from the DNA, allowing the next cycle. Sequencing by synthesis (SBS), like the "old style" dye-termination electrophoretic sequencing, relies on incorporation of nucleotides by a DNA polymerase to determine the base sequence. A DNA library with affixed adapters is denatured into single strands and grafted to a flow cell, followed by bridge amplification to form a high-density array of spots onto a glass chip. Reversible terminator methods use reversible versions of dye- terminators, adding one nucleotide at a time, detecting fluorescence at each position by repeated removal of the blocking group to allow polymerization of another nucleotide. The signal of nucleotide incorporation can vary with fluorescently labeled nucleotides, phosphate-driven light reactions and hydrogen ion sensing having all been used. Examples of SBS platforms include Illumina GA and HiSeq 2000. The MiSeq® personal sequencing system (Illumina, Inc.) also employs sequencing by synthesis with reversible terminator chemistry. In contrast to the sequencing by synthesis method, the sequencing by ligation method uses a DNA ligase to determine the target sequence. This sequencing method relies on enzymatic ligation of oligonucleotides that are adjacent through local complementarity on a template DNA strand. This technology employs a partition of all possible oligonucleotides of a fixed length, labeled according to the sequenced position. Oligonucleotides are annealed and ligated and the preferential ligation by DNA ligase for matching sequences results in a dinucleotide encoded color space signal at that position (through the release of a fluorescently labeled probe that corresponds to a known nucleotide at a known position along the oligo). This method is primarily used by Life Technologies’ SOLiDTM sequencers. Before sequencing, the DNA is amplified by emulsion PCR. The resulting beads, each containing only copies of the same DNA molecule, are deposited on a solid planar substrate. -17- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 SMRT^ sequencing is based on the sequencing by synthesis approach. The DNA is synthesized in zero-mode wave-guides (ZMWs)-small well-like containers with the capturing tools located at the bottom of the well. The sequencing is performed with use of unmodified polymerase (attached to the ZMW bottom) and fluorescently labeled nucleotides flowing freely in the solution. The wells are constructed in a way that only the fluorescence occurring at the bottom of the well is detected. The fluorescent label is detached from the nucleotide at its incorporation into the DNA strand, leaving an unmodified DNA strand. A. Association of Intratumoral Escherichia with improved survival to single- agent immune checkpoint inhibition in patients with advanced non-small cell lung cancer The impact of the intratumoral microbiome on immune checkpoint inhibitor (ICI) efficacy in patients (pts) with non-small cell lung cancer (NSCLC) is unknown. In preclinical studies, the presence of lung intratumoral Escherichia was associated with a proinflammatory tumor microenvironment and decreased metastases within lung tissue. It was sought to detect intratumoral bacteria in pts with advanced NSCLC using hybrid capture-based, next generation sequencing (NGS). 849 pts treated with ICI-based therapy who underwent NGS were studied. Unmapped reads were extracted from BAM files, and these were queried for bacteria (blastn alignment using the NCBI database). Putative environmental contaminants were subtracted from the analysis using “no template” controls (n=2,539) to exclude possible artifactual false positives. A custom E.Coli fluorescence in situ hybridization (FISH) probe was used to visualize Escherichia within the tumors after co-registration with H&E. In 849 pts, a median of 30 bacterial reads was detected per sample (inter-quartile range (18-85)). Among 68 pts with paired primary / metastatic samples, the bacterial spectra were similar in both sites, suggesting that tumor resident bacteria might travel with cancer cells to distant sites. Antibiotic use within 30 days of tumor sampling was associated with decreased intratumoral bacterial diversity (p=0.023 by Inverse Simpson, p=0.038 by Shannon). Intratumoral Escherichia was associated with better PFS (HR 0.78, 95% CI 0.62-0.98, p=0.036), and OS (HR 0.74, 95% CI 0.58-0.95, -18- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 p=0.017) in pts treated with single-agent ICI, but not combination Chemo / ICI. In a multivariable model adjusting for prognostic features in NSCLC including PD-L1 tumor proportion score, the presence of intratumoral Escherichia was associated with better PFS (p=0.040) and OS (p=0.045) upon single-agent ICI therapy. Escherichia appeared to be intracellular based on co- registration of FISH staining and serial H&E sections. These findings warrant further investigation of the possible inter-relationships between intratumoral Escherichia, tumor immune micro-environment, and ICI therapeutic outcomes. Purpose Consistent with its role in the pathogenesis of a variety of human diseases, the host microbiome is increasingly recognized as a hallmark of cancer and a determinant of therapeutic outcomes1. The gut microbiota profiling of patients with cancer has demonstrated bacterial signature associated with favorable and unfavorable response to ICI. Specifically, in patients with advanced non-small cell lung cancer (NSCLC), presence of Akkermansia muciniphila was found to be enriched in patients with response to ICI2, confirmed in a prospective, observational trial3. While there has been extensive literature and microbiota profiling of the gut microbiota, few studies have evaluated the role of the intratumor microbiota and response to ICI. Like the gut epithelium, the human airway and lungs interface with the external environment and are colonized by a complex variety of microorganisms known as the lung microbiome. The most extensively reported organisms are anaerobes and aerotolerant bacteria, with pathogenic gram-negative bacteria also commonly detected, particularly in immunocompromised individuals4,5. Seminal work characterized the lung microbiome of healthy subjects via broncho-alveolar lavage and found associations between certain taxa and elevated Th-17 lymphocyte response6, suggesting that the lung microbiome may influence host inflammation and lung carcinogenesis7,8. In contrast to the increasing data demonstrating associations between the gut microbiome and outcome to ICI in NSCLC, associations between the lung microbiome and outcome to ICI have not been previously described. -19- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Difficulties associated with detecting intratumoral microbiota9,10have hindered the ability to explore associations between intratumoral bacteria and ICI response. Characterization of the lung tissue microbiome has been a challenge, as 16s ribosomal (r)RNA sequencing has broad limitations when performed on clinically available formalin-fixed, paraffin-embedded (FFPE) samples. Culture-free techniques could facilitate the discovery of lung microbiome / disease interactions with the utilization of next-generation sequencing (NGS) platforms. Recent pre-clinical work in a B16 melanoma mouse model found that Escherichia within lung parenchyma was associated with decreased formation of lung metastases and a pro- inflammatory tumor microenvironment11. More recent data also demonstrated that the presence of intratumoral Escherichia was associated with favorable outcomes in patients with bladder cancer treated with neoadjuvant chemotherapy12. Therefore, the a priori hypothesis was that intratumoral Escherichia would be associated with improved outcomes to ICI in patients with NSCLC. The primary objective of this study was to determine whether detection of off-target reads intratumorally, specifically, Escherichia was associated with survival to ICI in patients with NSCLC. Patients and Methods Patients The initial discovery cohort consisted of patients with NSCLC treated at Memorial Sloan Kettering Cancer Center (MSK). A total of 6,900 patients with NSCLC who underwent NGS testing with MSK-IMPACT13were assessed for eligibility. Patients with advanced (defined as either stage IV or stage III not amenable to definitive treatment) NSCLC treated with anti-PD-L1-based therapy and who underwent NGS with the MSK-IMPACT platform at the institution were included. The validation cohort consisted of patients with NSCLC whose tumors were profiled at Foundation Medicine Inc. (FMI). For the FMI cohort, a total of 11,458 patients with -20- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 NSCLC who received tumor biopsy-based comprehensive genomic profiling (CGP) using the FoundationOne®CDx assay targeting ~324 cancer-related genes14were assessed for eligibility. This cohort included a subset of 1,345 patients from the Flatiron Health-Foundation Medicine NSCLC clinico-genomic database (FH-FMI CGDB)15. This nationwide (US-based) database contained data collected between January 2011 and December 2022 that originated from approximately 280 US cancer clinics (~800 sites of care). Additional details regarding clinical data abstraction from the FMI validation cohort are detailed in the Supplementary Methods. For both the MSK and FMI cohorts, survival analyses were assessed from start of treatment until progression (for real-world progression-free survival (rwPFS)) or death (for real- world overall survival (rwOS) and rwPFS). Next-Generation Sequencing Next-generation sequencing was performed as previously described for both the MSK-IMPACT and FoundationOne®CDx platforms14,16and summarized in the Supplementary Methods. Details for off-target bacterial read detection are detailed in the Supplementary Methods. Fluorescence in situ hybridization Details for fluorescence in situ hybridization are provided in the Supplementary Methods. RNA sequencing Details for RNA sequencing are provided in the Supplementary Methods. Statistical analysis Patients in the single-agent ICI and platinum-doublet combination with ICI (Chemo-ICI) cohorts were analyzed separately given known differences in baseline characteristics and clinical outcomes between these two groups17. Statistical analyses were -21- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 performed using R version 4.1.1. Survival rates were estimated using a Kaplan-Meier estimator, and survival curves were compared using a log-rank test. Consistent with prior studies18, a multivariable analysis was performed, adjusting for standard prognostic factors in both ICI monotherapy and Chemotherapy-ICI, which are age, sex, histology, line of therapy, ECOG performance, and PD-L1 tumor proportion score. Multivariable analyses were performed using the Cox regression model to determine HRs and 95% CIs for PFS and OS while adjusting for known clinicopathologic features using the coxph package and visualized using the forest plot package. Only patients with information available for all variables were included in the multivariable analysis. All tests were two sided, and statistical significance was set at p<0.05. Two patients in the MSK cohort were not evaluable for PFS. Tables were visualized using the gt_summary package.19Results Patient characteristics and detection of intratumoral Escherichia in non-small cell lung cancer samples A total of 1,462 patients with advanced NSCLC treated with anti-PD-L1-based immunotherapy with NGS testing were assessed for inclusion (FIG. 5). Median age was 67, and 52% (757) were female (Table 5). Most patients (1,283; 90%) had Eastern Cooperative Oncology Group performance status 0 – 1, harbored tumors with adenocarcinoma histology (1,137; 78%) and had a history of current or former smoking (1,127; 84%). Raw BAM files from NGS of biopsy samples were processed as described in the Supplemental Methods and in FIG. 6. Since environmental contamination is an established pitfall of microbiome studies, this limitation was addressed with analyses of no template controls (NTCs): samples with reagents but no tumor tissue, as detailed in the Methods. Escherichia contamination within the NTCs (N=2,378 samples, Table S2) did not exceed 6 bacterial reads in the most recent version of MSK-IMPACT (version 6, FIG. 7). Because later versions of MSK-IMPACT demonstrated significantly lower environmental contamination (Table S2, FIG. 7), all analyses were restricted to samples sequenced with the latest version of MSK-IMPACT, version 6 (FIG. 5). Baseline -22- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 characteristics were similar for patients in the MSK-IMPACT version 6 cohort (n=958) compared to the global cohort (Table S3). Intratumoral Escherichia is associated with favorable outcome to single-agent ICI Given pre-clinical data suggesting a beneficial role of intratumoral Escherichia in modulating the tumor microenvironment, the primary objective was to determine whether intratumoral Escherichia was associated with improved outcome to ICI. A threshold of readcount>6 (lack of background contamination above this threshold within MSK-IMPACT version 6 tested NTCs) was used to determine Escherichia status (positive or negative) (FIG. 7). Patients treated with anti-PD-L1 monotherapy (single-agent ICI) or with platinum-doublet chemotherapy in combination with anti-PD-L1 (Chemo-ICI) were included (FIG. 5). Aside from differences in histology (p=0.004 in the single-agent ICI group), there were no differences in baseline clinical characteristics between the Escherichia-positive and negative groups (Table 1). There were no differences between in the Escherichia-positive and negative groups for medications of interest including antibiotics use, inhaled corticosteroids, COPD, or post- progression therapy administration (Table S4). In patients treated with single-agent ICI, intratumoral Escherichia was associated with significantly longer OS (median OS 16 vs 11 months, respectively; HR 0.73, 95% CI 0.59, 0.92, p=0.0065; FIG. 1A) and marginally longer PFS (median PFS 3.2 vs 2.3 months, respectively; hazard ratio [HR] 0.82, 95% CI 0.66-0.101, p=0.057; FIG. 8A). To adjust for potential confounding features (Table 1) including previously defined prognostic features associated with ICI benefit in patients with NSCLC, a multivariable analysis was performed. Consistent with univariable survival results, among patients treated with single-agent ICI, intratumoral Escherichia was significantly associated with improved OS (HR 0.76, p=0.026; FIG. 1B), but not PFS (FIG. 8B). In patients treated with combination Chemo-ICI, intratumoral Escherichia was not associated with OS (median OS 13 vs 15 months, respectively; HR 0.97, 95% CI 0.78, 1.20, p=0.80) (FIG. 9A) or PFS (median PFS 5.5 vs 5.5 months, respectively; HR 1.00, 95% CI 0.83, 1.21, p=0.99) (FIG. 9B). There were no associations between presence of Escherichia in the Chemo-ICI-treated cohort for OS (HR 0.97, p=0.78) (FIG. 9C) or PFS (HR 1.02, p=0.89) (FIG. 9D) in multivariable analyses. Although -23- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 analyses were performed at the genus level, all of the Escherichia reads belonged to E. coli. These data suggest that intratumoral reads mapping to Escherichia, particularly E. coli, is associated with improved survival to single-agent ICI in patients with advanced NSCLC, even after adjusting for standard clinico-pathologic features. Intratumoral Escherichia is associated with favorable outcome to single-agent ICI in an independent cohort Next, the role of intratumoral Escherichia on outcome to therapy in an independent cohort was assessed. Among 1,345 patients with advanced NSCLC that underwent next-generation sequencing using FoundationOne CDx®assay, 779 were treated with anti-PD- L1-based regimens, alone (n=444) or in combination with platinum doublet chemotherapy (n=328) (FIG. 10). Clinical features in this cohort were similar to those of the MSK cohort (Table S5), with lower median age in the Escherichia-positive cohort. Escherichia-positivity was defined by the median split (>22 reads) in this cohort due to lack of no template control samples within this cohort. Presence of intratumoral Escherichia was associated with improved OS in univariable (median OS 16.0 vs 10.8 months, HR, 0.68, p=0.001) (FIG. 2A) and multivariable analyses (HR=0.73, 95%CI 0.54,0.97, p=0.032) (FIG. 2B). Furthermore, intratumoral Escherichia was associated with improved PFS in univariable (median PFS 7.3 vs 4.7 months, HR 0.74, 95%CI 0.61-0.91, p=0.004) (FIG. 11A) and multivariable analyses (HR 0.73, 95%CI 0.57-0.94, p=0.014) (FIG. 11B). Consistent with the MSK cohort, intratumoral Escherichia was not associated with outcome to combination Chemo-ICI for OS (p=0.3) (FIG. 12A) or PFS (p=0.7) (FIG. 12B) or in multivariable analyses for OS (FIG. 12C) or PFS (FIG. 12D). Visualization of intratumoral Escherichia coli To confirm the presence of E. coli in NSCLC samples positive by NGS using an orthogonal approach, FFPE sections from representative NSCLC cases, Escherichia-positive and Escherichia-negative samples by NGS were probed with E. coli 16s rRNA and fluorescence in -24- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 situ hybridization followed by H&E staining to co-register fluorescent signal. As a positive control, a colon biopsy was included. E. coli was demonstrated in the tumor acini of the colon adenocarcinoma, contiguous with the lumen of the colon (FIG. 3A) with 13,237 reads of Escherichia detected by the MSK-IMPACT method and subsequent strong positive fluorescent signal for Escherichia. No definitive fluorescent signal for E. coli was found in the LUAD with 0 reads by the MSK-IMPACT method (FIG. 3B). In contrast, a primary sample (LUAD resection, 26 reads of Escherchia detected by MSK-IMPACT method; FIG. 3C) and a metastatic LUAD sample collected from the bladder dome (165 reads of Escherichia detected by MSK- IMPACT method; FIG. 3D) exhibited fluorescence, with the bacteria appearing within tumor cells adjacent to the nucleus, suggesting that bacteria may be found within or immediately proximal to the tumor cells. These findings were consistent across a total of 32 samples, with read count detected by MSK-IMPACT associated with E. coli detection on FISH (p=0.017) (FIG. 13). Exploratory transcriptomic analysis reveals signal for enhanced immune infiltration among Escherichia-positive samples In an exploratory analysis, RNA sequencing data from 73 patients with NSCLC was examined to examine potential effects of the presence of Escherichia on the tumor immune microenvironment. Intratumoral Escherichia was assessed using the same NGS method described above. In this subset, 58 patients (79%) were Escherichia-positive. Genes associated with favorable response to ICI enriched in the Escherichia-positive samples included GZMB20,21, associated with enhanced response to ICI and increased immune cell infiltration in NSCLC and other malignancies, CCL20, encoding a cytokine with airway tropism22known to facilitate recruitment of T cells, B cells, and NK cells into lung tumors23. Other cytokine and chemokine genes such as CXCR2P124, CXCL1312, and IL12RB225associated with response to ICI were also differentially expressed in the Escherichia-positive group, and in the case of CXCL13, previously associated with intratumoral E. coli-specific CXCL13-Producing T follicular helper cells12. Genes implicated in the function of active regulatory T cells were differentially expressed in the Escherichia-positive group such as FOXP326and SERPINE227(FIG. 4A-B). -25- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Genes demonstrating differentially elevated expression in Escherichia-negative samples included FREM228, PRKCZ28and USP4429, all of which have previously defined immune-suppressive functions and have been associated with T cell-barren tumors. Other differentially expressed genes associated with the Escherichia-negative samples included CAMK1D30and TNIK31, previously associated with ICI resistance (FIG. 4A-B). Interestingly, marker genes for T cells, B cells, as well as specific subpopulations of T cells such as cytotoxic T cells were enriched in Escherichia-positive samples (FIG. 4C). Several relevant pathways to T cell activation were enriched in Escherichia-positive samples (FIG. 4D). Taken together, these exploratory results suggest that intratumoral Escherichia may promote an inflammatory tumor microenvironment. Discussion In this large analysis of patients with advanced NSCLC treated with ICI, off- target reads mapping to intratumoral Escherichia were associated with favorable survival to single-agent ICI, but not combination platinum-doublet Chemo-ICI in two independent cohorts. The results extend preclinical evidence that intratumoral Escherichia may contribute to a more immunostimulatory tumor microenvironment within the lung. These observations can also be placed within the context of prior investigations intentionally deploying Escherichia as cancer- targeting bacteria. E. coli expressing the bacterial cytolytic pore-forming toxin ClyA markedly suppressed metastatic tumor growth and prolonged survival time in mice, and E. coli colonized hypoxic and necrotic tumor regions32. Most recently, Goubet et al. reported that E. coli-specific CXCL13-producing follicular helper CD4+ T cells in the tumor microenvironment of bladder cancer are associated with clinical efficacy of neoadjuvant PD-1 blockade33. Manipulation of the tumor microbiome could ultimately be a novel strategy to induce or augment anti-tumor immunity. Landmark studies examining the pan-tumor microbiome have employed bacteria specific enrichment techniques such as PCR amplification of the 16S rRNA sequence9. These high-sensitivity analyses opened the field of intratumoral microbiome analysis and demonstrated the high prevalence of tumor-associated bacteria. However, bacterial detection in FFPE tissue -26- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 by 16S rRNA analysis is challenging because of the high level of nucleic acid degradation which often reduces the DNA fragment size below the span of amplification primers. Also, amplification can lead to spurious quantification as the efficiency of different PCR amplification products varies. Likewise, amplification of low-level contamination events can lead to false positive signals. The technique developed in this study allows for convenient secondary use of NGS data obtained through stringently controlled clinical NGS assays. However, differences in microbiome detection technique may account for differences in Escherichia prevalence observed across studies— Nejman et al found that the prevalence of Escherichia spp was 36% in patients with a smoking history (Table S12, Nejman et al). Laroumagne et al.34showed that gram- negative bacteria, such as Haemophilus influenzae, Enterobacter spp. and Escherichia coli, colonized 39.8% of 216 lung cancer bronchoscopic samples, with E. coli making up 8%. The differences in prevalence between studies, including within this study, could be due to detection method and sample source, which is an ongoing limitation and challenge in interpretation of intratumoral microbiota studies, especially with studies using NGS. Furthermore, DNA extraction for clinical sequencing may introduce a bias towards identification of gram-negative organisms: the thicker peptidoglycan layers of gram-positive bacteria may bias against detection of these organisms. Despite being a large study including a discovery and a validation cohort, the study has limitations with regards to lack of adjustment for contamination in the FMI cohort due to lack of no-template controls. Moreover, further studies will be needed to arrive at a consensus for definition of Escherichia positivity as different thresholds were used between both cohorts due to lack of NTCs in the FMI cohort. Therefore, prospective validation with orthogonal confirmation of Escherichia are warranted. In addition, given continuity with the airway and lung parenchyma, the presence of bacteria is expected within lung tumors, and E. coli has been described in settings such as hospital-acquired pneumonia or in immunocompromised individuals. However, total bacterial load is likely lower in lung tumors compared to other tumors such as the colon. The limitations of sequencing low biomass samples have been clearly documented and these limitations also apply to the study. An additional limitation for the MSK -27- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 cohort is that the precise date of sample collection was not available. However, among patients with available data, 95% of samples were obtained pre-treatment. This limitation was addressed in the FMI validation cohort, where date of sample collection was available for all samples, and this was adjusted for using left truncation analysis. With regards to the exploratory transcriptomic analyses, the limitation that transcriptomic results were not included in the multivariable analysis is acknowledged. In this study, data obtained from the clinical next-generation sequencing platform was leveraged. Alternative approaches for exploring the microbiome include independent 16S RNA sequencing or the development of a distinct hybridization capture assay, followed by the recapture of DNA libraries. Both methods necessitate resequencing, limited by DNA yield associated with the small size of lung cancer biopsies. Notably, 16S RNA sequencing is challenging for formalin-fixed paraffin-embedded tissue due to fragmented DNA, resulting in sizes significantly smaller than the distance between PCR probes. There are distinct advantages to employing clinical NGS for this purpose including high throughput on readily available clinical specimens. However, there are certain drawbacks to consider; the sequencing approach does not specifically target Escherichia. Additionally, while sterile procedures were rigorously maintained during FFPE tissue processing, paraffin blocks are not continuously stored in sterile environments, introducing the potential for contamination. In addition, there exist specific limitations to mapping to the NCBI database; including continuous updates and risk for inclusion of contaminated contigs. An additional limitation and point of debate in the literature in tumor microbiota studies are choice of database; future studies could consider performing analyses on different iterations of the NCBI database or the MetaPhlAn database, for example. Conclusion Intratumoral reads mapping to Escherichia were associated with improved OS in univariate and multivariate analyses, in patients treated with single-agent ICI, in two cohorts. Escherichia appeared to localize within tumor cells and was associated with transcriptomic -28- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 changes suggesting an enhanced immune cell infiltration. The association between intratumoral bacteria, specifically Escherichia, and response to single-agent ICI merits further study. Supplemental Methods MSK-IMPACT Next-generation sequencing Briefly, DNA was extracted from FFPE tumors and patient-matched blood samples. Bar-coded libraries were generated and sequenced, targeting all exons and select introns of a custom gene panel of 468 (version 6), 410 (version 5), or 341 (version 3) genes. Mean sequencing coverage across all tumor samples was 744×, with minimum depth of coverage of 91×. Samples were run through a custom pipeline12to identify somatic alterations, including mutations and copy number alterations. To normalize somatic nonsynonymous TMB, the total number of mutations was divided by the coding region captured in each panel, which covered 0.98, 1.06, and 1.22 megabases (Mb) in the 341-, 410-, and 468-gene panels, respectively12. FoundationOne®CDx Next-generation sequencing Briefly, DNA was extracted FFPE tissue sections, followed by whole-genome shotgun library construction and hybridization-based capture for the targeted genes to a median exon coverage depth of >500x. Genomic alterations including base substitutions, short insertions and deletions, copy number alterations and rearrangements, were identified using an approach described previously10,34. Tumor mutational burden (TMB, in mutations / megabase, mut / Mb) was estimated from ~1.1Mb of the genome35, and samples with a TMB of at least 10 mut / Mb were classified as TMB-high. As with the MSK cohort, only patients with advanced (defined as either stage IV or stage III not amenable to definitive treatment) were included. All data were de-identified to protect patient confidentiality. Retrospective longitudinal clinical data were derived from electronic health records and comprised of both patient-level structured (e.g. medication orders and administrations) and unstructured (e.g. smoking history) data. Treatment information was -29- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 abstracted from oncologist-defined, rule-based lines of therapy. Clinical data were linked to genomic data derived from FoundationOne® CDx results by de-identified, deterministic matching. Patients without a date of death or progression were censored at date of last clinic visit or structured activity. To account for immortal time in the FMI cohort, results were left- truncated and patients were treated as at risk of death only after the later of their first FMI report date and their second visit in the FH network, as both are requirements for inclusion in the database. Mitigating sources of contamination upstream of pipeline analysis Samples at MSK were processed in a CLIA certified diagnostics laboratory utilizing a standardized protocol for formalin fixation and paraffin embedding. Patient samples were processed by a stringent protocol to preclude sample-to-sample and organism-related contamination: all procedures were carried out in an appropriately maintained biological safety cabinet / fume-hood. Working surfaces and tools used were cleaned and decontaminated with CaviCide decontaminant which is effective against all bacterial, viral, and fungal species. Only one sample was processed at a time, and new instruments were used for each type of tissue sampled and each patient to avoid carryover. For DNA extraction, the laboratory used sterilized and / or autoclaved reagents as appropriate. The MSK-IMPACT analysis pipeline utilizes a matched peripheral blood normal sample to identify patient specific germline polymorphisms. Samples were monitored for minor allele contamination in homozygous polymorphism sites which allows for high sensitivity for even low levels of contamination. Bacterial read detection and filtering for potential contaminants in the MSK discovery cohort Samtools (version 1.7) was used to extract unmapped reads from processed BAM files of MSK-IMPACT clinical samples into FASTA files. Unmapped reads from each sample were queried for microbial content using BLASTN (2.9.0+) against all human bacteria from the NCBI database (last accessed December 17th, 2023). From prior iterations of the pipeline, it was found that the Escherichia coli assembly LM996054 was mapping with high identity to reads -30- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 also mapping with the PhiX assembly. The LM996054 assembly was removed from the database. BLASTN was run on the updated database using the following settings: word_size 28 and perc_identity 90. KronaTools was then used to annotate corresponding NCBI Taxonomy to the aligned hits using the GI identifier from output. A custom R script utilizing the NCBI Taxonomy Database was used to group Taxonomic ID by genus and species, tallying the total reads per sample for each genus / species category. To identify putative contaminants, N=2,378 no-template controls (NTCs) were processed through the same microbial detection pipeline described above. Several organisms detected in NTCs or previously classified on the contamination “blacklist”
[0057] are potentially pathogenic or commensal17-34. These mixed-evidence genera (both a pathogen and common contaminant) were allowed into the dataset only if they were detected above the threshold of contamination (defined as highest read count detected for each of the individual genera) found in the NTC analysis. Bacterial read detection in the FMI validation cohort An approach similar to that used for the MSK discovery cohort was applied for bacterial read detection in the FMI validation cohort. Briefly, reads left unmapped to the hg19 reference genome were extracted using Samtools (version 1.7). Unmapped reads were then queried for microbial content using BLASTN (v2.6.0) against all bacterial species in the NCBI nucleotide database, followed by annotation of the corresponding NCBI taxonomy. A median Escherichia read count of 22 was observed in the overall cohort. Due to the absence of NTCs, the median read count was used to dichotomize the FMI cohort. Samples with a read count >22 were classified as Escherichia positive, while the remaining were considered Escherichia negative. Although there were no NTC samples available for the FMI cohort, and while samples originate from real-world cases from hospitals with different anatomic pathology laboratories, the FMI molecular laboratory where FFPE extraction and sequencing were performed utilized standard CLIA protocols to mitigate potential for contamination. -31- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Bulk RNAsequencing Approximately 200–500 ng of FFPE RNA extracted from FFPE slides with a DV200 range between 3-99 or 65–100 ng of fresh frozen RNA (DV20098-99) per sample were used for RNA library construction using the KAPA RNA Hyper library prep kit (Roche, Switzerland). The number of pre-capture PCR cycles was adjusted based on the quality and quantity of RNA extracted from the samples. Customized adapters with 3bp unique molecular indexes (UMI) (Integrated DNA Technologies, USA) and sample-specific dual-index primers (Integrated DNA Technologies, USA) were added to each library. The quantity of libraries was measured with Qubit (Thermo Fisher Scientific, USA) and the quality was assessed by TapeStation Genomic DNA Assay (Agilent Technologies, USA). Approximately 500 ng of each RNA library were pooled for hybridization capture with IDT Whole Exome Panel V1 (Integrated DNA Technologies, USA) using a customized capture protocol modified from NimbleGen SeqCap Target Enrichment system (Roche, Switzerland). The captured DNA libraries were then sequenced on an Illumina NovaSeq6000 in paired ends (2X100bp) to a target 50 million read pairs per sample. RNAsequencing analysis Reads were initially mapped to hg19 / GRCh37 using STAR, followed by deduplication from the combination of UMI sequence and alignment coordinate using umi-tools (v1.0.1) Deduplicated FASTQs were extracted from the deduplicated BAMs and used for pseudoalignment. Transcript expression was quantified from RNA-seq reads using kallisto (version 0.46.2) and hg19 Ensembl gene models. Lowly expressed genes were eliminated using the filterByExpr function from the edgeR package using a minimum of 20 counts and a total minimum count of 50. Differential gene expression analysis was performed with limma (Escherichia pos – Escherichia neg) using the voomWithQualityWeights function in order to downweight low quality samples. Batch was added to the limma model in order to account for batch effects. Escherichia status was used as a categorical variable (>10 reads defined as positive). GSEA analysis was performed using the fgsea R package using C5 gene ontology -32- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 sets, and Gene set variation analysis and pathway / cell type expression scoring was performed using the GSVA R package. For cell type analysis, cell type markers were obtained from PanglaoDB (citation below) and converted into gene sets. These gene sets were used for GSEA and GSVA analysis as described above. Image visualization was done by ggplot2(v3.5.0). Beta diversity analysis Beta diversity was determined by Bray-Curtis method using vegan package, generating unsupervised clustering defined by principle component analysis. The first three principal component generate points for each sample in three dimensional vector space. The relationship between two samples is defined by Euclidean distance between sample points in vector space. For each primary and metastatic sample, the Euclidean distance is calculated. For control, random pairings of samples were generated and the Euclidean distances are calculated in the same fashion. The primary-metastasis pairs are compared to the control of randomly selected pairs and Wilcox statistical test was applied. Fluorescence In Situ Hybridization Thirty-two samples included a range of Escherichia reads (0-166 reads of Escherichia detected by MSK-IMPACT), primary lung adenocarcinoma samples, and metastatic samples. A positive control colon adenocarcinoma sample containing sigmoid colon (13,237 Escherichia reads detected by MSK-IMPACT) was also included. FFPE tissue sections (5 μm) were maintained at 4°C. Samples were loaded into Leica Bond RX, baked for 30 mins. at 60°C, dewaxed with Bond Dewax Solution (Leica, AR9222), and pretreated with EDTA-based epitope retrieval ER2 solution (Leica, AR9640) for 15 mins. at 95°C (no proteolytic retrieval). An E. coli 16S rRNA-03probe (Advanced Cell Diagnostics (ACD), Cat# 312248) with sequence: GCCTTCGGGTTGTAAAGTACTTTCAGCGGGGAGGAAGGGAGTAAAGTTAATACCTTT GCTCATTGACGTTACCCGCAGAAGAAGC) (SEQ ID NO: 1) was hybridized for 2 hours at 42°C. Hybridized probes were detected using RNAscope 2.5 LS Reagent Kit – Brown (ACD, Cat# 322100) according to manufacturer’s instructions with some modifications (DAB -33- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 application was omitted and replaced with fluorescent CF594 / tyramide (Biotium, B40953) for 20 minutes at room temperature). Slides were visualized using CaseViewer version 2.4. Red / E.Coli channel was adjusted black 0, gamma 1.0, white 30. Green / background was adjusted black 0, gamma 1.0, white 30. Blue / DAPI was adjusted to black 0, gamma 1.0, white 60 to determine signal-to- noise / autofluorescence. Slides were scanned on a P250 Panoramic Scanner (3DHistech, Budapest, Hungary) using a 40x / 0.95NA objective. Regions of interest around the tissues were then drawn and exported as TIF files using Slide Viewer (3DHistech, Hungary.) Slides were then stained with hematoxylin and eosin (H&E). The coverslip was removed with xylene (Epredia, Cat. 99-905-01) and scanned again using an AT-2 Aperio scanner at 40X. Due to the xylene, some cellular material was lost. Slides were visualized with MSK Smart Slide Viewer. Representative images were selected for visualization. Regions of interest were co-registered manually by a pathologist (CMV). Images were exported to the Indica Labs Halo image analysis platform, followed by cell segmentation and signal thresholding that were performed separately on each image similar to as previously described35. Table 1. Baseline characteristics Characteristic Chemo / IO IO Escherichia- Escherichia- p- Escherichia- Escherichia p- neg, N = pos, N = value neg, N = - pos, N = valu 31212151 214712841e2Age 69 (61, 75) 66 (59, 74) 0.3 69 (63, 75) 68 (60, 74) 0.06 7 Sex 0.9 0.08 3 F 152 (49%) 103 (48%) 72 (49%) 164 (58%) M 160 (51%) 112 (52%) 75 (51%) 120 (42%) ECOG PS 0.5 0.3 >=2 33 (11%) 27 (13%) 13 (8.8%) 35 (12%) 0-1 260 (89%) 175 (87%) 134 (91%) 249 (88%) Unknown 19 13 Histology 0.14 0.00 4 -34- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Adenocarcino 239 (77%) 178 (83%) 94 (64%) 220 (77%) ma Other 35 (11%) 14 (6.5%) 24 (16%) 21 (7.4%) Squamous cell 38 (12%) 23 (11%) 29 (20%) 43 (15%) carcinoma Smoking 0.085 0.3 history Current / Forme 251 (81%) 185 (86%) 133 (90%) 248 (87%) r Never 60 (19%) 29 (14%) 14 (9.5%) 36 (13%) Unknown 1 1 Percent PDL1 0 (0, 5) 0 (0, 10) 0.15 20 (0, 80) 25 (0, 80) >0.9 Unknown 34 26 22 18 1 Median (IQR); n (%) 2 Wilcoxon rank sum test; Pearson’s Chi-squared test Table S1: Baseline characteristics Characteristic N = 1,4621Age category 67 (60, 74) Sex F 757 (52%) M 705 (48%) ECOG PS 0-1 1,283 (90%) >=2 146 (10%) Unknown 33 Smoking status Current / Former 1,227 (84%) Never 233 (16%) Unknown 2 Histology Adenocarcinoma 1,137 (78%) Squamous cell carcinoma 188 (13%) Other 137 (9.4%) Line of therapy 1 796 (54%) >=2 666 (46%) PD-L1 status <1 557 (51%) >=50 296 (27%) -35- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 1-49 247 (22%) Unknown 362 PD-L1% 0 (0, 50) Unknown 367 Treatment ICI monotherapy 920 (63%) Chemo-ICI 542 (37%) MSK-IMPACT Version IM6 984 (67%) IM5 386 (26%) IM3 92 (6.3%) 1 Median (IQR); n (%) Table S2, No Template Control by MSK-IMPACT Version MSK-IMPACT N = 2,3781Read count (median, IQR) Version Version 3 192 (8.1%) 42 (23,86) Version 5 574 (24%) 12 (4, 39) Version 6 1,612 (68%) 4 (2,10) 1 n (%) Table S3. Baseline characteristics for MSK-IMPACT version 6 Characteristic N = 9581Age category 68 (60, 75) Sex F 491 (51%) M 467 (49%) ECOG PS 0-1 818 (88%) >=2 108 (12%) Unknown 32 Smoking status Current / Former 817 (85%) Never 139 (15%) Unknown 2 Histology Adenocarcinoma 731 (76%) Squamous cell carcinoma 133 (14%) -36- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Other 94 (9.8%) Line of therapy 1 661 (69%) >=2 297 (31%) PD-L1 status <1 449 (52%) >=50 220 (25%) 1-49 194 (22%) Unknown 95 PD-L1% 0 (0, 50) Unknown 100 Treatment Chemo-ICI 527 (55%) ICI monotherapy 431 (45%) 1 Median (IQR); n (%) Table S4. Additional medication, demographic information, and post-progression treatment information MSK Cohort Characteristic Chemotherapy-ICI ICI monotherapy Escherichia- Escherichia- p- Escherichia- Escherichia p- neg, N = pos, N = value neg, N = - pos, N = valu 31212151 214712841e2Antibiotics <60 0.6 >0.9 days prior to ICI Antibiotics 93 (30%) 59 (27%) 35 (24%) 67 (24%) No Antibiotics 219 (70%) 156 (73%) 112 (76%) 217 (76%) Antibiotics 0.6 0.6 during ICI Antibiotics 73 (23%) 55 (26%) 35 (24%) 75 (26%) No Antibiotics 239 (77%) 160 (74%) 112 (76%) 209 (74%) Inhaled 0.3 0.7 corticosteroid use Inhaled 39 (13%) 20 (9.3%) 20 (14%) 42 (15%) corticosteroids No inhaled 273 (88%) 195 (91%) 127 (86%) 242 (85%) corticosteroids -37- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 PPI use 0.088 No PPI 244 (78%) 181 (84%) 128 (87%) 243 (86%) PPI 68 (22%) 34 (16%) 19 (13%) 41 (14%) COPD >0.9 0.7 COPD 124 (40%) 86 (40%) 61 (41%) 112 (39%) No COPD 188 (60%) 129 (60%) 86 (59%) 172 (61%) Chemotherapy 145 (46%) 109 (51%) 0.3 55 (37%) 108 (38%) >0.9 post- progression Immunotherap 86 (28%) 54 (25%) 0.5 42 (29%) 90 (32%) 0.5 y post- progression Experimental 62 (20%) 50 (23%) 0.4 22 (15%) 55 (19%) 0.3 therapy post- progression- 1 n (%) 2 Pearson’s Chi-squared test FMI Cohort Characteristic Chemotherapy-ICI ICI monotherapy Escherichia Escherichia p- Escherichia Escherichia p- -neg, N = -pos, N = value -neg, N = -pos, N = value 16411641 223012141 2COPD >0.9 0.6 COPD 35 (21%) 35 (21%) 57 (25%) 58 (27%) No COPD 129 (79%) 129 (79%) 173 (75%) 156 (73%) Chemotherapy 44 (66%) 43 (60%) 0.5 56 (52%) 53 (46%) 0.4 post- progression Immunotherap 7 (10%) 9 (13%) 0.8 32 (30%) 38 (33%) 0.6 y post- progression Experimental 16 (24%) 20 (28%) 0.7 20 (19%) 23 (20%) 0.9 therapy post- progression 1 n (%) 2 Pearson’s Chi-squared test -38- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Table S5. Baseline characteristics – FMI validation cohort Characteristic Chemotherapy-ICI ICI monotherapy Escherichi Escherichi p- Escherichi Escherichi p- a Negative, a Positive,2a Negative, a Positive, N = 164 N= 1641val21ue N= 2301N = 2141value Age 69 (62, 76) 68 (61, 73) 0.11 71 (62, 77) 67 (61, 75) 0.024 Sex 0.4 0.8 M 91 (55%) 84 (51%) 111 (48%) 106 (50%) F 73 (45%) 80 (49%) 119 (52%) 108 (50%) ECOG PS 0.5 0.023 >=2 33 (23%) 29 (20%) 50 (24%) 30 (15%) 0-1 110 (77%) 115 (80%) 155 (76%) 166 (85%) Unknown 21 20 25 18 Histology 0.5 0.8 Squamous cell 27 (16%) 35 (21%) 68 (30%) 61 (29%) carcinoma Other 35 (21%) 31 (19%) 36 (16%) 39 (18%) Adenocarcinom 102 (62%) 98 (60%) 126 (55%) 114 (53%) a Smoking history 0.11 0.7 History of 146 (89%) 154 (94%) 208 (90%) 196 (92%) smoking No history of 18 (11%) 10 (6.1%) 22 (9.6%) 18 (8.4%) smoking Percent PDL1 0.8 0.8 Negative 48 (38%) 53 (41%) 45 (26%) 44 (27%) Low Positive 44 (35%) 47 (36%) 57 (33%) 48 (29%) High Positive 33 (26%) 30 (23%) 70 (41%) 71 (44%) Unknown 39 34 58 51 1 Median (IQR); n (%) 2 Wilcoxon rank sum test; Pearson’s Chi-squared test References Hanahan, D. Hallmarks of Cancer: New Dimensions. Cancer Discovery 12, 31-46 (2022). https: / / doi.org / 10.1158 / 2159-8290.Cd-21-1059 -39- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Routy, B. et al. Gut microbiome influences efficacy of PD-1-based immunotherapy against epithelial tumors. Science 359, 91-97 (2018). https: / / doi.org / 10.1126 / science.aan3706 Derosa, L. et al. Intestinal Akkermansia muciniphila predicts clinical response to PD-1 blockade in patients with advanced non-small-cell lung cancer. Nature Medicine 28, 315- 324 (2022). https: / / doi.org / 10.1038 / s41591-021-01655-5 Liu, Y. et al. Lung tissue microbial profile in lung cancer is distinct from emphysema. Am J Cancer Res 8, 1775-1787 (2018). Laroumagne, S. et al. Bronchial colonisation in patients with lung cancer: a prospective study. European Respiratory Journal 42, 220-229 (2013). https: / / doi.org / 10.1183 / 09031936.00062212 Segal, L. N. et al. Enrichment of the lung microbiome with oral taxa is associated with lung inflammation of a Th17 phenotype. Nature Microbiology 1, 16031 (2016). https: / / doi.org / 10.1038 / nmicrobiol.2016.31 Segal, L. N. et al. Enrichment of lung microbiome with supraglottic taxa is associated with increased pulmonary inflammation. Microbiome 1, 19 (2013). https: / / doi.org / 10.1186 / 2049-2618-1-19 Tsay, J.-C. J. et al. Lower Airway Dysbiosis Affects Lung Cancer Progression. Cancer Discovery 11, 293-307 (2021). https: / / doi.org / 10.1158 / 2159-8290.Cd-20-0263 Nejman, D. et al. The human tumor microbiome is composed of tumor type-specific intracellular bacteria. Science 368, 973-980 (2020). https: / / doi.org / 10.1126 / science.aay9189 -40- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Heymann, C. J. F., Bard, J.-M., Heymann, M.-F., Heymann, D. & Bobin-Dubigeon, C. The intratumoral microbiome: Characterization methods and functional impact. Cancer Letters 522, 63-79 (2021). https: / / doi.org / https: / / doi.org / 10.1016 / j.canlet.2021.09.009 Le Noci, V. et al. Modulation of Pulmonary Microbiota by Antibiotic or Probiotic Aerosol Therapy: A Strategy to Promote Immunosurveillance against Lung Metastases. Cell Rep 24, 3528-3538 (2018). https: / / doi.org / 10.1016 / j.celrep.2018.08.090 Goubet, A. G. et al. Escherichia coli-Specific CXCL13-Producing TFH Are Associated with Clinical Efficacy of Neoadjuvant PD-1 Blockade against Muscle-Invasive Bladder Cancer. Cancer Discov 12, 2280-2307 (2022). https: / / doi.org / 10.1158 / 2159-8290.Cd-22- 0201 Cheng, D. T. et al. Memorial Sloan Kettering-Integrated Mutation Profiling of Actionable Cancer Targets (MSK-IMPACT): A Hybridization Capture-Based Next- Generation Sequencing Clinical Assay for Solid Tumor Molecular Oncology. J Mol Diagn 17, 251-264 (2015). https: / / doi.org / 10.1016 / j.jmoldx.2014.12.006 Milbury, C. A. et al. Clinical and analytical validation of FoundationOne®CDx, a comprehensive genomic profiling assay for solid tumors. PLoS One 17, e0264138 (2022). https: / / doi.org / 10.1371 / journal.pone.0264138 Singal, G. et al. Association of Patient Characteristics and Tumor Genomics With Clinical Outcomes Among Patients With Non-Small Cell Lung Cancer Using a Clinicogenomic Database. Jama 321, 1391-1399 (2019). https: / / doi.org / 10.1001 / jama.2019.3241 Cheng, D. T. et al. Memorial Sloan Kettering-Integrated Mutation Profiling of Actionable Cancer Targets (MSK-IMPACT): A Hybridization Capture-Based Next- Generation Sequencing Clinical Assay for Solid Tumor Molecular Oncology. The -41- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Journal of Molecular Diagnostics 17, 251-264 (2015). https: / / doi.org / https: / / doi.org / 10.1016 / j.jmoldx.2014.12.006 Elkrief, A. et al. Outcomes of single-agent PD-(L)-1 versus combination with chemotherapy in patients with PD-L1-high (≥ 50%) lung cancer. Journal of Clinical Oncology 40, 9052-9052 (2022). https: / / doi.org / 10.1200 / JCO.2022.40.16_suppl.9052 Elkrief, A. et al. Efficacy of PD-(L)1 blockade monotherapy compared with PD-(L)1 blockade plus chemotherapy in first-line PD-L1-positive advanced lung adenocarcinomas: a cohort study. J Immunother Cancer 11 (2023). https: / / doi.org / 10.1136 / jitc-2023-006994 Daniel D. Sjoberg, K. W., Michael Curry, Jessica A. Lavery and Joseph Larmarange GT Summary Package. The R Journal (2021) 13:1, pages 570-580 Salti, S. M. et al. Granzyme B regulates antiviral CD8+ T cell responses. J Immunol 187, 6301-6309 (2011). https: / / doi.org / 10.4049 / jimmunol.1100891 Wang, X. Q. et al. Spatial predictors of immunotherapy response in triple-negative breast cancer. Nature 621, 868-876 (2023). https: / / doi.org / 10.1038 / s41586-023-06498-3 Starner, T. D., Barker, C. K., Jia, H. P., Kang, Y. & McCray, P. B., Jr. CCL20 is an inducible product of human airway epithelia with innate immune properties. Am J Respir Cell Mol Biol 29, 627-633 (2003). https: / / doi.org / 10.1165 / rcmb.2002-0272OC Vilgelm, A. E. & Richmond, A. Chemokines Modulate Immune Surveillance in Tumorigenesis, Metastasis, and Response to Immunotherapy. Front Immunol 10, 333 (2019). https: / / doi.org / 10.3389 / fimmu.2019.00333 Keenan, T. E., Burke, K. P. & Van Allen, E. M. Genomic correlates of response to immune checkpoint blockade. Nat Med 25, 389-402 (2019). https: / / doi.org / 10.1038 / s41591-019-0382-x -42- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Xue, D. et al. A tumor-specific pro-IL-12 activates preexisting cytotoxic T cells to control established tumors. Sci Immunol 7, eabi6899 (2022). https: / / doi.org / 10.1126 / sciimmunol.abi6899 Dykema, A. G. et al. Lung tumor-infiltrating Treg have divergent transcriptional profiles and function linked to checkpoint blockade response. bioRxiv, 2022.2012.2013.520329 (2022). https: / / doi.org / 10.1101 / 2022.12.13.520329 Grigoriou, M. et al. Regulatory T-cell Transcriptomic Reprogramming Characterizes Adverse Events by Checkpoint Inhibitors in Solid Tumors. Cancer Immunol Res 9, 726- 734 (2021). https: / / doi.org / 10.1158 / 2326-6066.Cir-20-0969 Routh, E. D. et al. Transcriptomic Features of T Cell-Barren Tumors Are Conserved Across Diverse Tumor Types. Front Immunol 11, 57 (2020). https: / / doi.org / 10.3389 / fimmu.2020.00057 Zhou, X. & Sun, S. C. Targeting ubiquitin signaling for cancer immunotherapy. Signal Transduct Target Ther 6, 16 (2021). https: / / doi.org / 10.1038 / s41392-020-00421-2 Volpin, V. et al. CAMK1D Triggers Immune Resistance of Human Tumor Cells Refractory to Anti-PD-L1 Treatment. Cancer Immunol Res 8, 1163-1179 (2020). https: / / doi.org / 10.1158 / 2326-6066.Cir-19-0608 Kim, J. et al. TNIK Inhibition Has Dual Synergistic Effects on Tumor and Associated Immune Cells. Adv Biol (Weinh) 6, e2200030 (2022). https: / / doi.org / 10.1002 / adbi.202200030 Jiang, S. N. et al. Inhibition of tumor growth and metastasis by a combination of Escherichia coli-mediated cytolytic therapy and radiotherapy. Mol Ther 18, 635-642 (2010). https: / / doi.org / 10.1038 / mt.2009.295 -43- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Goubet, A.-G. et al. Escherichia coli-specific CXCL13-producing TFH are associated with clinical efficacy of neoadjuvant PD-1 blockade against muscle-invasive bladder cancer. Cancer Discovery (2022). https: / / doi.org / 10.1158 / 2159-8290.Cd-22-0201 Laroumagne, S. et al. Bronchial colonisation in patients with lung cancer: a prospective study. Eur Respir J 42, 220-229 (2013). https: / / doi.org / 10.1183 / 09031936.00062212 B. System and Methods for Selecting Cancer Therapies in Cancer Subjects Using Escherichia Reads Referring now to FIG. 15, depicted is a block diagram of a system 100 for selecting anti-cancer therapies for subjects based on Escherichia reads from biological samples. In overview, the system 100 may include at least one data processing system 105, at least one sample profiler 110, and at least one administrative device 115, communicatively coupled with one another via at least one network 120. The data processing system 105 may include at least one dataset retriever 125, at least one sequence analyzer 130, at least one attribution filter 135, at least one score generator 140, at least one response predictor 145, at least one output handler 150, and at least one database 155, among others. The system 100 may include any of the components and may be used to implement or perform any of the methods as detailed in Section A. Each of the components in the system 100, as detailed herein, may be implemented using hardware (e.g., one or more processors coupled with memory) or a combination of hardware and software, as detailed herein in Section C. In further overview, the data processing system 105 (sometimes herein generally referred to as a computing system or a server) may be any computing device including one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The data processing system 105 can be in communication with the sample profiler 110, the administrative device 115, the database 155, and other devices, via the network 120. The data processing system 105 may be situated, located, or otherwise associated with at least one server group. The server group may correspond to a data center, a -44- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 branch office, or a site at which one or more servers corresponding to the data processing system 105 are situated. The data processing system 105 may include one or more subsystems, modules, or components to execute the various processes and tasks detailed herein. On the data processing system 105, the data retriever 125 may obtain a dataset of reads generated from a sample from a subject. The sequence analyzer 130 may identify a candidate portion of reads corresponding to candidate Escherichia reads. The attribution filter 135 may filter reads corresponding to a bacteriophage from the candidate portion of reads. The score generator 140 may determine a score indicating a number of Escherichia reads based on the filtered reads. The response predictor 145 may select an anti-cancer therapy for the subject based on the score. The output handler 150 may generate an output for presentation via the administrative device 115. The sample profiler 110 may be any device to perform genetic sequencing on a gene segment (e.g., deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sample) in a sample taken from a subject and to generate sequencing datasets using the genetic sequencing. The genetic sequencing carried out may be a high throughput, massively parallel sequencing technique (sometimes herein referred to as next-generation sequencing), such as whole genome sequencing (WGS), pyrosequencing, reversible dye-terminator sequencing, SOLiD sequencing, Ion semiconductor sequencing, or Helioscope single molecule sequencing, among others. Using the genetic sequencing data, the sample profiler 110 may generate a genomic dataset profiling the subject. In generating the genomic dataset, the sample profiler 110 may execute read alignment, variant calling, and gene expression quantification, among others. In some embodiments, the sample profiler 110 may be part of a genomic profiling platform, such as IMPACT, GENIE, or FoundationOne platform, among others. The sample profiler 110 may use the gene sequencing to generate a sequencing dataset. The sequencing dataset may be maintained using one or more files according to a format (e.g., FASTQ, BAM, SAM, BCL, or VCF formats). -45- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 The administrative device 115 (sometimes herein referred to as a computing device) may be any computing device comprising one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The administrative device 115 may be in communication with the data processing system 105 and the sample profiler 110 via the network 120. The administrative device 115 may have at least one display. The administrative device 115 may be associated with an entity (e.g., a clinician) examining the subject or gene data from the subject. The display may present information about the subject provided by the data processing system 105. The database 155 may store and maintain various resources and data associated with the data processing system 105, the sample profiler 110, and the administrative device 115, among others. The database 155 may include a database management system (DBMS) to arrange and organize the data maintained thereon. The database 155 may be in communication with the data processing system 105, the sample profiler 110, and the administrative device 115, via the network 120. While running various operations, the data processing system 105, the sample profiler 110, and the administrative device 115 may access the database 155 to retrieve identified data therefrom. The data processing system 105, the sample profiler 110, and the administrative device 115 may also write data onto the database 155 from running such operations. Referring now to FIG. 16, depicted is a block diagram of a process 200 for parsing genome sequence data in the system 100 for selecting cancer therapies based on Escherichia reads. Under the process 200, the data processing system 105 and the sample profiler 110 may be used to assess at least one subject 205 with respect to cancer in the subject 205. The subject 205 (sometimes referred to herein as a cancer subject) may be at risk of or diagnosed with cancer. The cancer may include lung cancer. For example, the cancer may include non-small cell lung cancer (NSCLC), such as adenocarcinoma, squamous cell carcinoma, or large cell carcinoma, among others. The cancer (e.g., lung cancer) may be at any level of disease progression in terms of tumor growth. The lung cancer may be, for example, at least one of an early stage (e.g., Stage I, with cancer confined to the lung), a localized stage (e.g., Stage II, -46- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 with cancer spread to lymph nodes or adjacent regions within the lung), an advanced stage (e.g., Stage III, with cancer spread throughout the chest region), or a metastasized stage (e.g., Stage IV, with cancer spread throughout the body), among others. The subject 205 may be at any age range, such as above or equal to 18 years old (e.g., an adult) or less than 18 years old (e.g., a minor). At least one biological sample 210 may be taken, isolated, or otherwise obtained from the subject 205. The biological sample 210 may be used to assess the subject 205 for selection of anti-cancer therapies. The biological sample 210 may include, for example, blood, plasma, sera, urine, feces, epidermal sample, vaginal sample, skin sample, cheek swab, sperm, amniotic fluid, cultured cells, bone marrow sample, tumor biopsies, aspirate and / or chorionic villi, cultured cells, among others. The biological sample 210 may be obtained from a primary cancer site or a metastatic site associated with the cancer. For example, the primary cancer site may correspond to a lung region from which the lung cancer originated in the subject 205. The metastatic site may correspond to another anatomical site in the subject 205 to which the lung cancer spread from the primary cancer site. In some embodiments, the biological sample 210 may undergo initial screening for the presence or absence of Escherichia, prior to additional processing by the data processing system 105 and the sample profiler110. In some embodiments, the biological sample 210 may be detected as having or lacking Escherichia using a probe, such as the hybridized probe (e.g., SEQ ID NO: 1 as detailed in Section A). With the application of the probe, the biological sample 210 may be imaged (e.g., using a fluorescence imaging device) to scan for indications of Escherichia. When Escherichia is detected in the biological sample 210 in the screening, the biological sample 210 may undergo additional processing using the sample profiler 110. Otherwise, when Escherichia is not detected in the biological sample 210 in the screening, additional processing using the sample profiler 110 may be refrained from use for the biological sample 210. -47- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 The sample profiler 110 may execute, carry out, or otherwise perform genome sequencing on the biological sample 210. From performing genome sequencing, the sample profiler 110 may create, produce, or otherwise generate at least one dataset 215 (sometimes herein referred to as transcriptomic data). The dataset 215 may profile, define, or otherwise characterize the genetic makeup of the gene segments from the biological sample 210 of the subject 205. The dataset 215 may identify or include a set of genome sequence reads 220A–N (hereinafter generally referred to as genome sequence reads 220). Each genome sequence read 220 may correspond to a base pair position (or coordinate) within the gene segment in the biological sample 210. The set of genome sequence reads 220 may have any number of sequence reads, for example, ranging anywhere from 50 to 10,000,000 base pairs (bps). Each genome sequence read 220 may correspond to at least one respective position of the gene segment from the biological sample 210. In some embodiments, each genome sequence read 220 may have at least one nucleotide base, such as adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U). Each genome sequence read 220 may have at least one alphanumeric character to indicate a base type at a position in the gene segment. For instance, the identifier may have the alphanumeric character “A” for the nucleotide base type adenine, “C” for cytosine, “G” for guanine, “T” for thymine, and “U” for uracil, among others. In some embodiments, the dataset 215 may include multiple sets of genome sequence reads 220 corresponding to the multiple sequencing of the gene segments from the biological sample 210 of the subject 205. With the generation of the dataset 215, the sample profiler 110 may send, transmit, or otherwise provide the dataset 215 to the data processing system 105. In some embodiments, the sample profiler 110 may send the dataset 215 for storage on the database 155. The dataset retriever 125 may identify, receive, or otherwise obtain the dataset 215 including the set of genome sequence reads 220 from the sample profiler 110 or the database 155. In some embodiments, the data retriever 125 may obtain dataset 215, when the biological sample 210 is identified as having the candidate Escherichia reads via E. coli fluorescence in situ hybridization (FISH). In some embodiments, the FISH may be performed to detect E. coli nucleic acids using -48- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 a probe (e.g., of SEQ. ID. NO. 1). With the obtaining, the dataset retriever 125 may process or parse the dataset 215 to extract or identify the set of genome sequence reads 220 from the dataset 215. The sequence analyzer 130 may select, extract, or otherwise identify at least one candidate portion of genome sequence reads 220’A–N (hereinafter generally referred to as candidate portion of genome sequence reads 220’ or more generally as a first portion) from the set of genome sequence reads 220. The candidate portion of genome sequence reads 220’ can correspond to a subset of the initial set of genome sequence reads 220 that potentially have Escherichia reads. The candidate portion of the genome sequence reads 220’ may correspond to candidate Escherichia reads. The Escherichia reads may include, for example, genome sequence reads of E. coli found in the intratumoral microbiome of tumors associated with lung cancer. In some embodiments, the candidate portion of genome sequence reads 220’ may correspond to predefined genome sequence reads associated with presence of Escherichia in the biological sample 210. The associated, predefined genome sequence reads may include, for example, Granzyme B (GZMB), Chemokine (C-C motif) ligand 20 (CCL20), C-X-C chemokine receptor type 2 pseudogene 1 (CXCR2P1), Chemokine (C-X-C motif) ligand 13 (CXCL13), Interleukin 12 receptor, beta 2 subunit (IL12RB2), forkhead box P3 (FOXP3), or Serpin family E member 2 (SERPINE2), among others. The attribution filter 135 may remove or filter genome sequence reads from the candidate portion of genome sequence reads 220’ to generate at least one remnant portion of genome sequence reads 220”A–N (hereinafter generally referred to as remnant portion of genome sequence reads 220” or more generally as a second portion). The remnant portion of genome sequence reads 220” may correspond to true Escherichia reads and may include a subset of the candidate portion of sequence reads 220’ subsequent to filtering. At least some of the candidate portion of sequence reads 220’ may include sequence reads that are not true (or actual) Escherichia reads, but may be correlated with other types of reads, such as bacteriophage. In some embodiments, the attribution filter 135 may select or identify the sequence reads that are attributable to both the bacteriophage (e.g., ΦX such as PhiX174) and the true Escherichia reads -49- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 (e.g., associated with E. coli or predefined genome sequence reads). With the identification, the attribution filter 135 may exclude or remove the sequence reads that are attributable to both the bacteriophage and the true Escherichia reads from the candidate portion of genome sequence reads 220’ to generate the remnant portion of genome sequence reads 220”. Conversely, the attribution filter 135 may add or include the sequence reads that are not attributable to the bacteriophage and the true Escherichia reads from the candidate portion of genome sequence reads 220’ to the remnant portion of genome sequence reads 220”. In some embodiments, the attribution filter 135 may apply a filter to the candidate portion of genome sequence reads 220’ to generate the remnant portion of genome sequence reads 220”. The filter may include, for example, a filterByExpr as detailed herein. The filter may remove genome sequence reads based on expression levels (e.g., in terms of count, count-per-million (CPM), or other measures) of gene types in the candidate portion of genome sequence reads 220’. For each gene type, the attribution filter 135 may identify a respective expression level for the gene type in the candidate portion of genome sequence reads 220’. The attribution filter 135 may compare the expression level with a cutoff threshold defined by the filter. The threshold may define a value for the expression level at which to suppress or remove genome sequences of the gene type from the candidate portion of genome sequence reads 220’. If the expression level for the gene type satisfies (e.g., greater than or equal to) the cutoff threshold, the attribution filter 135 may keep or maintain the genome sequence reads associated with the gene type in the candidate portion of genome sequence reads 220’. Otherwise, if the expression level for the gene type does not satisfy (e.g., less than) the cutoff threshold, the attribution filter 135 may remove the genome sequence reads associated with the gene type from the candidate portion of genome sequence reads 220’. Referring now to FIG. 17, depicted is a block diagram of a process 300 for providing outputs based on classifications from genome sequence data in the system for selecting anti-cancer therapies based on Escherichia reads. Under the process 300, the score generator 140 may calculate, generate, or otherwise determine at least one score 305 using the remnant portion of genome sequence reads 220”. The score 305 may identify or indicate a number of true reads -50- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 for Escherichia. The score 305 may, for example, be a numerical value (e.g., an integer between 0 and 10,000) identifying a count of Escherichia reads in the biological sample 210. The score 305 may indicate a degree of presence of intratumoral Escherichia in the biological sample 210 taken from the subject 205. For instance, the score 305 may identify a level of presence of intratumoral Escherichia in the biological sample 210 taken from the subject 205. To determine, the score generator 140 may process or parse through the remnant portion of genome sequence reads 220”. In parsing, the score generator 140 may maintain a counter to keep track of the number of true reads for Escherichia (e.g., E. coli.). When there is a match between genome reads of the remnant portion of genome sequence reads 220” with true reads for Escherichia, the score generator 140 may update the counter. The response predictor 145 may obtain, identify, or determine at least one threshold 310 to compare the score 305 against. The threshold 310 may identify or define a value for the score 305 at which to determine a presence or absence of Escherichia in the subject 205 (or by extension the biological sample 210). In some embodiments, the response predictor 145 may calculate, generate, or otherwise determine the threshold 310 for a control group based on the number of true reads for Escherichia in subjects of the control group. The control group may include one or more non-cancer subjects (e.g., subjects lacking or not diagnosed with lung cancer). The threshold 310 may correspond to a combined value (e.g., mean, median, or weighted average) of numbers of true reads for Escherichia across the subjects in the control group. In some embodiments, the response predictor 145 may assign, set, or otherwise identify a fixed value for the number of true reads for Escherichia as the threshold 310. The fixed value may be assigned or configured by an administrator of the data processing system 105. With the identification of the threshold 310, the response predictor 145 may compare the score 305 with the threshold 310. Based on the comparison of the score 305 with the threshold 310, the response predictor 145 may identify or select at least one anti-cancer therapy from a set of candidate anti-cancer therapies 315A and 315B (hereinafter generally referred to as anti-cancer therapies 315). The set of candidate anti-cancer therapies 315 may include, for example, at least one of immune checkpoint inhibitor (ICI) therapy 315A or -51- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 chemotherapy 315B, among others. The ICI therapy 315A may include, for example, at least one of a PD-L1 antibody, a PD-1 antibody, or CTLA-4 treatment, among others. In some embodiments, the ICI therapy may include one or more of pembrolizumab, nivolumab, cemiplimab, atezolizumab, avelumab, durvalumab, ipilimumab, tremelimumab, ticlimumab, JTX-4014, Spartalizumab (PDR001), Camrelizumab (SHR1210), Sintilimab (IBI308), Tislelizumab (BGB-A317), Toripalimab (JS 001), Dostarlimab (TSR-042, WBP-285), INCMGA00012 (MGA012), AMP-224, AMP-514, KN035, CK-301, AUNP12, CA-170, or BMS-986189, among others. In some embodiments, the ICI therapy 315A may exclude CTLA-4 treatment and may include at least one of PD-L1 antibody or a PD-1 antibody. The chemotherapy 315B may include, for example, at least one of cisplatin or carboplatin, among others. In some embodiments, the response predictor 145 may classify the subject 205 (or the biological sample 210) into a first group corresponding to a presence of Escherichia (e.g., E. coli) or a second group corresponding to an absence of Escherichia (e.g., E. coli). Based on the classification, the response predictor 145 may select the anti-cancer therapy for the subject 205. If the score 305 is determined to satisfy (e.g., greater than or equal to) the threshold 310, the response predictor 145 may select the anti-cancer therapy 315 to include the ICI therapy 315A for the subject 205. The response predictor 145 may exclude chemotherapy 315B for administration for the subject 205. In some embodiments, the response predictor 145 may identify the subject 205 as a predicted responder to the ICI therapy 315A. In some embodiments, the response predictor 145 may identify the subject 205 as a candidate for the administration of the ICI therapy 315A. In some embodiments, the response predictor 145 may classify the subject 205 into the first group corresponding to the presence of Escherichia (e.g., E. coli), when the score 305 is determined to satisfy (e.g., greater than or equal to) the threshold 310. The response predictor 145 may select the anti-cancer therapy 315 to include ICI therapy 315A, based on the classification of the subject 205 into the first group. On the other hand, if the score 305 is determined to not satisfy (e.g., less than) the threshold 310, the response predictor 145 may select the anti-cancer therapy 315 to include a combination of the ICI therapy 315A and the chemotherapy 315B for administration for the -52- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 subject 205. In some embodiments, the response predictor 145 may select the anti-cancer therapy 315 to include the chemotherapy 315B or other therapies, such as surgical removal (e.g., lobectomy or resection), radiation therapy (e.g., external beam radiation therapy (EBRT), stereotactic body radiation therapy (SBRT), or internal beam radiation therapy (IBRT)), or targeted therapy (e.g., EGRF or ALK inhibitors), among others. In some embodiments, the response predictor 145 may identify the subject 205 as a predicted responder to a combination of the ICI therapy 315A and the chemotherapy 315B. In some embodiments, the response predictor 145 may identify the subject 205 as a candidate for the administration of the combination of the ICI therapy 315A and the chemotherapy 315B. In some embodiments, the response predictor 145 may classify the subject 205 into the second group corresponding to the absence of Escherichia (e.g., E. coli), when the score 305 is determined to not satisfy (e.g., less than) the threshold 310. The response predictor 145 may select the anti-cancer therapy 315 to include the combination of the ICI therapy 315A and the chemotherapy 315B, based on the classification of the subject 205 into the second group. The output handler 150 may create, produce, or otherwise generate at least one instruction 320. The instruction 320 may indicate or identify the anti-cancer therapy 315 selected for the subject 205. If the selected anti-cancer therapy 315 includes the ICI therapy 315A, the instruction 320 may identify the ICI therapy 315A as to be administered to the subject 205. In contrast, if the selected anti-cancer therapy 315 includes the combination of the ICI therapy 315A and the chemotherapy 315B, the instruction 320 may identify the combination of the ICI therapy 315A and the chemotherapy 315B as to be administered to the subject 205. If the selected anti-cancer therapy 315 includes the chemotherapy 315B or other therapies, the instruction 320 may identify the chemotherapy 315B or other therapies to be administered to the subject 205. In some embodiments, the instruction 320 may identify the subject 205 as a predicted responder to the ICI therapy 315A or a predicted responder to the combination of ICI therapy 315A and the chemotherapy 315B. In some embodiments, the instruction 320 may identify the subject 205 as a candidate for administration of the ICI therapy 315A or as a candidate for the administration of the combination of ICI therapy 315A and the chemotherapy -53- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 315B. In some embodiments, the instruction 320 may identify or include one or more of the following: the score 305, an identifier for the subject 205 (e.g., using an anonymized identifier), the classification of the subject 205 (or the biological sample 210), or information about the subject 205 (e.g., age, gender, or histology), among others. With the generation of the instruction 320, the output handler 150 may send, transmit, or otherwise provide the instruction 320 to the administrative device 115. The administrative device 115 may retrieve, obtain, or otherwise receive the instruction 320 from the data processing system 105. Upon receipt, the administrative device 115 may render, display, or otherwise present the instruction 320 via the display. For instance, the administrative device 115 may present the instruction 320 via a graphical user interface of an application. A user (e.g., a clinician examining the subject 205) of the administrative device 115 may use the information of the instruction 320 presented via the display to decide whether to administer the selected anti-cancer therapy 315. The subject 205 may be administered with a therapeutically effective amount of the selected anti-cancer therapy 315 (e.g., the ICI therapy 315A or the combination of the ICI therapy 315A and chemotherapy 315B) in accordance with the instruction 320. When the selected anti-cancer therapy 315 includes the ICI therapy 315A, the subject 205 may be administered (e.g., by the clinician) with a therapeutically effective amount of the ICI therapy 315A (e.g., a PD-L1 antibody, a PD-1 antibody, or CTLA-4 antibody). In some embodiments, when the subject 205 is identified as a candidate for the administration of the ICI therapy 315A, the subject 205 may be administered with a therapeutically effective amount of the ICI therapy 315A. The administration of the ICI therapy 315A may be performed exclusive of the administration of the chemotherapy 315B. The administration of the ICI therapy 315A to the subject may be altered, changed, or otherwise adjusted after the initial administration. For example, the amount (e.g., dosage or strength of the ICI therapy 315A) may be increased or decreased. In some embodiments, the ICI therapy 315A may be adjusted over a specific time period. The time period may include, for example, at least one of 1 month, 2 -54- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, or more, among others. When the selected anti-cancer therapy 315 includes the combination of the ICI therapy 315A and chemotherapy 315B, the subject 205 may be administered (e.g., by the clinician) with a therapeutically effective amount of the ICI therapy 315A (e.g., a PD-L1 antibody, a PD-1 antibody, or CTLA-4 antibody). In some embodiments, when the subject 205 is identified as a candidate for the administration of the combination of the ICI therapy 315A and chemotherapy 315B, the subject 205 may be administered with a therapeutically effective amount of the combination of the ICI therapy 315A and chemotherapy 315B. The administration of the combination of the ICI therapy 315A and chemotherapy 315B to the subject may be altered, changed, or otherwise adjusted after the initial administration. For example, the amount (e.g., dosage or strength of the combination of the ICI therapy 315A and chemotherapy 315B) may be increased or decreased. In some embodiments, the combination of the ICI therapy 315A and chemotherapy 315B may be adjusted over a specific time period. The time period may include, for example, at least one of 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, or more, among others. In some embodiments, when the selected anti-cancer therapy 315 includes the chemotherapy 315B or the other therapies, the subject 205 may be administered with the chemotherapy 315B or the other therapies as identified in the instruction 320. The administration of the anti-cancer therapy 315 (e.g., the ICI therapy 315A) to treat, alleviate, or address the cancer in the subject 205 may result in various positive outcomes. In some embodiments, the administration of the selected anti-cancer therapy 315, including the ICI therapy 315A to the subject 205 may result in an improved overall survival (OS), relative to an OS of a second cancer subject administered with ICI therapy 315A associated with a second score indicating a number of true reads for Escherichia not satisfying the threshold 310 (e.g., as depicted in FIG. 1A). In some embodiments, the administration of the selected anti-cancer therapy 315, including the ICI therapy 315A to the subject 205 may result in an improved progression-free survival (PFS), relative to a PFS of a second cancer subject administered with -55- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 ICI therapy 315A and associated with a second score indicating a number of true reads for Escherichia not satisfying the threshold 310 (e.g., as depicted in FIG. 8A). In addition, the administration of the anti-cancer therapy 315 (e.g., the combination of ICI therapy 315A and chemotherapy 315B) to treat, alleviate, or address the lung cancer in the subject 205 is anticipated or expected to result in various positive outcomes. In some embodiments, the administration of the selected anti-cancer therapy, including the combination of ICI therapy 315A and chemotherapy 315B to the subject may result in an improved overall survival (OS), relative to the OS of the subject prior to administration of the anti-cancer therapy. The administration of the selected anti-cancer therapy, including the combination of ICI therapy 315A and chemotherapy 315B to the subject results in an improved progression-free survival (PFS), relative to the PFS of the subject prior to administration of the anti-cancer therapy. (Larroquette et al., European Journal of Cancer, vol. 158, Nov. 2021, pp. 47–62.) In this manner, the number of Escherichia reads in filtered genomic data can be used to select therapies and provide significant improvements to clinical outcomes in subjects with cancers (e.g., lung cancer). The use of the filter (e.g., the attribution filter 135) can allow for precise discrimination of signals in the genomic sequence reads between Escherichia and bacteriophages (e.g., PhiX) to accurately characterize the tumor microbiota in the subject. By accurately determining the level of the presence of intratumoral Escherichia in biological samples from individuals, the data processing system 105 may generate and provide instructions for personalized and effective treatment of cancer. In particular, it is shown that the administration of monotherapy ICI for subjects with Escherichia results in significantly higher overall survival (OS) and progression free survival (PFS), relative to those without Escherichia (e.g., as seen in Figs 1A and 8A). The selection of anti-cancer therapies in this manner can also reduce the instances of delivery of anti-cancer therapies that would not be effective for such subjects. -56- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 From a computing perspective, the data processing system 105 may make use of the filter (e.g., the attribution filter 135) to precisely remove false genomic sequence reads that are attributable to reads of both true Escherichia and bacteriophages (e.g., PhiX). The filter may be highly selective (e.g., able to discriminate), precise (e.g., capably consistent at isolating true reads), and specific (e.g., accurately categorizing true reads as Escherichia and false reads as false) with respect to discrimination between true Escherichia and bacteriophage reads. With the filtering out of such genomic sequence reads, the data processing system 105 can conserve and reduce the consumption of computing resources, such as processors and memory, that would otherwise have been wasted when processing the genomic sequence data. The data processing system 105 can provide for a more reliable identification of intratumoral Escherichia from genomic sequence data. In addition, the outputs provided by the data processing system 105 may result in more accurate treatment administered to the subject, thereby lessening repeated uses of computing resources that would have been spent in providing less reliable and inaccurate outputs. Referring now to FIG. 18, depicted is a flow diagram of a method 400 of selecting anti-cancer therapies for subjects based on Escherichia reads from biological samples. The method 400 may be implemented using or performed by any of the components described herein, such as the system 100 or the system 500. Under the method 400, a computing system may receive a genomic dataset of a biological sample from a subject (405). The computing system may identify a candidate portion from the transcriptomic dataset corresponding to candidate Escherichia reads (410). The computing system may filter reads corresponding to bacteriophage from the candidate portion to generate a remaining portion (415). The computing system may determine a score indicating a number of true reads for Escherichia from the remaining portion (420). The computing system may identify a threshold for the number of reads of Escherichia (425). The computing system may determine whether the score satisfies (e.g., greater than or equal to) the threshold (430). If the score satisfies the threshold, the computing system may classify the subject as group corresponding to a presence of Escherichia (435). The -57- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 computing system may also select a monotherapy (e.g., immune checkpoint inhibitor (ICI) therapy) for the subject (440). On the other hand, if the score does not satisfy (e.g., less than) the threshold, the computing system may classify the subject as group corresponding to an absence of Escherichia (445). The computing system may also select a combination therapy (e.g., ICI therapy and chemotherapy) for the subject (450). The computing system may provide an output indicating the selected anti-cancer therapy (455). The subject may be administered with a therapeutically effective amount of the selected anti-cancer therapy (460). C. A Network and Computing Environment Various operations described herein can be implemented on computer systems. FIG. 19 shows a simplified block diagram of a representative server system 500, computing system 514, and network 526 usable to implement certain embodiments of the present disclosure. In various embodiments, server system 500 or similar systems can implement services or servers described herein or portions thereof. Computing system 514 or similar systems can implement clients described herein. The system 100 described herein can be similar to the server system 500. Server system 500 can have a modular design that incorporates a number of modules 502 (e.g., blades in a blade server embodiment); while two modules 502 are shown, any number can be provided. Each module 502 can include processing unit(s) 504 and local storage 506. Processing unit(s) 504 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 504 can include a general-purpose primary processor as well as one or more special-purpose co-processors such as graphics processors, digital signal processors, or the like. In some embodiments, some, or all processing units 504 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 504 can execute instructions stored in local storage 506. Any type of processors in any combination can be included in processing unit(s) 504. -58- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 Local storage 506 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic, or optical disk, flash memory, or the like). Storage media incorporated in local storage 506 can be fixed, removable, or upgradeable as desired. Local storage 506 can be physically or logically divided into various subunits such as a system memory, a read-only memory (ROM), and a permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 504 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 504. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module 502 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections. In some embodiments, local storage 506 can store one or more software programs to be executed by processing unit(s) 504, such as an operating system and / or programs implementing various server functions such as functions of the system 100 or any other system described herein, or any other server(s) associated with system 100 or any other system described herein. “Software” refers generally to sequences of instructions that, when executed by processing unit(s) 504, cause server system 500 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 504. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 506 (or non-local storage described below), processing unit(s) 504 -59- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 can retrieve program instructions to execute and data to process in order to execute various operations described above. In some server systems 500, multiple modules 502 can be interconnected via a bus or other interconnect 508, forming a local area network that supports communication between modules 502 and other components of server system 500. Interconnect 508 can be implemented using various technologies, including server racks, hubs, routers, etc. A wide area network (WAN) interface 510 can provide data communication capability between the local area network (e.g., through the interconnect 508) and the network 526, such as the Internet. Other technologies can be used to communicatively couple the server system 500 with the network 526, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.15 standards). In some embodiments, local storage 506 is intended to provide working memory for processing unit(s) 504, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 508. Storage for larger quantities of data can be provided on the local area network by one or more mass storage 512 that can be connected to interconnect 508. Mass storage 512 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage 512. In some embodiments, additional data storage resources may be accessible via WAN interface 510 (potentially with increased latency). Server system 500 can operate in response to requests received via WAN interface 510. For example, one of modules 502 can implement a supervisory function and assign discrete tasks to other modules 502 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 510. Such operation can generally be automated. Further, in some -60- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 embodiments, WAN interface 510 can connect multiple server systems 500 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation. Server system 500 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG. 15 as computing system 514. Computing system 514 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on. For example, computing system 514 can communicate via WAN interface 510. Computing system 514 can include computer components such as processing unit(s) 516, storage device 518, network interface 520, user input 522, and user output 524. Computing system 514 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like. Processing unit 516 and storage device 518 can be similar to processing unit(s) 504 and local storage 506 described above. Suitable devices can be selected based on the demands to be placed on computing system 514. For example, computing system 514 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Computing system 514 can be provisioned with program code executable by processing unit(s) 516 to enable various interactions with server system 500. Network interface 520 can provide a connection to the network 526, such as a wide area network (e.g., the Internet) to which WAN interface 510 of server system 500 is also connected. In various embodiments, network interface 520 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, 5G, LTE, etc.). -61- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 User input device 522 can include any device (or devices) via which a user can provide signals to computing system 514; computing system 514 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 522 can include any or all of a keyboard, touch pad, touch screen, mouse, or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on. User output device 524 can include any device via which computing system 514 can provide information to a user. For example, user output device 524 can include display-to- display images generated by or delivered to computing system 514. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light-emitting diode (LED) display including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog-to- digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both input and output device. In some embodiments, other user output devices 524 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on. Some embodiments include electronic components, such as microprocessors, storage, and memory that store computer program instructions in a computer readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer readable storage medium. When one or more processing units execute these program instructions, they cause the processing unit(s) to perform various operations indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 504 and 516 can provide various functionality for server system 500 and computing system 514, including any of the functionality described herein as being performed by a server or client, or other functionality. -62- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 It will be appreciated that server system 500 and client computing system 514 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 500 and client computing system 514 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be, but need not be, located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations (e.g., by programming a processor or providing appropriate control circuitry) and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software. While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies, including but not limited to specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished (e.g., by designing electronic circuits to perform the operation) by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be -63- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa. Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, and other non-transitory media. Computer readable media encoded with the program code may be packaged with a compatible electronic device, or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium). Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims. EQUIVALENTS The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. -64- 4908-5251-1299.1 Atty. Dkt. No.: 115872-3221 In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group. As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 cells refers to groups having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to groups having 1, 2, 3, 4, or 5 cells, and so forth. All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification. -65- 4908-5251-1299.1
Claims
Atty. Dkt. No.: 115872-3221 WHAT IS CLAIMED IS:
1. A method of selecting an anti-cancer therapy in a cancer subject, comprising: obtaining, by one or more processors, for a subject at risk of or diagnosed with cancer, a dataset comprising genome sequence reads of a biological sample obtained from the cancer subject; identifying, by the one or more processors, from the dataset identifying the genome sequence reads, a first portion of the genome sequence reads corresponding to candidate Escherichia reads; filtering, by the one or more processors, genome sequence reads corresponding to a bacteriophage from the first portion of the genome sequence reads to generate a second portion of the genome sequence reads corresponding to true Escherichia reads; determining, by the one or more processors, using the second portion of the genome sequence reads, a score indicating a number of true reads for Escherichia; selecting, by the one or more processors, from a plurality of anti-cancer therapies, the anti-cancer therapy for the cancer subject based on the score relative to a threshold; and storing, by the one or more processors, using one or more data structures, an association between the cancer subject and the anti-cancer therapy.
2. The method of claim 1, further comprising determining, by the one or more processors, a threshold for a control group based on a plurality of number of reads of Escherichia from a corresponding plurality of non-cancer subjects.
3. The method of claims 1 or 2, further comprising identifying, by the one or more processors, a fixed value for the number of true reads for Escherichia as the threshold.
4. The method of any one or more of claims 1–3, further comprising determining, by the one or more processors, that the score satisfies the threshold, and -66- 4908-5251-1299.1Atty. Dkt. No.: 115872-3221 wherein selecting the anti-cancer therapy further comprises selecting the anti-cancer therapy including immune checkpoint inhibitor (ICI) therapy, responsive to determining that the score satisfies the threshold.
5. The method of any one or more of claims 1–3, further comprising determining, by the one or more processors, that the score does not satisfy the threshold, and wherein selecting the anti-cancer therapy further comprises selecting the anti-cancer therapy including a combination therapy of the ICI therapy and chemotherapy, responsive to determining that the score does not satisfy the threshold.
6. The method of any one or more of claims 1–3, further comprising classifying, by the one or more processors, the cancer subject into one of (i) a first group corresponding to a presence of intratumoral Escherichia or (ii) a second corresponding to an absence of the intratumoral Escherichia, based on comparing the number of reads with the threshold, and wherein selecting the anti-cancer therapy further comprises selecting the anti-cancer therapy based on classifying the cancer subject into one of the first group or the second group.
7. The method of any one or more of the preceding claims, wherein the plurality of anti-cancer therapies comprises at least one of (i) an ICI therapy or (ii) a combination therapy of the ICI therapy and chemotherapy, optionally wherein the ICI therapy comprises at least one of a PD-L1 antibody, a PD-1 antibody, or a CTLA-4 antibody.
8. The method of claim 7, wherein the ICI therapy comprises one or more of pembrolizumab, nivolumab, cemiplimab, atezolizumab, avelumab, durvalumab, ipilimumab, tremelimumab, ticlimumab, JTX-4014, Spartalizumab (PDR001), Camrelizumab (SHR1210), Sintilimab (IBI308), Tislelizumab (BGB-A317), Toripalimab (JS 001), Dostarlimab (TSR-042, WBP-285), INCMGA00012 (MGA012), AMP-224, AMP-514, KN035, CK-301, AUNP12, CA-170, or BMS-986189. -67- 4908-5251-1299.1Atty. Dkt. No.: 115872-3221 9. The method of any one or more of claims 4–8, further comprising administering the cancer subject with a therapeutically effective amount of the anti-cancer therapy selected from the plurality of anti-cancer therapies for the cancer subject.
10. The method of claim 4, wherein administration of the selected cancer therapy including the ICI therapy to the cancer subject results in an improved overall survival (OS), relative to an OS of a second subject administered with ICI therapy associated with a second score indicating a number of true reads for Escherichia not satisfying the threshold.
11. The method of one or more of claims 4 or 10, wherein administration of the selected cancer therapy including the ICI therapy to the cancer subject results in an improved progression-free survival (PFS), relative to a PFS of a second cancer subject administered with ICI therapy and associated with a second score indicating a number of true reads for Escherichia not satisfying the threshold.
12. The method of any one or more of the preceding claims, wherein filtering the sequence reads further comprises identifying, from the first portion of the genome sequence reads, the sequence reads attributable to the bacteriophage and the true Escherichia reads, wherein the bacteriophage comprises ΦX, and wherein the true Escherichia reads are associated with E. coli.
13. The method of any one or more of the preceding claims, wherein obtaining the dataset further comprises obtaining the dataset comprising the genome sequence reads of the biological sample identified as having the candidate Escherichia reads via E. coli fluorescence in situ hybridization (FISH).
14. The method of claim 13, wherein FISH comprises detecting E. coli nucleic acids using a probe of SEQ ID NO:
1. -68- 4908-5251-1299.1Atty. Dkt. No.: 115872-3221 15. The method of any one or more of the preceding claims, wherein the cancer comprises lung cancer, and wherein the lung cancer optionally comprises non-small cell lung cancer (NSCLC), and is one of an early stage, a localized stage, an advanced stage, or a metastasized stage in the cancer subject.
16. The method of any one or more of the preceding claims, wherein the biological sample is obtained from at least one of a primary cancer site or a metastatic cancer site from the cancer subject.
17. A system for selecting a cancer therapy in a cancer subject, comprising: one or more processors coupled with memory, configured to: obtain, for a subject at risk of or diagnosed with cancer, a dataset comprising genome sequence reads of a biological sample obtained from the cancer subject; identify, from the dataset identifying the genome sequence reads, a first portion of the genome sequence reads corresponding to candidate Escherichia reads; filter sequence reads corresponding to a bacteriophage from the first portion of the genome sequence reads to generate a second portion of the genome sequence reads corresponding to true Escherichia reads; determine, using the second portion of the genome sequence reads, a score indicating a number of true reads for Escherichia; select, from a plurality of anti-cancer therapies, the anti-cancer therapy for the cancer subject based on the score relative to a threshold; and store, using one or more data structures, an association between the cancer subject and the anti-cancer therapy.
18. The system of claim 17, wherein the one or more processors are configured to determine a threshold for a control group based on a plurality of number of reads of Escherichia from a corresponding plurality of non-cancer subjects. -69- 4908-5251-1299.1Atty. Dkt. No.: 115872-3221 19. The system of any one or more of claims 17 or 18, wherein the one or more processors are configured to identify a fixed value for the number of true reads for Escherichia as the threshold.
20. The system of any one or more of claims 17–19, wherein the one or more processors are configured to: determine that the score satisfies the threshold, and select the anti-cancer therapy including immune checkpoint inhibitor (ICI) therapy, responsive to determining that the score satisfies the threshold.
21. The system of any one or more of claims 17–20, wherein the one or more processors are configured to: determine that the score does not satisfy the threshold, and select the anti-cancer therapy including a combination therapy of the ICI therapy and chemotherapy, responsive to determining that the score does not satisfy the threshold.
22. The system of any one or more of claims 17–21, wherein the one or more processors are configured to: classify the cancer subject into one of (i) a first group corresponding to a presence of intratumoral Escherichia or (ii) a second corresponding to an absence of the intratumoral Escherichia, based on comparing the number of reads with the threshold, and select the anti-cancer therapy based on classifying the cancer subject into one of the first group or the second group.
23. The system of any one or more of claims 17–22, wherein the plurality of anti-cancer therapies comprises at least one of (i) an ICI therapy or (ii) a combination therapy of the ICI therapy and chemotherapy, wherein the ICI therapy comprises at least one of a PD-L1 antibody or CTLA-4 treatment. -70- 4908-5251-1299.1Atty. Dkt. No.: 115872-3221 24. The system of claim 23, wherein the ICI therapy comprises one or more of pembrolizumab, nivolumab, cemiplimab, atezolizumab, avelumab, durvalumab, ipilimumab, tremelimumab, ticlimumab, JTX-4014, Spartalizumab (PDR001), Camrelizumab (SHR1210), Sintilimab (IBI308), Tislelizumab (BGB-A317), Toripalimab (JS 001), Dostarlimab (TSR-042, WBP-285), INCMGA00012 (MGA012), AMP-224, AMP-514, KN035, CK-301, AUNP12, CA-170, or BMS-986189.
25. The system of any one or more of claims 20–23, wherein the cancer subject is administered with a therapeutically effective amount of the anti-cancer therapy selected from the plurality of anti-cancer therapies for the cancer subject.
26. The system of claim 20, wherein administration of the selected cancer therapy including the ICI therapy to the cancer subject results in an improved overall survival (OS), relative to an OS of a second cancer subject administered with ICI therapy associated with a second score indicating a number of true reads for Escherichia not satisfying the threshold.
27. The system of claims 20 or 26, wherein administration of the selected cancer therapy including the ICI therapy to the cancer subject results in an improved progression-free survival (PFS), relative to a PFS of a second cancer subject administered with ICI therapy and associated with a second score indicating a number of true reads for Escherichia not satisfying the threshold.
28. The system of any one or more of claims 17–27, wherein the one or more processors are configured to identify, from the first portion of the genome sequence reads, the sequence reads attributable to the bacteriophage and the true Escherichia reads, wherein the bacteriophage comprises ΦX, and wherein the true Escherichia reads are associated with E. coli.
29. The system of any one or more of claims 17–28, wherein the one or more processors are configured to obtaining the dataset comprising the genome sequence reads of the biological -71- 4908-5251-1299.1Atty. Dkt. No.: 115872-3221 sample identified as having the candidate Escherichia reads via E. coli fluorescence in situ hybridization (FISH).
30. The system of claim 29, wherein FISH comprises detecting E. coli nucleic acids using a probe of SEQ ID NO:
1.
31. The system of any one or more of claims 17–30, wherein the cancer comprises lung cancer, and wherein the lung cancer optionally comprises non-small cell lung cancer (NSCLC) and is one of an early stage, a localized stage, an advanced stage, or a metastasized stage in the cancer subject.
32. The system of any one or more of claims 17–31, wherein the biological sample is obtained from at least one of a primary cancer site or a metastatic cancer site from the cancer subject. -72- 4908-5251-1299.1
Citation Information
Patent Citations
Intestinal flora marker for judging intestinal cancer and detection method thereof
CN112725457A
Systems and methods for identifying morphological patterns in tissue samples
US20230081613A1