Methods and systems for whole-spectrum analysis of disseminated and circulating tumor cells
A novel method using tetrameric antibodies and scRNA-seq for DTCs/CTCs enrichment and analysis addresses sensitivity and accuracy issues, enabling precise detection and characterization of DTCs/CTCs across cancer types, enhancing early cancer diagnosis and drug target discovery.
Patent Information
- Application Number
- PCT/US2025/034492
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-20
- Filing Date
- 2025-06-20
- Publication Date
- 2025-12-26
AI Technical Summary
Current methods for detecting and analyzing disseminated and circulating tumor cells (DTCs/CTCs) in liquid biopsies suffer from low sensitivity, high false-positive rates, and inability to accurately determine tissue-of-origin and tumor type due to phenotypic diversity and rarity of these cells, hindering their clinical and diagnostic applications.
A label-free method using tetrameric antibodies and density gradient centrifugation to enrich DTCs/CTCs, followed by single-cell RNA sequencing (scRNA-seq) to identify genome-wide copy number alterations (CNAs) and apply recursive logistic regression for tumor type tracing, enabling accurate identification and characterization of DTCs/CTCs across various cancer types.
The method achieves precise detection and unbiased molecular profiling of DTCs/CTCs, enhancing early cancer diagnosis and drug target discovery with robust performance in low-depth sequencing, reducing costs and improving clinical applicability.
Smart Images

Figure US2025034492_26122025_PF_FP_ABST
Abstract
Description
[0001]METHODS AND SYSTEMS FOR WHOLE-SPECTRUM ANALYSIS OF DISSEMINATED AND CIRCULATING TUMOR CELLS CROSS REFERENCE TO RELATED APPLICATION This application claims priority to U.S. Provisional Application No.63 / 662,045, filed June 20, 2024, which is incorporated by reference in its entirety. FIELD This disclosure relates to methods of enriching or identifying circulating tumor cells or disseminated tumor cells in a sample and methods of determining tissue of origin or tumor type of the cells. ACKNOWLEDGMENT OF GOVERNMENT SUPPORT This invention was made with government support under R33CA256112 and R01CA296872 awarded by the National Cancer Institute of the National Institutes of Health. The government has certain rights in the invention. BACKGROUND Tumor-derived materials in liquid biopsies, such as disseminated and circulating tumor cells (DTCs / CTCs), circulating tumor DNAs (ctDNAs), and tumor-derived exosomes (TEXs), are garnering increasing recognition for their transformative potential in cancer diagnostics, therapeutic monitoring, and prognosis prediction. CTCs and DTCs, as metastatic precursor cells, detach from primary tumors and disseminate in blood and other body fluids, respectively. Unlike ctDNAs and TEXs, these intact tumor cells not only encapsulate the entire spectrum of molecular information but also have the capability to provide crucial phenotypic and functional insights. This positions DTCs / CTCs as valuable surrogates for primary tumor lesions and as crucial windows into understanding tumor metastasis. However, the detection of DTCs / CTCs is fraught with challenges. Current methods, typically reliant on epithelial markers like epithelial cell adhesion molecule (EpCAM) and cytokeratin (CK), or morphological characteristics, suffer from low sensitivity and high false-positive rates due to the phenotypic diversity of DTCs / CTCs and the presence of non-tumor cells with epithelial traits in liquid biopsies. Furthermore, traditional approaches fall short in tracing the Tissue-of-Origin and Tumor Types (TOTT) and obtaining accurate, unbiased single-cell transcriptome data of DTCs / CTCs in a high-throughput fashion, crucial for liquid biopsy- based cancer diagnostics and in-depth molecular investigations of these metastatic cells. The research and clinical applications of DTCs / CTCs have long been impeded by their extreme rarity and the lack of an accurate, generic method for their unbiased detection and analysis across different cancer types through a shared oncogenic trait. Such limitations hinder the full exploitation of the vast information potential inherent in DTCs / CTCs, arguably the most informative tumor-derived material obtainable in liquid biopsies. Cancer fundamentally arises from genomic instability. As a salient feature, somatic copy number alternations (CNAs) are nearly ubiquitous in solid tumors but occur sporadically in normal cells, positioning it as a robust, generic marker for identifying rare DTCs / CTCs in complex normal cell populations. Nevertheless, the extreme rarity of DTCs / CTCs precludes applying single-cell sequencing directly to these samples for genomic elucidation. Emerging computational endeavors leverage machine-learning algorithms to analyze single-cell RNA sequencing (scRNA-seq) data for DTCs / CTCs detection and TOTT tracing. However, these methods have mostly been validated on synthetic datasets, usually abundant in DTCs / CTCs (>100), derived from prior studies and remain untested against real clinical specimens, which are characteristically scarce in DTCs / CTCs. Moreover, some of these algorithms necessitate prior training on specific cancer types, thus circumscribing their applicability in heterogeneous clinical landscapes. SUMMARY Disclosed herein are label-free methods leveraging genome-wide CNAs for unbiased detection, comprehensive molecular profiling, and TOTT tracing of DTCs / CTCs in diverse liquid biopsy specimens across various cancer types. In some aspects, methods of enriching DTCs or CTCs from a biological sample are provided. In some examples, the methods include contacting the sample with a plurality of tetrameric antibodies, wherein each tetrameric antibody comprises a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by a selected cell type; applying the sample contacted with the plurality of tetrameric antibodies to a first density gradient medium and performing a first density gradient centrifugation to obtain a first population of enriched DTCs or CTCs; resuspending the remaining cells following the first density gradient centrifugation, applying the resuspended cells to a second density gradient medium, and performing a second density gradient centrifugation to obtain a second population of enriched DTCs or CTCs; and combining the first and second populations of enriched DTCs or CTCs, thereby enriching the DTCs or CTCs. In other aspects, methods for identifying DTCs or CTCs in a biological sample are provided. In some examples, the methods include enriching DTCs or CTCs from the biological sample; performing single cell RNA-sequencing (scRNA-seq) on the enriched DTCs or CTCs to obtain scRNA-seq data; and identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells, thereby identifying DTCs or CTCs in the biological sample. In other examples, the methods include receiving a single cell RNA sequencing (scRNA-seq) dataset for the sample; and identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells, thereby identifying DTCs or CTCs in the biological sample. Methods of identifying a tumor type or tissue of origin of a tumor in a biological sample from a subject are also provided. In some examples, the methods include identifying DTCs or CTCs in the biological sample (for example utilizing the methods disclosed herein); constructing signature gene sets for a plurality of solid tumor types for tumor type-specific score calculations; determining tumor type-specific gene signature scores of the DTCs or CTCs from the scRNA-seq data obtained from the sample; and identifying the tumor type or tissue of origin based on recursive logistic regression model prediction using tumor type specific gene signature scores as input. In other examples, the methods include receiving a single cell RNA sequencing (scRNA-seq) dataset from a sample; determining a gene signature score of the DTCs or CTCs from the scRNA-seq data obtained from the sample; comparing the gene signature score of the DTCs or CTCs in the sample with one or more tumor type-specific signature scores; and identifying the tumor type or tissue of origin based on the gene signature score. Systems and computer-readable media including computer-executable instructions for performing a process for identifying DTCs or CTCs or identifying a tumor type or tissue of origin of a tumor are also provided. In other aspects, methods of identifying a druggable cancer target and a cancer drug against the target are provided. In some examples, the methods include receiving RNA-seq data of DTCs or CTCs; determining metastatic transcriptional states in the RNA-seq data of the DTCs or CTCs encoding a drug target; and identifying a drug targeting the drug target. The foregoing and other features of this disclosure will become more apparent from the following detailed description of several aspects, which proceeds with reference to the accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS FIGS.1A-1G are an overview of DTCFinder and core data modalities. FIG.1A: Schematic representation of DTCFinder. FIG.1B: A flowchart illustrating the main steps of DTCFinder. FIG.1C: Contrasting CNA burdens between normal cells or tissues and eight representative tumor types from TCGA. NAT: Normal Adjacent to Tumor. FIG.1D: Histogram showcasing accumulated outlier bins (AOBs) for baseline immune cells (blue) alongside a Gaussian fitting and spike-in lung cancer cells (red) as discerned by DTCFinder. The dashed line demarcates the AOB cut-off p value of 10-8. FIG.1E: ROC curve with sensitivity, specificity, precision (PPV), and AUC values for DTC / CTC identification in the cell line spike-in sample, utilizing the cut-off p value of 10-8(denoted by the red cross in the ROC curve). FIG.1F: Comparative analysis of genome-wide CNAs of spike-in H1975 cells as assessed by LC-WGS against those inferred by DTCFinder and InferCNV. Genomic segments manifesting consistent CNAs are encircled by blue dashed boxes. CNAs at specific genomic loci, accurately inferred by DTCFinder but overlooked by InferCNV, are marked by yellow stars. FIG.1G: Spearman correlations between CNAs measured by WGS and those inferred by DTCFinder across three spike-in lines are represented. The correlation coefficients are symbolized by the circle sizes, and p-values are conveyed through color coding. FIGS.2A-2J show exemplary analyses of clinical liquid biopsy samples by DTCFinder. FIG.2A: Alignment of genome-wide CNAs of CTCs inferred by DTCFinder with those of the paired primary tumor tissue experimentally measured by WGS for patient DTC-S1. Genomic segments with consistent CNAs are outlined by blue dashed boxes. FIG. 2B: A heatmap displaying pairwise Spearman’s correlations of inferred CNAs across all identified DTCs / CTCs in patient DTC-S1 (CC: correlation coefficient). FIG.2C: Proportions of DTCs / CTCs identifiable (CellSearch+) or unidentifiable (CellSearch-) by markers (CD45- / EPCAM+ / PanCK+) used in the CellSearch system, across the nine lung cancer patients studied (PB: peripheral blood, MPE: malignant pleural effusion, CSF: cerebrospinal fluid). FIG.2D: Top and bottom 30 genes showing differential expression in DTCs / CTCs from LUAD patients DTC-N1 and DTC-N2, contrasted with primary tumor cells from 16 LUAD patients and 4 normal lung samples. Averaged gene expression levels of each sample are displayed based on the normalized z-scores. FIG.2E: Representative differentially enriched transcriptome signatures between DTCs / CTCs from LUAD patients DTC-N1 and DTC-N2 and primary tumor cells from 16 LUAD patients are color-coded and illustrated with normalized z-scores. For illustrative clarity, the data bars for DTCs / CTCs in the heatmap are enlarged five-fold. FIGS.2F and 2G: Comparative analysis of average expression frequencies across all the cancer-testis antigens (CTA) curated in CTDatabase (FIG.2F) alongside expression levels and frequencies of representative CTAs in DTCs / CTCs from two NSCLC patients (DTC-N1 and DTC-N2), primary tumor cells from 16 LUAD patients, and epithelial cells from 4 normal lung tissues (FIG.2G). In FIG.2F, mean ± SEM is represented by dots and error bars, with statistical significance assessed via Brown-Forsythe and Welch ANOVA tests, indicated by red (DTC-N1) and orange (DTC N2) stars for pairwise comparisons, corrected for multiple comparisons using the Benjamini, Krieger, and Yekutieli method (*p<0.05, **p<0.005, ***p<0.0005, ns: not significant). In FIG.2G, normalized z- scores show averaged CTA expression across single cells within each sample, with circle sizes indicating frequencies (the percentage of cells expressing the CTA). FIG.2H: Predictions of the tissue-of-origin and tumor type (TOTT) for DTCs / CTCs from patients with different tumor types are color coded as red for positive prediction and blue for negative prediction. The tumor types listed in parentheses next to the sample IDs indicate the diagnosed tumor types. Stars on the sample IDs highlight the samples with 3 or fewer DTCs / CTCs identified. FIG.2I: Prediction of the TOTT for different forms of DTCs / CTCs (color coded by blue, green, and red respectively) identified from a breast cancer patient in a prior study. FIG.2J: The sensitivity and precision of DTC identification using downsampled, low-depth scRNA-seq data from 3 lung cancer patients are depicted across 3 independent downsampling trials (mean ± SD), with accompanying TOTT predictions shown above. FIG.3 shows evaluation of CNA burdens across varied stages for four exemplary tumor types. The depicted CNA burdens for distinct tumor samples from TCGA, classified according to their respective stages, are subjected to pairwise comparisons employing the Kruskal-Wallis test, supplemented by Dunn's test for corrections pertaining to multiple comparisons. The red dashed line indicates the average CNA burden of normal immune cells (*p<0.05, **p<0.005, ***p<0.0005, ns: not significant). FIGS.4A-4H show assessment of DTCFinder performance by the cell line spike-in sample. FIGS.2A-2C: Histograms of accumulated outlier bins (AOBs) for baseline immune cells, depicted with Gaussian fittings, juxtaposed with spike-in H1975 (FIG.4A), H1650 (FIG.4B), and A549 (FIG.4C) cells as identified by DTCFinder. Dashed lines represent the AOB cut-off p-value of 10-8. Ratios of identified to actual present tumor cells are displayed above each histogram for individual cell lines. FIG.4D: Histogram illustrating AOBs for baseline immune cells (blue) and spike-in A549 cells (red) not detected by DTCFinder, highlighting proximity of AOBs of missed A549 cells to the DTC / CTC identification cut-off threshold. FIG.4E: Normalized z-scores represent expression levels of cell line-specific markers across spike-in A549, H1650, H1975 cells, and randomly selected immune cells, with color coding for cell line-specific gene markers: green for A549, red for H1650, and blue for H1975. FIG.4F: DTCFinder’ s performance at varying bin sizes is illustrated by the balanced accuracy and false discovery rate (FDR) of DTC identification, establishing 150 genes / bin as the default setting due to near-perfect balanced accuracy and minimal FDR. FIGS.4G and 4H: Comparative illustrations of genome-wide CNAs as measured by LC- WGS and inferred by DTCFinder and InferCNV for H1650 (FIG.4G) and A549 (FIG.4H). Genomic segments manifesting consistent CNAs are encircled by blue dashed boxes. Copy number profiles at specific genomic loci, accurately inferred by DTCFinder but imprecisely inferred by InferCNV, are marked by yellow stars. FIGS.5A-5C illustrate DTCFinder’s specificity with normal lung tissues and non- cancerous liquid biopsy samples. FIGS.5A and 5B: Histograms depict accumulated outlier bins (AOBs) for baseline immune cells (blue) with Gaussian fitting contrasted with non- immune epithelial cells from normal lung tissues (green) from patients Normal-L1 (FIG.5A) and Normal-L2 (FIG.5B). No cells in these samples surpassed the threshold to be falsely identified as DTCs / CTCs. FIG.5C: Total cell number and false positive cells identified by DTCFinder in PBMC and CSF samples collected from a cohort of patients diagnosed with multiple sclerosis (MS) and idiopathic intracranial hypertension (IIH). FIGS.6A-6D show in-depth analysis of genome-wide CNA by DTCFinder for DTCs / CTCs identified in liquid biopsy samples from lung cancer patients in this study. FIG. 6A: Alignments of genome-wide CNAs of DTCs / CTCs inferred by DTCFinder with those of paired primary tumor tissues ascertained by LC-WGS are illustrated for patients DTC-N1, DTC-S2, and DTC-N2. Genomic segments with consistent CNAs are outlined by blue dashed boxes. FIG.6B: Depictions of inferred genome-wide CNA profiles by DTCFinder for DTCs / CTCs from patients DTC-N3, DTC-U1, DTC-N4, DTC-N5, and DTC-S3. FIG. 6C: Heatmaps showing pairwise Spearman’s correlations of inferred CNAs across all identified DTCs / CTCs in patients DTC-S1, DTC-S2, and DTC-S3, respectively. FIG.6D: Genome-wide CNAs measured by LC-WGS for 11 randomly selected single CTCs isolated from the peripheral blood of patient DTC-S1 reveals highly consistent (correlated) CNA patterns across different CTCs, corroborating the high correlations of the inferred CNA profiles. FIGS.7A-7C show evaluation of CTC identification by direct clustering analysis. FIGS.7A and 7B: The Louvain clustering method applied to scRNA-seq data from patients DTC-S1 (FIG.7A) and DTC-S2 (FIG.7B) is illustrated, with CTC-residing clusters (left) and CTCs (right) color-coded and represented in UMAP projections. FIG.7C: The inferred genome-wide CNAs by DTCFinder for CTCs and remaining cells within the CTC-residing clusters are displayed for patients DTC-S1 (top) and DTC-S2 (bottom). FIGS.8A-8E show transcriptome characterization of DTCs / CTCs and corresponding primary tumors. FIG.8A: A heatmap illustrates the expression levels of marker genes delineating different SCLC subtypes across identified CTCs from peripheral blood samples of patients DTC-S1 and DTC-S2. FIGS.8B and 8C: The top and bottom 30 differentially expressed genes in CTCs from SCLC patients DTC-S1 (FIG.8B) or DTC-S2 (FIG.8C) are juxtaposed with primary tumor cells from patients harboring the same SCLC subtype as well as normal lung samples, with the averaged gene expression levels of each sample represented based on normalized z-scores. FIGS.8D and 8E: Representative differentially enriched transcriptome signatures between CTCs from SCLC patients DTC-S1 (FIG.8D) or DTC-S2 (FIG.8E) and primary tumor cells from patients with the same SCLC subtype are color- coded and illustrated with normalized z-scores. The data bars for CTCs are magnified fivefold for enhanced clarity. FIGS.9A-9D illustrate increased expression of cancer-testis antigens (CTAs) in DTCs / CTCs. FIG.9A: Comparative analysis of expression levels and frequencies of representative CTAs amongst CTCs from SCLC patient DTC-S1, primary tumor cells from six patients of the same SCLC subtype, and epithelial cells from four normal lung tissues. FIG.9B: Comparative analysis of expression levels and frequencies of representative CTAs between CTCs from SCLC patient DTC-S2, primary tumor cells from 13 patients with the same SCLC subtype, and epithelial cells from 4 normal lung tissues. For both FIG.9A and FIG.9B, normalized z-scores show averaged CTA expression across single cells within each sample, with circle sizes indicating expression frequencies (the percentage of cells expressing the CTA). FIGS.9C and 9D: Genomic segments with DTC / CTC-specific CNAs coinciding with loci of overexpressed CTAs are outlined by green dashed boxes, with associated CTAs enumerated below for patients DTC-S1 (FIG.9C) and DTC-N1 (FIG.9D). FIGS.10A-10C show design and evaluation of DTCFinder’s model for tissue-of- origin and tumor type (TOTT) tracing. FIG.10A: A flowchart elucidates the construction of the recursive logistic regression (LR)-based Originomics Mapper, designed to trace TOTT utilizing single-cell transcriptomic data of DTCs / CTCs. FIG.10B: The representation of TOTT prediction outcomes, utilizing transcriptome data from patients across 31 solid tumor types within the TCGA database, is conducted in a 10-fold cross-validation setting. Scores exceeding zero indicate a positive prediction corresponding to a specific tumor type. FIG. 10C: Sensitivities and precisions of TOTT predictions rendered by DTCFinder spanning 31 TCGA solid tumor types with or without the integration of the recursive process. FIGS.11A-11B show performance of DTCFinder on low-depth sequencing data. FIG.11A: Comparison of inferred genome-wide CNA profiles between original sequencing depth and low-depth (5K reads / cell) data across three lung cancer samples. FIG.11B: DTCFinder’s TOTT prediction using downsampled, low-depth (5K reads / cell) data across three lung cancer samples. FIG.12 is a graph showing improvement in CTC recovery rate using the modified enrichment protocol. FIG.13 is a block diagram of an example computing system in which exemplary described aspects can be implemented. FIG.14 is a block diagram of an example cloud computing environment that can be used in conjunction with the technologies described herein. FIGS 15A-15E show schematic depiction of differential transcriptional states and drug sensitivities of DTCs and CTCs identified and characterized using DTCFinder. FIG. 15A: A dimensionality reduction plot (Uniform Manifold Approximation and Projection, UMAP) illustrating that DTCs / CTCs isolated from peripheral blood (PB) and malignant pleural effusion (MPE) samples from a lung cancer subject (designated DTC-U2) are distributed among six distinct clusters (C0-C5), each corresponding to a unique transcriptional state (or clone). Levels of expression for representative epithelial, mesenchymal, and mesothelial marker genes are shown in the lower panel. FIG.15B: A stacked bar chart showing the fraction of CTC (from PB) or DTC (from MPE) for each of the clusters C0-C5. FIG.15C: A dot plot showing the eight most differentially expressed genes in each transcriptional cluster, wherein the color and size of each dot represent, respectively, the average gene expression level across all single cells within a cluster and the proportion of cells within said cluster expressing the gene. FIG.15D: Analysis of the enrichment of representative molecular signatures within individual DTCs / CTCs. FIG.15E: Assessment of differential drug sensitivities among each cell cluster (representing distinct transcriptional clones) to selected FDA-approved or investigational oncology compounds. FIGS 16A-16B show comparison of ITC identification performance between a dual- component GMM with CD45+ cells as baseline (FIG.16A) and a three-component GMM with CD45- cells as baseline for outlier identification (FIG.16B). DETAILED DESCRIPTION The methods disclosed herein provide for the unbiased discovery, tumor type tracing, and deep molecular analysis of rare, genuine disseminated or circulating tumor cells (DTCs / CTCs) across various types of liquid biopsy samples (e.g., peripheral blood, pleural effusion, cerebrospinal fluid). Notably, these methods are applicable across nearly all solid tumor types, transcending the limitations imposed by the need for cancer-type specific molecular markers. Traditional approaches, typically reliant on epithelial markers or morphological characteristics, are hampered by low sensitivity and high false-positive rates due to the phenotypic diversity of DTCs / CTCs and the presence of non-tumor cells with epithelial traits in liquid biopsies. Moreover, existing methods fall short in determining their tissue-of-origin and obtaining accurate, unbiased single-cell transcriptome data of DTCs / CTCs in a high- throughput fashion, crucial for predictive cancer diagnostics and in-depth molecular investigations of these metastatic cells. Improvements made by the present methods include the following: Label-free DTC / CTC Discovery Approach: The disclosed methods demonstrate remarkable accuracy in distinguishing DTCs / CTCs from normal cells in a wide array of liquid biopsy samples, independent of their epithelial traits. Accurate CNA inference and Validation: The disclosed methods demonstrate a superior ability to infer genome-wide copy number alterations (CNAs) of DTCs / CTCs, which have been rigorously validated against whole-genome sequencing data of DTCs / CTCs and bulk tumor tissues. This performance surpasses existing methods in terms of precision and reliability. Accurate Tissue Origin and Tumor Type Tracing: The disclosed methods leverage DTC / CTC transcriptome data to trace the tissue-of-origin and tumor types of DTCs / CTCs, with remarkable accuracy even in samples with minimal DTC / CTC presence. This is very important for the clinical utilization of DTCs / CTCs in early cancer diagnosis. Comprehensive DTC / CTC Transcriptome Characterization: The disclosed methods excel in deriving accurate and unbiased single-cell transcriptome profiles of DTCs / CTCs. This capability enables comprehensive transcriptome analysis, revealing distinct transcriptomic signatures characteristic of these metastatic precursors and allowing for in silico discovery of potential drug targets to eliminate these metastatic cells. Robust Performance in Low-Depth Sequencing: The disclosed methods demonstrate the robust performance in identifying DTCs / CTCs, inferring genome-wide CNAs, and tracing tissue origins and tumor types with low-depth sequencing data, enhancing the practical applicability in diverse clinical and research settings where deep sequencing is not always viable, and significantly reducing the cost of identifying and analyzing DTCs / CTCs. I. Summary of Terms Unless otherwise noted, technical terms are used according to conventional usage. Definitions of many common terms in molecular biology may be found in Krebs et al. (eds.), Lewin’s genes XII, published by Jones & Bartlett Learning, 2017. As used herein, the singular forms “a,” “an,” and “the,” refer to both the singular as well as plural, unless the context clearly indicates otherwise. For example, the term “a cell” includes singular or plural cells and can be considered equivalent to the phrase “at least one cell.” As used herein, the term “comprises” means “includes.” It is further to be understood that any and all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for nucleic acids or polypeptides are approximate, and are provided for descriptive purposes, unless otherwise indicated. Although many methods and materials similar or equivalent to those described herein can be used, particular suitable methods and materials are described herein. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. To facilitate review of the various aspects, the following explanations of terms are provided: CD45: Also referred to as “leukocyte common antigen.” A cell surface protein tyrosine phosphatase expressed by most immune and hematological cells, except for mature erythrocytes and platelets. CD45 nucleic acid and amino acid sequences are publicly available and include GenBank Accession Nos. NM_002838.5 and NP_002829.3, respectively. CD61: The integrin subunit beta 3, expressed in platelets. CD61 nucleic acid and amino acid sequences are publicly available and include GenBank Accession Nos. NM_000212.3 and NP_000203.2, respectively. CD66b: Also known as CEA cell adhesion molecule 8 (CEACAM8), involved in cell-cell adhesion. CD66b is expressed on granulocytes. CD66b nucleic acid and amino acid sequences are publicly available and include GenBank Accession Nos. NM_001816.4 and NP_001807.2, respectively. Circulating tumor cell (CTC) / Disseminated tumor cell (DTC): Cells derived from a primary tumor that are present in blood (CTCs) or other body fluids or parts of the body (DTCs). Isolated tumor cells (ITCs), including single cells or cell clusters typically having a size of not more than 0.2 mm, are a type of DTCs that have detached from the primary tumor and spread through direct invasion or intra-tumoral lymphatic vessels to peritumoral tissues and regional lymph nodes (LNs). In some examples, CTCs / DTCs are metastatic precursor cells. Unlike circulating tumor DNA or tumor-derived exosomes, CTCs / DTCs are intact tumor cells that include the entire spectrum of molecular information of the cells. DTCs encompass all tumor cells that have spread from the primary tumor to other parts of the body, including those invading peritumoral regions, local lymph nodes, and those circulating in the blood and other body fluids. CTC refers specifically to tumor cells that are found circulating in the blood. In some examples, CTCs or DTCs are present in whole blood or peripheral blood mononuclear cell (PBMC) fraction. In other examples, DTCs are present in body fluids such as cerebrospinal fluid, pleural effusion, urine, or bile. Copy number alteration (CNA): Gain or loss of genetic material, such as gain or loss of one or more copies of a gene or portion thereof in the autosomal genome. In some examples, CNA burden for a sample is determined by calculating the cumulative genetic length of segments of gain or loss and expressed as a percentage of total autosomal genome size. Erythrocyte: Also referred to as red blood cell (RBC), erythrocytes are typically biconcave disks consisting mainly of hemoglobin. Mature erythrocytes do not include a nucleus. Erythrocytes express a variety of membrane proteins, including the blood group antigens such as the A, B, and Rh antigens, and glycophorin A and B. Glycophorin A: One of the major sialoglycoproteins of the erythrocyte membrane. Glycophorin A nucleic acid and amino acid sequences are publicly available and include GenBank Accession Nos. NM_002099.8 and NP_002090.4, respectively. Granulocyte: A type of immune cell including cytoplasmic granules that are released in response to infection, allergic reactions, or asthma. Basophils, eosinophils, neutrophils, and mast cells are all types of granulocytes. Leukocyte: White blood cells, including granulocytes, lymphocytes (including B cells and T cells), natural killer (NK) cells, macrophages, and dendritic cells. Platelet: Blood cells derived from megakaryocytes whose major function is hemostasis. Platelets are also involved in inflammatory processes, for example, in response to infection or injury. Sample: Any sample that contains or could contain CTCs or DTCs. In some the sample is obtained from a subject (such as a human or veterinary subject). Biological samples, include, for example, liquid biopsy samples, and solid biopsy samples. Liquid biopsy samples include, but are not limited to, whole blood, serum, plasma, peripheral blood mononuclear cells, urine, saliva, cerebral spinal fluid (CSF), bronchoalveolar lavage (BAL) fluid, pleural effusion or malignant pleural effusion, ascites, or bile. Solid biopsy samples include, but are not limited to, biopsy of peritumoral tissues (referring to non-tumor tissues located in close proximity to a tumor), or lymph node biopsy. Subject: Living multi-cellular vertebrate organisms, a category that includes both human and non-human animals. In some aspects herein, the subject has or is suspected of having a tumor or cancer. Tissue-of-Origin and tumor type (TOTT): The tissue from which a CTC or DTC is derived is referred to as the “tissue-of-origin.” For example, the tissue-of-origin may be lung in the case of CTCs or DTCs from a primary lung tumor. The “tumor type” is a histopathological classification of the type of tumor from the tissue. In some examples, the tumor type of a lung tumor may include small cell lung carcinoma (SCLC) or non-small cell lung carcinoma (NSCLC). II. Methods of Enriching Circulating Tumor Cells or Disseminated Tumor Cells Methods of enriching CTCs or DTCs from a sample are provided herein. In some aspects, the sample is a blood sample (such as whole blood or peripheral blood mononuclear cell sample). In some aspects, the methods include contacting the sample with a plurality of tetrameric antibodies, wherein each tetrameric antibody includes a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by a selected cell type; applying the sample contacted with the plurality of tetrameric antibodies to a first density gradient medium and performing a first density gradient centrifugation to obtain a first population of enriched DTCs or CTCs; resuspending the remaining cells following the first density gradient centrifugation, applying the resuspended cells to a second density gradient medium, and performing a second density gradient centrifugation to obtain a second population of enriched DTCs or CTCs; and combining the first and second populations of enriched DTCs or CTCs, thereby enriching the DTCs or CTCs (for example, compared to the starting sample). In some examples, the methods result in recovery of about 60-90% of CTCs or DTCs in the starting sample (such as about 60%, 65%, 70%, 75%, 80%, 85%, 90%, or more of the CTCs or DTCs in the starting sample). In some examples, the methods result in recovery of greater than 80% of CTCs or DTCs in the starting sample. In some examples, the sample is contacted with the plurality of tetrameric antibodies for a sufficient time for the first and second portions of the antibodies to bind to the antigens. In some examples, the sample is incubated with the plurality of tetrameric antibodies for about 20-30 minutes at 15-25ºC. In one non-limiting example, the sample is incubated with the plurality of tetrameric antibodies for about 20 minutes at room temperature. In some aspects of the methods, the second portion of the plurality of tetrameric antibodies specifically binds to an antigen expressed by leukocytes, an antigen expressed by platelets, or an antigen expressed by granulocytes. In some examples, the plurality of tetrameric antibodies includes each of a tetrameric antibody including a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by leukocytes, a tetrameric antibody including a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by platelets, and a tetrameric antibody including a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by granulocytes. In particular examples, the antigen expressed by leukocytes is CD45, the antigen expressed by platelets is CD61, and / or the antigen expressed by granulocytes is CD66b. In additional examples, antigen expressed by erythrocytes is glycophorin A. In some examples, the plurality of tetrameric antibodies includes or consists of tetrameric antibodies that specifically bind to glycophorin A and CD45, tetrameric antibodies that specifically bind to glycophorin A and CD61, and tetrameric antibodies that specifically bind to glycophorin A and CD66b. In some examples, the enrichment methods do not include incubating the sample with tetrameric antibodies that target antigens that are not glycophorin A, CD45, CD61, or CD66b. A tetrameric antibody refers to a molecule or complex formed by linking together any two antibodies (including full-length antibodies and any antigen-binding fragments such as scFv, Fv, Fab, Fab', Fab'-SH, and F(ab')2), such that the resulting molecule or complex includes four antigen-binding domains. Methods to make tetrameric antibodies are known in the art, e.g., as described in US 6872567 B2 and US 4868109 A (incorporated by reference herein in their entireties). In some aspects, a first antibody and a second antibody (corresponding to the first and second portions of a tetrameric antibody, respectively) are covalently linked together to form the tetrameric antibody. Such linkage can be formed through a crosslinking reaction, which, for example, may involve the use of a bifunctional crosslinker like N-Succinimidyl 3-(2-pyridyldithio)propionate (SPDP) . In some other aspects, the first antibody and second antibody are linked together through non-covalent interactions. In some examples, a third antibody (e.g., an F(ab')2) that recognizes the Fc portion of the first and second antibodies is used to link the first and second antibodies together. Exemplary tetrameric antibodies for use in the disclosed methods can be produced by one of ordinary skill in the art or are commercially available (such as RosetteSep™ reagents from StemCell™ Technologies). In some aspects of the methods, one or more steps are carried out in low protein binding tubes or ultra-hydrophilic polymer coated tubes. In some examples, contacting the sample with the plurality of tetrameric antibodies is carried out in low protein binding tubes or ultra-hydrophilic polymer coated tubes. In other examples the first density gradient centrifugation, the second density gradient centrifugation, or both are carried out in ultra- hydrophilic polymer coated centrifuge tubes. Low protein binding tubes can reduce cell loss during antibody labeling and improve recovery of CTCs and DTCs. Ultra-hydrophilic polymer coated centrifuge tubes can minimize loss of cells due to non-specific binding to the walls of the tubes and further improve recovery of CTCs and DTCs. Exemplary low protein binding tubes include Protein LoBind® tubes (Eppendorf, catalog number 0030122240). One of ordinary skill in the art can select other appropriate low protein binding tubes, for example, Pierce low protein binding tubes or Corning Costar low binding tubes. Exemplary ultra-hydrophilic polymer coated tubes include STEMFULL™ centrifuge tubes (S-bio, catalog number MS-90150Z), OOSAFE® centrifuge tubes (Astec Bio USA), and PROTEOSAVE™ centrifuge tubes (Sbio, catalog number MS- 52150Z). One of ordinary skill in the art can select other appropriate ultra-hydrophilic polymer coated tubes. In some aspects, the disclosed methods include two density gradient centrifugation steps. Inclusion of a second density gradient centrifugation step permits collection of residual unlabeled cells not collected in the first density gradient centrifugation step, further enriching the DTCs or CTCs recovered from the sample. The density gradient centrifugation steps utilize a density gradient medium, such as Lymphoprep™ density gradient medium (StemCell™ Technologies). One of ordinary skill in the art can select other appropriate density gradient medium reagents, including Ficoll-Paque™ medium, Percoll™ medium, Ficoll™ medium, or Histopaque™ medium. In some aspects, following contacting the sample with the plurality of tetrameric antibodies, the sample is transferred to a centrifuge tube containing a first density gradient medium and a first density gradient centrifugation is performed. The sample may be diluted prior to transferring to the centrifuge tube containing the density gradient medium. In some examples, the first density gradient centrifugation is at about 1200 x g for about 20-30 minutes and does not include braking. Following the first density gradient centrifugation, the cell layer including the first population of enriched DTCs or CTCs, is collected. In some aspects, following the first density gradient centrifugation, the remaining cells are collected and applied the to a second density gradient medium and a second density gradient centrifugation is performed. In some examples, the second density gradient centrifugation is at about 1200 x g for about 20-30 minutes and does not include braking. Following the second density gradient centrifugation, the cell layer including the second population of enriched DTCs or CTCs, is collected. The first and second populations of enriched DTCs or CTCs are combined, thereby enriching the DTCs or CTCs. In other aspects, the sample is a liquid biopsy sample that does not include erythrocytes or platelets, such as cerebrospinal fluid (CSF), pleural effusion (for example, malignant pleural effusion (MPE)), ascites, urine, bile, pleural lavage fluid, or peritoneal lavage fluid. In some examples, the sample is MPE, ascites, urine, or lavage fluid and the methods include filtering the sample using a membrane filter with a pore size of 70 microns, and centrifuging the filtered sample at 300 g for 10 minutes to separate the cell pellets from the supernatant. The method further includes treating the cell pellet with red blood cell (RBC) lysing buffer to lyse and remove RBCs from the sample. Leukocyte depletion is performed, for example with EasySep™ Human CD45 Depletion Kit II from STEMCELL Technologies and magnetic separation to isolate the labeled immune cells (leukocytes). After magnetic separation, the unwanted immune cells remain adhered to the tube due to the magnetic particles and the supernatant, which contains the desired circulating tumor cells (CTCs) along with any unlabeled immune cells, is separated from the pellet. The cells in the supernatant are resuspended in phosphate-buffered saline (PBS) for single-cell RNA sequencing or other analysis. In other examples, the sample is CSF or bile and the methods include centrifugation of the samples at 300 g for 10 minutes to separate the cell pellet and resuspending the cell pellet in 0.5 mL of PBS for single-cell RNA sequencing or other analysis. This type of sample is not subjected to further enrichment processes. In some examples, the methods result in recovery of about 60-90% of CTCs or DTCs in the starting sample (such as about 60%, 65%, 70%, 75%, 80%, 85%, 90%, or more of the CTCs or DTCs in the starting sample). The disclosed methods may include additional steps, including but not limited to, washing (for example, rinsing tubes to maximize retention of cells or washing cell pellets), diluting samples or sample fractions, lysis of red blood cells, and / or other steps. In additional examples, the methods may include resuspension of cells (such as CTCs or DTCs) in a buffer or other appropriate solution for downstream steps, such as RNA-sequencing or other analysis. III. Systems and Methods for Identifying Circulating Tumor Cells or Disseminated Tumor Cells Methods for identifying DTCs or CTCs in a biological sample are provided. In some aspects, the methods include receiving an RNA sequencing dataset (such as a single cell RNA-seq dataset) for the sample and identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells, thereby identifying DTCs or CTCs in the biological sample. In some examples, the RNA-seq dataset is from a sample obtained from a subject having or suspected of having a tumor. In some examples, the RNA-seq dataset is from a sample that has been enriched for DTCs or CTCs, for example utilizing the methods described in Section II. In further examples, single cell RNA sequencing is performed on the enriched DTCs or CTCs to obtain a scRNA-seq dataset. In some examples, the scRNA-seq dataset is low-depth scRNA-seq data. In other examples, the scRNA-seq dataset is high- depth scRNA-seq data. In some aspects, identifying the genome-wide CNA signature in one or more individual cells includes i) segmenting the scRNA-seq data to a plurality of genomic bins each comprising a pre-defined number of genes from a genome; ii) determining a mean expression intensity for the genes in each bin; iii) determining genome-wide CNA inference from the expression intensity for each bin using Gaussian mixture modeling (GMM); iv) determining copy number deviation from a baseline for each bin; and v) identifying a number of outlier bins wherein copy number deviates from the baseline. In some examples, a dual-component GMM is used (for example, when analyzing liquid biopsy samples, such as blood samples). In some examples, a GMM with a higher number of components, such as three or more components, is used for analyzing samples including more diverse cell types, such as ITCs (for example, in samples from peritumoral tissues or lymph nodes). In some examples, for identifying DTCs or CTCs in liquid biopsy samples or ITCs in LNs, the copy number deviation is determined by utilizing a baseline copy number profile for each genomic bin based on the posterior probabilities of CD45+ immune cells (excluding those with PTRPC expression levels within the first quartile of all CD45+ cells). In some examples, for identifying ITCs in peritumoral tissues where most of the cells are CD45 negative (CD45-) normal tissue cells, the copy number deviation is determined by utilizing a baseline copy number profile for each genomic bin based on the posterior probabilities of CD45- cells. The median absolute deviation (MAD) of these posterior probabilities is computed for each bin. The copy number deviation of each cell from this baseline for a specific genomic bin is determined by calculating the difference between its posterior probability and the sample median of the posterior probabilities for that bin. In some examples, if this deviation exceeds 5×MAD, the cell is classified as a copy number outlier for that particular bin. In some examples, the method further includes identifying the DTCs or CTCs when the number of outlier bins where copy number deviates from the baseline is greater than a threshold p-value. In one example, the threshold p-value is 10-8. In additional examples, the method further includes one or more pre-processing steps prior to segmenting the scRNA-seq data. The pre-processing steps may include gene / cell filtering, data normalization, or other steps. A non-limiting example of the method is illustrated in FIG.1B and includes receiving an scRNA-seq data set 110; data preprocessing 120; segmentation, feature selection, and expression intensity calculation 130; genome-wide CNA inference using Gaussian mixture modeling 140; determining CNA outlier statistics using MAD estimator 150; and DTC or CTC identification based on accumulated outlier bins (AOB) 160. The methods can further include additional steps illustrated in FIG.1B for identifying tissue origin and tumor type, which are discussed in Section IV. IV. Systems and Methods for Identifying Tumor Type or Tissue of Origin of Circulating Tumor Cells or Disseminated Tumor Cells Also provided are methods for identifying tumor type or tissue of origin (TOTT) of CTCs or DTCs. In some aspects, the method includes applying a gene signature set (or marker gene set) to DTC or CTC transcriptome data (such as scRNA-seq data). Each marker gene set identifies and distinguishes a tumor type. An aggregate expression level of a marker gene set is calculated and converted to an arbitrary scale (such as 0 to 1). A higher expression level results in a higher score, indicating the same feature for a particular tumor type. A machine learning model (such as machine learning regression model), which has been trained on aggregate gene marker signature scores for predicting tumor type, is applied. In one non-limiting example, the machine learning model is a recursive logistic regression. In some examples, a recursive logistic regression increases accuracy of identifying the TOTT, for example, to 99%. However, one of skill in the art can select other machine learning models that could be utilized in the disclosed methods. In some aspects, the methods include identifying disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) in a biological sample, for example by the methods described in Section III, constructing signature gene sets for a plurality of solid tumor types for tumor type-specific score calculations, determining tumor type-specific gene signature scores of the DTCs or CTCs from scRNA-seq data obtained from the sample; and identifying the tumor type or tissue of origin based on a machine learning model (for example, a recursive logistic regression model) prediction using tumor type specific gene signature scores as input. In other aspects, the methods include receiving a DTC or CTC single cell RNA sequencing (scRNA-seq) dataset from a sample; determining a gene signature score of DTCs or CTCs from the scRNA-seq dataset obtained from the sample; comparing the gene signature score of the DTCs or CTCs in the sample with one or more tumor type-specific signature scores; and identifying the tumor type or tissue of origin based on the gene signature score. In some examples, identifying the tumor type or tissue of origin utilizes recursive logistic regression. In some examples, the RNA-seq dataset is from a sample obtained from a subject having or suspected of having a tumor. In some examples, the RNA-seq dataset is from a sample that has been enriched for DTCs or CTCs, for example utilizing the methods described in Section II. In further examples, single cell RNA sequencing is performed on the enriched DTCs or CTCs to obtain a scRNA-seq dataset. In some examples, the scRNA-seq dataset is low-depth scRNA-seq data. In other examples, the scRNA-seq dataset is high- depth scRNA-seq data. In additional examples, the method further includes one or more pre-processing steps. The pre-processing steps may include gene / cell filtering, data normalization, or other steps. In some examples, cells exhibiting low library complexity, characterized by either gene counts less than 200 or total feature UMI counts less than 500 are filtered out. In additional examples, cells displaying a substantial fraction of mitochondrial transcripts, exceeding 1 / 3 of their content, are excluded. Erythroid cells are also removed. In further examples, each cell’s feature UMI counts are normalized by dividing by its total UMI counts, scaling by a factor of 105. After this normalization, genes are retained for subsequent analysis if their total normalized counts summed over all cells exceed 10. Finally, the normalized UMI counts were log2-transformed, with a pseudocount of 1. A non-limiting example of the method is illustrated in FIG.10A and includes receiving a DTC transcriptome dataset (e.g., a DTC RNA-seq dataset) 1020, preprocessing and normalization 1030, prediction of cancer of origin (e.g., TOTT) 1040 utilizing a core prediction unit (CPU) 1050, and determining if the prediction of cancer of origin is unique 1070. If the prediction of origin is unique, the predicted cancer type is reported 1080 and the process ends 1090. If the prediction of origin is not unique (e.g., if multiple possible predictions of origin are determined) 1085, a recursive modeling process is applied to the CPU 1050 and prediction of cancer of origin 1040 is repeated. This process is iterated until a unique prediction of origin is obtained 1070, at which point the predicated cancer type is reported 1080 and the process ends 1090. In some non-limiting examples, the CPU 1050 includes integrated tumor transcriptome data 1060 (such as transcriptome data from one or more databases such as The Cancer Genome Atlas (TCGA), Molecular Signatures Database (MSigDB), or from one or more particular cancer types), preprocessing and normalization 1062, generating a gene signature prediction for each tumor type 1064, calculating a signature score 1066, and logistic regression model creation 1068. This process can be performed iteratively, incorporating new transcriptome data from databases (such as Human Tumor Atlas Network (HTAN)) and / or through iterative processing of datasets using the disclosed methods. V. Methods of Identifying a Druggable Cancer Target and Drug Against the Cancer Methods of identifying a druggable cancer target and a cancer drug against the target are also provided. In some examples, the methods include a) receiving RNA-seq data of disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) (for example, DTCs or CTCs identified by the methods provided herein); b) determining transcriptional states of the DTCs or CTCs encoding a drug target based on the RNA-seq data; and c) identifying a drug targeting the drug target. In some examples, the methods include applying a drug and drug target identification module utilizing a functional module states (FM-states) framework. In some examples, the FM-states framework utilizes public databases including (i) cancer cell line databases such as Cancer Cell Line Encyclopedia (CCLE) database, and (ii) drug sensitivity for cancer cell lines databases such as Genomic Drug Sensitivity for Cancer (GDSC), Cancer Therapeutics Response Portal (CTRP, e.g., v2), or PRISM, to identify ML-molecular signatures for drug targets. In some examples, the RNA-seq data is aggregated scRNA-seq data or is non- aggregated scRNA-seq data. VI. Methods of Assessing Drug Sensitivity Methods of assessing sensitivity of disseminated tumor cells (DTCs) and circulating tumor cells (CTCs) to a compound or agent (e.g., a drug, a biologic, a gene therapy, a cell therapy, etc.) are provided. The compound or agent includes those that have an established therapeutic (including cytotoxic, oncolytic, etc.) effect (such as an FDA-approved drug or treatment), and those that are under investigation. In some examples, the methods include: a) receiving RNA-seq data of disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) (for example, DTCs or CTCs identified by the methods provided herein); b) determining transcriptional states of the DTCs or CTCs based on the RNA-seq data; c) computationally assessing sensitivity to a compound or agent based on the transcriptional states; and d) identifying the compound or agent as having or not having predicted increased efficacy against the transcriptional states. In some examples, the above method comprises in b) sorting the transcription states into two or more clusters, in c) computationally assessing the compound or agent sensitivity based on the transcription states of the two or more clusters, and in d) identifying the compound as having or not having increased predicted efficacy against the transcription state of at least one of the clusters. In some examples, the method further comprises selecting a subject with cancer for treatment with the compound, wherein the compound is selected for treating the subject if the compound is determined as having increased predicted efficacy against the transcriptional states, or against the transcription state of at least one of the clusters. In some examples, the method further comprises administering the selected compound to the subject. In some examples, the DTCs or CTCs are from the subject. Transcription state refers to a profile including expression levels of a pre-selected panel of genes. In some examples, the panel of genes include epithelial cell markers (e.g., EPCAM and CDH1), mesothelial cell markers (e.g., WT1 and PDPN), mesenchymal cell markers (e.g., CDH2, VIM, and COL1A2), and / or epithelial-to-mesenchymal transition (EMT) markers. Increased efficacy means the compound or agent is expected to or will produce a therapeutical effect greater than a baseline level that is considered as no efficacy. In some examples, the computationally assessing includes applying a drug sensitivity prediction algorithm (for example, the PERCEPTION algorithm or FM-states). In some examples, the algorithm utilizes public pharmacological screening and drug sensitivity databases (such as GDSC, CTRP, or PRISM) and reference molecular signatures derived from cancer cell line databases (such as TCGA or CCLE). In some examples, the RNA-seq data is aggregated scRNA-seq data or is non-aggregated scRNA-seq data. VII. Samples and Subjects In several aspects, the biological sample is from a subject having or suspected of having cancer or a tumor. In some aspects, the biological sample is a liquid biological sample from a subject, also referred to as a liquid biopsy sample. In some examples, the sample is blood or a blood fraction (for example, whole blood or peripheral blood mononuclear cells) or anther body fluid, such as cerebrospinal fluid, pleural effusion (such as malignant pleural effusion), ascites, urine, bile, pleural lavage fluid, or peritoneal lavage fluid. In some aspects, the biological sample is a solid biological sample from a subject, such as a tissue biopsy sample, such as a peritumoral tissue biopsy, or lymph node biopsy. In some aspects, the subject has or is suspected of having a solid tumor. Examples of solid tumors, include sarcomas (such as fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, and other sarcomas), synovioma, mesothelioma, Ewing sarcoma, leiomyosarcoma, rhabdomyosarcoma, colon cancer, colorectal cancer, peritoneal cancer, esophageal cancer, pancreatic cancer, breast cancer (including basal breast carcinoma, ductal carcinoma and lobular breast carcinoma), lung cancer, ovarian cancer, prostate cancer, liver cancer (including hepatocellular carcinoma), gastric cancer, squamous cell carcinoma (including head and neck squamous cell carcinoma), basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, medullary thyroid carcinoma, papillary thyroid carcinoma, pheochromocytoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, Wilms tumor, cervical cancer, fallopian tube cancer, testicular tumor, seminoma, bladder cancer, kidney cancer (such as renal cell cancer), melanoma, and CNS tumors (such as a glioma, glioblastoma, astrocytoma, medulloblastoma, craniopharyrgioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma and retinoblastoma). Solid tumors also include tumor metastases (for example, metastases to the lung, liver, brain, or bone). In particular examples, the subject has or is suspected of having a lung tumor, a colorectal tumor, a liver tumor, a prostate tumor, a breast tumor, or a melanoma. In some aspects, the disclosed methods can be used, for example, to determine if surgical removal of tumor tissue is appropriate, and / or if certain treatments or treatment methods are appropriate for use in the subject. In particular examples, the methods further include treating the subject. In some examples, the treatment includes one or more of surgery, radiation therapy, chemotherapy, immunotherapy, or other therapies. A skilled clinician can select appropriate therapies for the subject, depending on factors such as the subject, the cancer being treated, treatment history, and other factors. VIII. Aspects of the Disclosure Aspect 1. A method for identifying disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) in a biological sample, comprising: a) enriching DTCs or CTCs from the biological sample; b) performing single cell RNA-sequencing (scRNA-seq) on the enriched DTCs or CTCs to obtain scRNA-seq data; and c) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells, thereby identifying DTCs or CTCs in the biological sample. Aspect 2. A method for identifying disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) in a biological sample, comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset for the sample; and b) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells, thereby identifying DTCs or CTCs in the biological sample. Aspect 3. The method of any one of aspects 1 or 2, wherein identifying the genome-wide CNA signature in one or more individual cells comprises: i) segmenting the scRNA-seq data to a plurality of genomic bins each comprising a pre-defined number of genes from a genome; ii) determining a mean expression intensity for the genes in each bin; iii) determining genome-wide CNA inference from the expression intensity for each bin using Gaussian mixture modeling; iv) determining copy number deviation from a baseline for each bin; and v) identifying a number of outlier bins wherein copy number deviates from the baseline. Aspect 4. The method of aspect 3, further comprising identifying the DTCs or CTCs when the number of outlier bins where copy number deviates from the baseline is greater than a threshold p-value. Aspect 5. The method of aspect 4, wherein the threshold p-value is 10-8. Aspect 6. The method of any of the prior aspects, wherein the scRNA-seq dataset is low- depth scRNA-seq data or high-depth scRNA-seq data. Aspect 7. The method of any of the prior aspects, wherein the biological sample comprises a liquid biological sample. Aspect 8. The method of aspect 7, wherein the liquid biological sample comprises whole blood, peripheral blood mononuclear cells, cerebrospinal fluid, pleural effusion, urine, or bile. Aspect 9. The method of any one of aspects 1 and 3 to 8, wherein the biological sample is a blood sample, and enriching the CTCs from the blood sample comprises: i) contacting the sample with a plurality of tetrameric antibodies, wherein each tetrameric antibody comprises a first portion that specifically binds to erythrocytes and a second portion that specifically binds to an antigen expressed by a selected cell type; ii) applying the sample contacted with the plurality of tetrameric antibodies to a first density gradient medium and performing a first density gradient centrifugation to obtain a first population of enriched cells; and iii) resuspending the first population of enriched cells, applying the resuspended cells to a second density gradient medium, and performing a second density gradient centrifugation to obtain a second population of enriched cells comprising DTCs or CTCs. Aspect 10. The method of any one of aspects 1-6, wherein the biological sample comprises a solid biological sample. Aspect 11. The method of aspect 10, wherein the solid biological sample comprises a peritumoral tissue biopsy, or a lymph node biopsy. Aspect 12. The method of aspect 10 or aspect 11, wherein the DTCs are isolated tumor cells (ITCs), and the Gaussian mixture model includes at least three components. Aspect 13. A method of identifying a tumor type or tissue of origin of a tumor in a biological sample from a subject, comprising: a) identifying disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) in the biological sample by the method of any one of aspects 1 to 12; b) constructing signature gene sets for a plurality of solid tumor types for tumor type- specific score calculations; c) determining tumor type-specific gene signature scores of the DTCs or CTCs from the scRNA-seq data obtained from the sample; and d) identifying the tumor type or tissue of origin based on recursive logistic regression model prediction using tumor type specific gene signature scores as input. Aspect 14. A method of identifying a tumor type or tissue of origin of a tumor in a biological sample from a subject, comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset from a sample; b) determining a gene signature score of the DTCs or CTCs from the scRNA-seq data obtained from the sample; c) comparing the gene signature score of the DTCs or CTCs in the sample with one or more tumor type-specific signature scores; and d) identifying the tumor type or tissue of origin based on the gene signature score. Aspect 15. The method of aspect 14, wherein identifying the tumor type or tissue of origin utilizes recursive logistic regression. Aspect 16. The method of any one of aspects 13 to 15, wherein the scRNA-seq dataset is low-depth scRNA-seq data or high-depth scRNA-seq data. Aspect 17. The method of any one of aspects 13 to 16, wherein the biological sample comprises a liquid biological sample or a solid biological sample. Aspect 18. The method of aspect 17, wherein the liquid biological sample comprises whole blood, peripheral blood mononuclear cells, cerebrospinal fluid, pleural effusion, urine, or bile. Aspect 19. The method of aspect 17, wherein the solid biological sample comprises a peritumoral tissue biopsy or a lymph node biopsy. Aspect 20. The method of any one of aspects 13 to 19, wherein the tumor type or tissue of origin is lung, colorectal, liver, prostate, breast, or melanoma Aspect 21. A method of enriching disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) from a biological sample, comprising: a) contacting the sample with a plurality of tetrameric antibodies, wherein each tetrameric antibody comprises a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by a selected cell type; b) applying the sample contacted with the plurality of tetrameric antibodies to a first density gradient medium and performing a first density gradient centrifugation to obtain a first population of enriched DTCs or CTCs; c) resuspending the remaining cells following the first density gradient centrifugation, applying the resuspended cells to a second density gradient medium, and performing a second density gradient centrifugation to obtain a second population of enriched DTCs or CTCs; and d) combining the first and second populations of enriched DTCs or CTCs, thereby enriching the DTCs or CTCs. Aspect 22. The method of aspect 21, wherein the second portion of the plurality of tetrameric antibodies specifically binds to an antigen expressed by leukocytes, an antigen expressed by platelets, or an antigen expressed by granulocytes. Aspect 23. The method of aspect 22, wherein the plurality of tetrameric antibodies comprises each of a tetrameric antibody comprising a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by leukocytes, a tetrameric antibody comprising a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by platelets, and a tetrameric antibody comprising a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by granulocytes. Aspect 24. The method of aspect 22 or aspect 23, wherein the antigen expressed by leukocytes is CD45, the antigen expressed by platelets is CD61, and the antigen expressed by granulocytes is CD66b. Aspect 25. The method of any one of aspects 21 to 24, wherein the antigen expressed by erythrocytes is glycophorin A. Aspect 26. The method of any one of aspects 21 to 25, wherein the biological sample comprises a liquid biological sample. Aspect 27. The method of aspect 26, wherein the liquid biological sample comprises whole blood or peripheral blood mononuclear cells. Aspect 28. The method of any one of aspects 21 to 27, wherein the contacting the sample with the plurality of tetrameric antibodies is carried out in low protein binding tubes or ultra- hydrophilic polymer coated tubes. Aspect 29. The method of any one of aspects 21 to 28, wherein the first and / or second density gradient centrifugation is carried out in ultra-hydrophilic polymer coated centrifuge tubes. Aspect 30. The method of any one of aspects 21 to 29, further comprising diluting the sample contacted with the plurality of tetrameric antibodies prior to applying to the first density gradient medium. Aspect 31. The method of any one of aspects 21 to 30, wherein the first and / or second density gradient centrifugation is at 1200 g for 20 minutes with no braking. Aspect 32. The method of any one of aspects 21 to 31, wherein the biological sample is from a subject with a tumor or suspected of having a tumor. Aspect 33. The method of aspect 32, wherein the tumor is a solid tumor. Aspect 34. The method of aspect 33, wherein the solid tumor is a lung tumor, a colorectal tumor, a liver tumor, a prostate tumor, a breast tumor, or a melanoma. Aspect 35. A disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) identification system comprising: one or more processors; and memory coupled to the one or more processors, wherein the memory comprises computer-executable instructions causing the one or more processors to perform a process comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset for a sample; and b) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells. Aspect 36. The system of aspect 35, wherein identifying the genome-wide CNA signature in one or more individual cells comprises: i) segmenting the scRNA-seq data to a plurality of genomic bins each comprising a pre-defined number of genes from a genome; ii) determining a mean expression intensity for the genes in each bin; iii) determining genome-wide CNA inference from the expression intensity for each bin using Gaussian mixture modeling; iv) determining copy number deviation from a baseline for each bin; and v) identifying a number of outlier bins wherein copy number deviates from the baseline. Aspect 37. One or more computer-readable media having encoded thereon computer- executable instructions that, when executed, cause a computing system to perform a disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) identification method comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset for a sample; and b) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells. Aspect 38. The computer-readable media of aspect 37, wherein identifying the genome- wide CNA signature in one or more individual cells comprises: i) segmenting the scRNA-seq data to a plurality of genomic bins each comprising a pre-defined number of genes from a genome; ii) determining a mean expression intensity for the genes in each bin; iii) determining genome-wide CNA inference from the expression intensity for each bin using Gaussian mixture modeling; iv) determining copy number deviation from a baseline for each bin; and v) identifying a number of outlier bins wherein copy number deviates from the baseline. Aspect 39. A tumor type or tissue of origin of a tumor identification system comprising: one or more processors; and memory coupled to the one or more processors, wherein the memory comprises computer-executable instructions causing the one or more processors to perform a process comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset from a sample; b) constructing signature gene sets for a plurality of solid tumor types for tumor type- specific score calculations. c) determining tumor type-specific gene signature scores of the DTCs or CTCs from the scRNA-seq data obtained from the sample; d) identifying the tumor type or tissue of origin based on the recursive logistic regression model predictions using tumor type specific gene signature scores as input. Aspect 40. One or more computer-readable media having encoded thereon computer- executable instructions that, when executed, cause a computing system to perform a tumor type or tissue of origin of a tumor identification system method comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset from a sample; b) determining a gene signature score of the DTCs or CTCs from the scRNA-seq data obtained from the sample; c) comparing the gene signature score of the DTCs or CTCs in the sample with one or more tumor type-specific signature scores; and d) identifying the tumor type or tissue of origin based on the gene signature score. Aspect 41. A method of identifying a compound effective against a drug target, comprising: a) receiving an RNA-seq dataset of disseminated tumor cells (DTCs) or circulating tumor cells (CTCs); b) determining transcriptional states of the DTCs or CTCs based on the RNA-seq dataset; c) computationally assessing the compound sensitivity based on the transcriptional states; and d) identifying the compound as having or not having increased predicted efficacy against the transcriptional states. Aspect 42. The method of aspect 41, wherein the method comprises in b) sorting the transcription states into two or more clusters, in c) computationally assessing the compound sensitivity based on the transcription states of the two or more clusters, and in d) identifying the compound as having or not having increased predicted efficacy against the transcription state of at least one of the clusters. Aspect 43. The method of aspect 41 or 42, further comprising selecting a subject with cancer for treatment with the compound, wherein the compound is selected for treating the subject if the compound is determined as having increased predicted efficacy against the transcriptional states, or against the transcription state of at least one of the clusters . Aspect 44. The method of aspect 43, further comprising administering the selected compound to the subject. Aspect 45. The method of any one of aspects 41-44, wherein the determining and computationally assessing comprises applying a drug and drug target identification model comprising a functional module states (FM-States) framework or drug sensitivity prediction algorithm. Aspect 46. The method of aspect 45, wherein the DTC or CTC RNA-seq data is aggregated single cell RNA-seq data. Aspect 47. The method of any one of aspects 41 to 46, wherein the DTCs or CTCs are identified according to the method of any one of aspects 1 to 20. EXAMPLES The following examples are provided to illustrate particular features of certain aspects of the disclosure, but the scope of the claims should not be limited to those features exemplified. Example 1 Materials and Methods This example describes the methods utilized in the experiments described in Example 2. Patient information and sample collection: Liquid biopsy samples including peripheral blood, malignant pleural effusion (MPE), and cerebrospinal fluid (CSF) in this study were obtained from Shanghai Chest Hospital between November 2019 and July 2022. The protocols were performed according to the principles of the Helsinki Declaration and approved by the Institutional Review Board (#KS1973, #IS21109). Cell lines and reagents: Lung cancer cell lines H1975, H1650 and A549 were obtained from American Type Culture Collection (ATCC). Cell lines were routinely maintained in ATCC-formulated cell culture medium containing 10% fetal bovine serum (FBS, Gibco) and 1× Penicillin-Streptomycin-Glutamine (Gibco) in a humidified atmosphere of 5% CO2 and 95% air at 37ºC. Reagents used in this study are listed in Table 1 below. Table 1. List of reagents REAGENT or RESOURCE SOURCE IDENTIFIER Anti-HK2 primary antibody Abcam #ab209847 0 0 DTC / CTC enrichment and single-cell RNA sequencing: Peripheral blood (5 ml), MPE (10 ml), and CSF (1 ml) were collected and sent to the lab at 4 ºC immediately after collection for sample processing and high-throughput scRNA-seq. MPE samples were filtered by a membrane with a pore size of 70 μm and centrifuged at 300 g for 10 min to separate cell pellets. Cells were then treated with red blood cell (RBC) lysing buffer (BD Biosciences) to remove RBCs, followed by depletion of leukocytes with EasySep™ Human CD45 Depletion Kit II (STEMCELL Technologies). The cell pellets were re-suspended in 0.5 mL of PBS. CSF samples were centrifuged at 300 g for 10 min to separate cell pellet, and re-suspended in 0.5 mL of PBS without further enrichment. Peripheral blood samples were centrifuged at 500 g for 5 min, and the cell pellets were re-suspended in an equivalent volume of HBSS and mixed with 75 μl tetrameric antibody cocktail (customized RosetteSep™ CTC Enrichment Cocktail, STEMCELL Technologies) at room temperature for 20 min, followed by adding 15 ml of HBSS with 2% FBS and mixing well. The mixture was carefully added along the wall of the Sepmate tube (SepMate™-50) after adding 15 ml density gradient liquid (Lymphoprep™) into the tube through the middle hole. After centrifuging at 1200 g for 20 min, the topmost supernatant (~10 ml) was discarded, and the remaining liquid (~10 ml) above the barrier of the Sepmate tube was rapidly poured out into a new centrifuge tube with addition of 20 ml of HBSS. After centrifuging at 600 g for 8 min, the supernatant was removed and 1 ml of RBC lysing buffer was then added for 5 min to lyse RBCs. After centrifuging at 450 g for 5 min, the nucleated cell pellet was re-suspended in 100 μl of HBSS for single-cell RNA sequencing. For tissue biopsies, they were dissociated using the human tumor dissociation kit (Miltenyi Biotec). Specifically, 50 μL Enzyme H, 25 μL Enzyme R, and 6.25 μL Enzyme A were added to 1.1 mL serum-free RPMI 1640 media (Thermo Fisher) to prepare the dissociation solution. Dissected biopsies were rinsed once with PBS, transferred to the dissociation solution, and incubated at 37°C with 1800 rpm shaking for 15 minutes. The resulting cell suspensions were filtered through a 30 μm strainer (Miltenyi Biotec) to remove undigested tissue. The cell suspension was further spun down at 400×g for 5 minutes to collect all the dissociated cells. The supernatant was removed, and the cells were washed once with PBS to remove all the residual dissociation solution. After 400×g for 5 minutes centrifuge, the cells were resuspended with HBSS for single-cell RNA sequencing. High-throughput scRNA-seq was conducted with a target recovery of 8,000 cells. Cell counting and cell viability were measured with a Countess II Automated Cell Counter, and the cell concentration was adjusted to 1000 cells / μL. Cells were loaded into a Chromium Single-Cell 3 Chip Kit v2 and v3 (10× Genomics). Reverse transcription, cDNA recovery, cDNA amplification, and library construction were performed using Chromium Next GEM Single Cell 3’ Reagent Kits v2 and v3 (10x Genomics) according to manufacturer’s instructions. Single-cell library sequencing was performed using the Illumina NovaSeq 6000 platform, with 150 bp paired-end reads. Isolation of genomic DNA from primary tumors and copy number alternation profiling: Two serial sections from each individual formalin-fixed and paraffin-embedded (FFPE) tumor tissue block were used to isolate genomic DNA (gDNA) by GeneRead DNA FFPE Kit (Qiagen) according to manufacturer’s instructions. Concentrations of gDNA were quantified with Qubit dsDNA HS Assay Kit (Invitrogen). WGS library was constructed with the NEBNext® Ultra™ DNA Library Prep Kit for Illumina (New England Biolabs) according to manufacturer’s protocols. Libraries were analyzed by Illumina Novaseq 6000 platform with 150 bp pair-end reads (Genewiz, China). CTC identification and retrieval, whole genome amplification, and copy number alternation profiling for patient sample DTC-S1: After blood sample processing and CTC enrichment, the enriched cell suspension was applied onto a 3% BSA-treated poly-L-lysine glass slide. Cells on the slide were fixed (2% PFA, 10 min), permeabilized (0.5% Triton X- 100, 15 min) and blocked with a solution containing 3% BSA and 10% Normal Goat Serum for 1 hour. On-chip immunostaining was then performed by incubation with APC- conjugated mouse anti-human CD45 antibody, eFluor 570-conjugated mouse anti-human PanCK antibody, and rabbit anti-human hexokinase 2 (HK2) antibody in PBS overnight at 4 ºC. After extensive washing with PBS, cells on the chip were treated with Alexa Fluor 488- conjugated goat-anti-rabbit secondary antibody in PBS for 1 h and DAPI for 10 min followed by washing with PBS. ImageXpress Micro XLS Wide field High Content Screening System (Molecular Devices) scanned the chip and imaged all cells in bright field and fluorescent channels (CD45: CY5; HK2: FITC, CK: PE, Nucleus: DAPI). A computational algorithm analyzed the images to identify CTCs adhering to previously validated criteria reported before (Yang et al., PNAS 118:e2012228118, 2021). Specifically, it identified putative CTCs characterized by distinct marker profiles, including DAPIpos / CD45neg / HK2high / CKpos, DAPIpos / CD45neg / HK2high / CKneg, and DAPIpos / CD45neg / HK2low / CKposcells, based on the calculated fluorescence cutoffs generated from HK2 and CK fluorescence intensities of CD45posleukocytes in the same sample. Single-cell LC-WGS was used to characterize genome-wide CNA profiles. Single- cell whole genome amplification (WGA) was firstly conducted with the MALBAC® Single Cell WGA Kit (Yikon Genomics). To assess the WGA coverage of amplified product, 22 primer pairs were designed to target 22 loci located on different chromosomes. Six primer pairs were randomly selected for PCR and a successful amplification of at least four out of six primer pairs generated a positive quality control (QC) PCR. WGA products that passed QC-PCR were then used to construct WGS library with the NEBNext® Ultra™ DNA Library Prep Kit for Illumina (New England Biolabs) or MGIEasy Universal DNA Library Prep Kit (MGI Tech) according to manufacturer’s protocols. The concentrations of purified fragmented DNA or libraries were quantified with Qubit dsDNA HS Assay Kit (Invitrogen). Libraries were sequenced by Illumina Navoseq 6000 with 150 bp pair-end reads (Genewiz, China) or MGI2000 sequencer with 100 bp single-end reads (JunHealth, China). Pre-processing of single-cell RNA sequencing data: FASTQ files of patient samples were individually processed to generate a feature-barcode matrix (genes × cells) using CellRanger version 6.1.2. The processing was performed with the refdata-cellranger- GRCh38-3.0.0 human genome reference, employing default parameters tailored for the 10X single-cell 3’ library and setting expected cell counts to 10,000. Genuine cells were distinguished from empty droplets based on the cumulative distribution of total transcript counts. Cells exhibiting low library complexity, characterized by either gene counts less than 200 or total transcript counts less than 500 were filtered out. Additionally, cells displaying a substantial fraction of mitochondrial transcripts, exceeding 1 / 3 of their content, were excluded. Prior to inferring single-cell CNAs using DTCFinder, erythroid cells were also removed. Construction and implementation of the DTCFinder algorithm: The overall workflow of DTCFinder is depicted in FIG.1A and includes the following steps: (1) Data normalization and transformation Each cell's feature unique molecular identifier (UMI) counts were normalized by dividing by its total UMI counts, scaling by a factor of 105. After this normalization, genes were retained for subsequent analysis if their total normalized counts summed over all cells exceeded 10. Finally, the normalized UMI counts were log2-transformed, with a pseudocount of 1. (2) Segmentation, feature selection, and expression intensity calculation The entire human genome (GRCh38.p13) was sequentially segmented by chromosomes, ensuring each segment (or bin) exclusively encompassed a pre-defined number of genes. A default of 150 genes per bin was chosen based on its optimal balanced accuracy and minimal FDR (FIG.4F). Gene positional information was procured using the R package biomaRt, referencing the ENSEMBL_MART_ENSEMBL dataset within the BioMart database and the hsapiens_gene_ensembl dataset. For every genomic bin per cell, the expression intensity was computed as the mean expression of the encompassed 150 genes, centralized to the cell's average gene expression, and subsequently utilized for CNA inference. (3) Genome-wide CNA inference using Gaussian mixture modeling (GMM) To discern genome-wide CNAs for individual cells, a GMM was implemented. Given that normal cells and DTCs / CTCs in liquid biopsy samples have distinct CNA signatures, a dual-component GMM was computed for each genomic segment over the calculated expression intensities. A three-component GMM was used for identifying ITCs from peritumoral tissue biopsy. Explicitly for each bin, the Expectation-Maximization (EM) algorithm acted on the z-score standardized expression intensities across cells, leveraging the R package mixtools, to determine maximum likelihood estimates and posterior probabilities. The posterior probabilities, as derived from the EM algorithm for Gaussian distribution mixtures, were subsequently utilized to simulate the cells’ CNAs over each genomic bin (Muller et al., Bioinformatics 34:3217-3219, 2018). This simulation procedure was reiterated eight times, selecting the GMM that maximized output likelihood estimates. Convergence of the EM algorithm was confirmed if the data log-likelihood varied by less than ε = 10-8during iteration and was deemed divergent if iterations exceeded 1000 times. Any genomic bin manifesting consistent divergence across all 8 simulations was excluded from further scrutiny. This methodology culminated in a posterior probability matrix (cells × bins) for single cells over the genomic bins modeled with definitive GMMs. These posterior probabilities were z-score standardized across cells and can be visualized as a heatmap post- standardization using pheatmap 1.0.12 package. Given that in liquid biopsies, immune cells and other normal cells considerably outnumber DTCs / CTCs and manifest homogenous CNA patterns, their z-scored posterior probabilities are approximated to zero, barring occasional outlier gene expressions in specific cells. This represents the diploid copy number baseline. In contrast, DTCs / CTCs typically display pronounced z-scores in altered chromosomal segments due to substantial expression variations of genes in those bins with respect to normal immune cells, highlighting potential amplified or deleted chromosomal regions. (4) Copy number outlier statistics and DTC identification DTCs / CTCs typically exhibit much more pronounced genome-wide CNAs compared to normal cells. The extent of these CNAs across the genome aids in DTC identification. To establish a baseline copy number profile for each genomic bin, we utilized the posterior probabilities of CD45+ immune cells (excluding those with PTRPC expression levels within the first quartile of all CD45+ cells). The median absolute deviation (MAD) of these posterior probabilities was then computed for each bin. The copy number deviation of each cell from this baseline for a specific genomic bin was determined by calculating the difference between its posterior probability and the sample median of the posterior probabilities for that bin. If this deviation exceeded 5×MAD, the cell was classified as a copy number outlier for that particular bin. Subsequently, we quantified the degree of genome- wide CNAs for each cell by tallying the number of bins identified as outliers. The accumulated outlier bins (AOBs) for the baseline immune cells were minimal and conformed to a Gaussian distribution. Conversely, DTCs / CTCs typically displayed significantly heightened numbers of AOBs. These can be pinpointed from the extreme right tail of the Gaussian distribution, correlating with a notably low p-value indicative of statistical significance. A cell was designated as a DTC if its AOB value surpassed the threshold set by a p-value of 10-8relative to the Gaussian distribution of baseline immune cells (FIG.1C). For ITC identification, CD45-negative cells were used as the baseline for outlier statistics and AOB calculations. (5) Recursive Originomics Mapper for tracing DTCs / CTCs’ tissue-of-origin and tumor type To ascertain the TOTT of DTCs / CTCs, we devised a machine learning-based Recursive Originomics Mapper (ROM) module, constructing 31 classifiers corresponding to 31 distinct solid tumor types. This model was trained using an integrated transcriptomic dataset, which included 30 solid tumor transcriptomes (comprising 8880 samples) from TCGA and an SCLC dataset from a previous study (George et al., Nature 524:47-53, 2015), inclusive of 81 human SCLC samples across various subtypes. For each tumor type, we curated a unique gene signature. This signature was formulated by leveraging the integrated transcriptomic data to pinpoint genes manifesting elevated expression in a specific tumor type while maintaining subdued expression in other types. To homogenize these diverse transcriptomic datasets and negate batch effects, we started with Illumina HiSeq percentile datasets sourced from the TCGA hub (UCSC Xena). These datasets rank gene RSEM values within a 0% to 100% spectrum for each specimen. The SCLC datasets were similarly transformed into percentile ranks, mirroring the methodology applied to the TCGA data. Subsequently, a signature gene set was curated for each tumor type (Table 2). This was achieved by selecting genes that met the following differential expression criteria in the tumor type of interest: (1) the gene's average percentile rank across all samples of the target tumor type exceeded 50%, (2) this percentile was at minimum double the average percentile of that gene across samples from other tumor types, and (3) a statistical significance threshold of p < 10-20was met, as determined by the Mann- Whitney U test. Table 2. List of signature genes for 31 solid tumor types used in the Recursive Originomics Mapper (ROM) module Tumor Signature genes Type , 2, 3, 5, , , 6, , , , Tumor Signature genes Type , , 1, , , Tumor Signature genes Type , P, 4, 7, L, , Tumor Signature genes Type 3, , , , Tumor Signature genes Type , 6, 1, , A, 2, 3, 5, Tumor Signature genes Type 2, , A, 3, , , , , , 1, , C, , 3, , 1, Tumor Signature genes Type 1, 2, , , , 5, , , , 2, Tumor Signature genes Type , , 3, 3, , 3, , Tumor Signature genes Type B, , L, 4, , , 1, , , Tumor Signature genes Type H, 3, 5, 1, 3, Tumor Signature genes Type , D, 9, 4, , 6, 2, Tumor Signature genes Type , , 2, 1, , , , Tumor Signature genes Type 6, E, , , , C, Tumor Signature genes Type 3, adenocarcinoma; CHOL: cholangiocarcinoma; COAD: colon adenocarcinoma; GBM: glioblastoma multiforme; HNSC: head and neck squamous cell carcinoma; KICH: kidney chromophobe; KIRC: kidney renal clear cell carcinoma; KIRP: kidney renal papillary cell carcinoma; LGG: brain lower grade glioma; LIHC: liver hepatocellular carcinoma; LUAD: lung adenocarcinoma; LUSC: lung squamous cell carcinoma; MESO: mesothelioma; OV: ovarian serous cystadenocarcinoma; PAAD: pancreatic adenocarcinoma; PCPG: pheochromocytoma and paraganglioma; PRAD: prostate adenocarcinoma; READ: rectum adenocarcinoma; SARC: sarcoma; SCLC: small cell lung cancer; SKCM: skin cutaneous melanoma; STAD: stomach adenocarcinoma; TGCT: testicular germ cell tumor; THCA: thyroid carcinoma; THYM: thymoma; UCEC: uterine corpus endometrial carcinoma; UCS: uterine carcinosarcoma; UVM: uveal melanoma Having derived the gene signature set for each tumor type, we computed a signature score for every sample within the integrated transcriptomic dataset by averaging the sample percentile expressions for all signature genes. Consequently, each sample was endowed with a unique signature score for each of the 31 tumor types, leading to an aggregate of 31 scores for every sample. By extending this computation to the entire sample set, we generated a matrix delineating samples against tumor-type scores. Using this matrix as input, we formulated and validated a classifier for each tumor type. This was executed via logistic regression modeling, employing the R package Epi 2.47 and harnessing the iteratively reweighted least squares (IRLS) methodology for deducing maximum likelihood estimates during the model fitting. To counteract potential overfitting, we instituted a constraint on feature (tumor type) selection, mandating that the number of features in each classifier be capped at five. For every tumor type, its classifier incorporated the signature score of that tumor alongside scores of up to four other tumor types. This integration resulted in a myriad of tumor type combinations for each classifier. For every tumor type's classifier, model training and parameter tuning were executed through a repeated cross-validation approach. We elected the combination boasting the highest balanced accuracy when distinguishing a specific tumor type from the cumulative 8961 samples spanning the 31 solid tumor categories. These refined 31 classifiers collectively formed the core prediction unit (CPU) of the ROM module (FIG.10A, Table 3). Table 3. List of classifiers of 31 solid tumor types used in first round of TOTT prediction. Tumor Classifiers type ACC -141091640127725+406822622324007+0225809771138399*ACC - - Tumor Classifiers type PRAD -51.8690316079766+16.4504533687825+2.85404234513628*PRAD- D D - 1 With the CPU in place, the algorithm processes a given sample's transcriptomic data, determines signature scores for the 31 solid tumor types, and employs the CPU to predict the sample's TOTT. A sample is designated a particular tumor type (positive prediction) if its type-specific logit ≥ 0. Owing to the overlapping transcriptomic signatures observed between liver (LIHC) and bile duct (CHOL) cancers, as well as between colon (COAD) and rectal (READ) cancers, the algorithm consolidates these analogous cancer types into a unified class for predictive purposes. If a singular positive prediction appears, that specific cancer type is reported. In the absence of positive predictions, the outcome is labeled “undetermined.” However, for scenarios with multiple positive predictions, the algorithm initiates a recursive process. In this recursive phase, the ROM module redefines the signature gene sets and logistic regression model solely for the tumor types positively identified in the initial round. Utilizing only the transcriptomic data of these specified tumor types from the integrated transcriptomic dataset, the algorithm recalibrates gene signatures based on our earlier criteria, computes updated signature scores, and then reconstructs the classifiers via logistic regression modeling using these refreshed scores. Using this overhauled CPU, the algorithm then revisits the task of determining the sample's TOTT. If a singular positive prediction emerges, that specific cancer type is reported. Otherwise, if no new predictions or the same number of positive predictions as the previous round arise, the cancer types identified in the initial round are reported. Should a reduced number of positive predictions manifest, the recursive procedure is invoked once more, refining the signature gene sets and logistic regression models tailored to the tumor types positively predicted in the second round. This iterative process persists until a definitive prediction is secured or until the predictive outcomes stabilize (FIG.10A). To gauge the ROM module's efficacy, we undertook a 10-fold cross-validation on the integrated transcriptomic dataset encompassing the 31 solid tumor types. In this evaluation, samples were arbitrarily segmented into ten equally sized groups. The model was trained using nine of these groups and tested on the remaining one, cycling through all groups until each had served as the testing dataset. This validation yielded a sensitivity of 92.9±5.8% and a precision of 94.8±6.4%, attesting to the model’s robust performance (FIG.10B). The finalized model, trained on the entire transcriptomic dataset of the 31 tumor types, was then primed for predicting DTCs / CTCs' TOTT. Using the algorithm, DTC / CTC transcriptomic data from a specific sample underwent preprocessing described above, with gene expression levels averaged across all discerned DTCs / CTCs to fashion a percentile-ranked dataset. Leveraging this data, the algorithm determined signature scores for the 31 solid tumor types and employed the established ROM module to predict that sample's TOTT. Visualization and clustering of single-cell RNA sequencing data: To test whether unsupervised clustering can directly separate DTCs / CTCs into distinct clusters (FIGS.7A and 7B), we employed Seurat to process, visualize, and cluster scRNA-seq data from patient samples DTC-S1 and DTC-S2 (Hao et al., Cell 184:3573-3587, 2021). These specific samples were chosen due to their origin from peripheral blood, containing a relatively large number of DTCs / CTCs, which aptly represent the heterogeneity of DTCs / CTCs in circulation and provide a sufficient sample size of DTCs / CTCs for clustering analysis. Unless otherwise specified, default parameters were maintained throughout the process. Raw count data underwent normalization in Seurat using the “LogNormalize” approach followed by the selection of top 10,000 variable features using vst method and data scaling and centralization over these selected features. We then used UMAP projections to generate lower dimensional representations on the first 50 principal components of the chosen features using knn = 20. Original Louvain algorithm was used to optimize the 451 modularity function to determine clusters. Copy number alteration analysis from low-coverage whole genome sequencing data: FASTQ files were aligned to the major chromosomes of human (GRCh38) using bowtie2-2.3.5.1 with default parameters. PCR duplicates were removed with Samtools (version 1.11). Aligned reads were counted in fixed bins averaging 500 kb. Bin counts were normalized for GC content with lowess regression and bin-wise ratios were calculated by computing the ratio of bin counts to the sample mean bin count. The diploid regions were determined using HMMcopy (version 0.1.1). Segmentation was performed with circular binary segmentation (CBS) method (alpha = 0.0001 and undo.prune = 0.05) from R Bioconductor DNAcopy package. Copy number noise was quantitated using the mean absolute pairwise difference (MAPD) algorithm. Samples with MAPD ≤ 0.45 passed the MAPD QC and were included in single-cell genome-wide CNA analyses. Identification of spiked-in cancer cells via cell line-specific markers: To accurately discern and quantify spiked-in H1975, H1650, and A549 cells, thereby establishing a benchmark for evaluating DTCFinder's classification precision, we derived cell line-specific signature genes from existing RNA-seq profiles of non-small cell lung cancer cells (Van Der Steen et al., Cancers (Basel) 12:1091, 2020). This dataset (GSE160683) features deep RNA sequencing of 60 human lung cell lines alongside a human lung epithelial cell line and two human lung fibroblast cell lines, each with three replicates. Differentially expressed genes (DEGs) between H1975 cells and an ensemble of other cell lines, including H1650, A549, BEAS-2B (lung epithelial cells), IMR-90 (lung fibroblasts), and WI-38 (lung fibroblasts) cells, were discerned. The criteria for DEGs were an expression fold change > 2 and p < 0.05 as determined by a two-tailed Welch's t-test. Following this, the DEGs were ranked by the magnitude of expression fold change. A manual evaluation of the top 30 ranked DEGs against all other lung cancer cell lines in the dataset was conducted. This led to the selection of 7 signature genes as shown below, characterized by pronounced expression in H1975 cells and minimal expression in H1650, A549, BEAS-2B, IMR-90, WI-38, and other lung cancer cell lines. An additional requirement for these signature genes was negligible expression in immune cells, as verified via The Human Protein Atlas (proteinatlas.org / ). Using a similar methodology, seven signature genes for H1650 and A549 were also identified, respectively. When clustering the expression matrix of the 21 identified signature genes across all single-cell transcriptome data of the spiked-in sample, the spiked-in H1975, H1650, and A549 cells distinctly clustered into three separate groups (Table 4), clearly demarcated from immune cells (FIG.4E). Table 4. Lung tumor cell line gene signatures Cell line Gene Signature A549 RSPO3 ARG2 INSL4 GABRB3 GABRA5 HOXB8 NF0B1 rous patients: To gauge the false positive rate of DTCFinder when evaluating normal tissue or liquid biopsy samples from individuals without cancer, we examined scRNA-seq datasets derived from two prior studies (Schafflick et al., Nat. Commun.11:247, 2020; Vieira Braga et al., Nat. Med.25:1153-1163, 2019). These datasets encompass scRNA-seq data from four freshly excised human lung tissues, specifically parenchymal lung and distal airway specimens, taken from uninvolved sections during tumor resections in 4 patients. Additionally, the datasets include scRNA-seq data from 22 CSF and PBMC samples harvested from 6 multiple sclerosis (MS) patients and 6 idiopathic intracranial hypertension (IIH) patients. Following the preprocessing steps described above, the data underwent analysis using DTCFinder. In this evaluation, a mere 2 out of approximately 8,000 cells from a CSF sample of an MS patient were incorrectly identified as DTCs (FIG.5C). Notably, the AOBs of these misclassified cells were 494 marginally proximate to the cut-off AOB threshold, defined by a p-value of 10-8. Analysis of differential transcriptome signatures and cancer-testis antigens between DTCs / CTCs and matched primary tumor cells: To discern differential transcriptomic signatures between DTCs / CTCs and the primary tumor cells from the same cancer types, we sourced scRNA-seq datasets for NSCLC and SCLC from two prior studies (Wu et al., Nat. Commun.12:2540, 2021; Chan et al., Cancer Cell 39:1479-1496, 2021). Specifically, scRNA-seq data was extracted from 16 LUAD patient tumors (Primary-N1- N16) and 19 SCLC patient tumors spanning various subtypes (Primary-SA1-SA13 and Primary-SN1-SN6) as detailed in Table 6. For LUAD samples, UMI count data from LUAD patient samples DTC-N1 and 502 DTC-N2 were amalgamated with data from primary tumor cells of 16 LUAD patients (Primary-N1-N16). In the SCLC dataset, subtype determination for samples DTC-S1 and DTC-S2 was performed based on marker gene expression for ASCL1, NEUROD1, POU2F3, and YAP1. DTC-S1 was classified as SCLC-505 N (NEUROD1+), and DTC-S2 as SCLC-A (ASCL1+) (FIG.8A). Thereafter, UMI count data for sample DTC-S1 was paired with data from SCLC-N patient tumors (Primary-SN1-SN6), and for sample DTC-S2 with SCLC-A tumors (Primary-SA1-SA13). Table 5. Summary of patient samples collected in this study Study Age Sex Histology Stage Treatment Sample type DTC / CTC ID range (volume) number s Study Age Sex Histology Stage Treatment Sample type DTC / CTC ID range (volume) number for hepatocellular carcinoma; PRAD, prostate adenocarcinoma; BRCA, breast carcinoma; PAAD, pancreatic adenocarcinoma. Table 6. Summary of patient samples from prior studies Study ID Original sample Histology DTC Sample Publication Data ID number type source source 0 ) 5 0 94 7 97 Study ID Original sample Histology DTC Sample Publication Data ID number type source source an . 71 48 , liver hepatocellular carcinoma; PRAD, prostate adenocarcinoma; BRCA, breast carcinoma. The integrated datasets underwent preprocessing as previously described, followed by normalization by library sizes, scaling by a factor of 10,000, and natural log-transformation with a pseudocount of 1. Batch effect was assessed for the combined datasets to ensure normalized expression of immune cells of shared phenotypes well-mixes across samples. Cell type annotations for primary tumor samples were applied to identify tumor cells following the same procedures used in the original studies (Wu et al., Nat. Commun. 12:2540, 2021; Chan et al., Cancer Cell 39:1479-1496, 2021). DEGs between DTCs / CTCs and their matched primary tumor cells was ascertained using the Seurat 4.3.0 package26, with significance determined by the non-parametric Wilcoxon rank-sum test. The top 30 DEGs for both DTCs / CTCs and primary tumor cells, exhibiting a significance level of p<10-30, were visualized via a heatmap alongside epithelial cells from the aforementioned normal lung samples (Normal-L1-L4) using pheatmap 1.0.12 package based on normalized z-scores of averaged expressed levels per sample. For differential transcriptomic signature assessment, the SingleCellSignatureExplorer (v.3.6) was harnessed against gene sets of HALLMARK, C2, C3.TFT, and C5.BP from Molecular Signature Database (MSigDB 2023.1.Hs). Then limma (v.3.42.0) R package facilitated differential analysis of these signature scores, considering only gene sets with p.adj < 0.05 as differentially enriched transcriptome signatures (DETSs). Selected DETSs were depicted in a heatmap using pheatmap 1.0.12 package based on normalized z-scores of single-cell signature scores. For a nuanced exploration of cancer-testis antigen (CTA) expression differences between DTCs / CTCs and their matched primary tumor cells, we examined expression metrics of CTAs listed in the CTDatabase. CTAs absent in all samples were excluded from analysis. Both expression frequency (percentage of cells expressing a specific CTA) and mean expression levels of CTAs were computed. Representative CTAs, demonstrating at least a 10% elevated expression frequencies in DTCs / CTCs relative to matched primary tumor cells, were illustrated in FIG.2G and FIGS.9A and 9B through dot plots using ggcorrplot 0.1.4.1 package. TOTT prediction using existing data: To augment the validation of the TOTT predictive accuracy across a broader spectrum of tumor types not collected in our patient cohort, we leveraged existing DTC / CTC transcriptome datasets from various prior studies. These encompassed tumor types such as liver cancer (LIHC), melanoma (SKCM), prostate cancer (PRAD), and breast cancer (BRCA) (Ebright et al., Science 367:1468-1473, 2020; Diamantopoulou et al., Nature 607:156-162, 2022; Sun et al., Nat. Commun.12:4091, 2021; Ramskold et al., Nat. Biotechnol.30:777-782, 2012; Miyamoto et al., Science 349:1351- 1356, 2015; Aceto et al., Cell 158:1110-1122, 2014). Detailed information regarding the original data sources, sample IDs, tumor types, and DTC / CTC counts can be found in Tables 5 and 6. For each sample, the transcriptomic data was averaged across all DTCs / CTCs, with the exception of sample DTC-B3, which utilized single-CTC transcriptome data. This consolidated data was then transformed into percentile ranks as their gene expression estimations as previously described. Subsequently, this transformed dataset served as the input for the ROM module within DTCFinder for the prediction of the tissue-of-origin and specific tumor type of the DTCs / CTCs. CNA inference of DTCs / CTCs via InferCNV: To juxtapose the genome-wide CNAs inferred by DTCFinder with those from established tools, we employed InferCNV to deduce CNAs from the single-cell transcriptome data of the spiked-in lung cancer cells. InferCNV necessitates the designation of a cohort of non-malignant cells to serve as a reference baseline for the CNA estimation in malignant cells. To facilitate the comparison with DTCFinder’s results, we opted for the same group of immune cells used as the baseline in DTCFinder's copy number outlier statistics. Specifically, this comprised all CD45+ immune cells, with an exclusion criterion set for those with PTRPC expression levels within the first quartile of all CD45+ cells. This baseline was utilized in InferCNV to deduce the genome-wide CNA profiles for the spiked-in cells: H1975, H1650, and A549. In the procedure, genes were arranged based on their genomic locations for each chromosome. Function infercnv::run() is called for the CNA inference of DTCs / CTCs via InferCNV with default settings except for cutoff set to 0.1 and denoise set to TRUE. A sliding window, encompassing 150 genes, was adopted to smooth the relative expression across each chromosome, thereby mitigating gene-specific expression influences. The relative expression values were subsequently centered to 1. The ceiling and floor for visualization were set using default settings, facilitating the portrayal of the data on a heatmap with a blue-red gradient, representing deletion to amplification, respectively. Assessment of DTCFinder’s performance on low-depth single-cell RNA sequencing data: To gauge the efficacy of DTCFinder when analyzing low-depth scRNA- seq datasets, we executed a downsampling of the original fastq files associated with DTC-N1, DTC-S1, and DTC-S2 to achieve an equivalent read depth of 5,000 reads / cell using toolkit seqtk-1.4. In detail, the original fastq file of each sample underwent downsampling into three lower-depth replicates by extracting reads randomly from the original files. The derived fastq files underwent the aforementioned preprocessing to yield feature-barcode matrices, which were subsequently used for DTC / CTC identification and TOTT tracing via DTCFinder. DTCs / CTCs discerned from the original datasets functioned as the benchmark for evaluating the sensitivity and precision of DTC / CTC identification within the downsampled datasets. The single-cell transcriptome data of DTCs / CTCs, sourced from the downsampled fastq files, was channeled into DTCFinder for TOTT predictions. Analysis of CNA burden across various tumor types using TCGA data: We performed a comprehensive analysis of CNA burdens in diverse tumor types, leveraging data from TCGA. The copy number segment data of primary tumor tissues, matched normal tissues and PBMCs were obtained from the Genomic Data Commons Database (portal.gdc.cancer.gov / ). CNA burden is defined as the percentage of the tumor autosomal genome with copy number altered. To calculate CNA burden for a sample, segments manifesting copy number gains and losses were identified. The cumulative genomic length of these segments was then calculated and expressed as a percentage of the total autosomal genome size. Statistical Analysis: Descriptive statistics, group comparisons, and correlation tests were performed using GraphPad PRISM 10 (GraphPad Software, Inc) unless noted elsewhere. Example 2 Development and Implementation of DTCFinder DTCFinder includes an experimental component to enrich DTCs / CTCs for high- throughput scRNA-seq and a computational pipeline that harnesses an amalgam of machine- learning models and statistical techniques for DTC / CTC identification and analysis (FIG. 1A). Specifically, we employed tetrameric antibody complexes to co-pellet immune cells and platelets along with red blood cells (RBCs) upon density gradient centrifugation, enriching DTCs / CTCs with leftover immune cells amenable for scRNA-seq. Genome-wide CNAs often manifest at early neoplastic stages and act as a pivotal catalyst in both the onset and advancement of malignancies. Analysis of TCGA data revealed that, even at clinical stage I, CNA burdens of tumor cells markedly surpass those of normal cells adjacent to tumors or in circulation, validating CNAs as reliable DTC / CTC markers (FIG.1B and FIG.3). Utilizing the single-cell transcriptomic data, DTCFinder segments the genome into contiguous genomic bins and applies a Gaussian mixture model (GMM) to infer relative copy numbers for each bin in every single cell. The inference is predicated on the distribution of average gene expression levels of each genomic segment across different cells. By establishing the inferred CNAs of a subset of immune cells in the liquid biopsy as the baseline, an outlier detection algorithm identifies bins with anomalous copy numbers and calculates the accumulated number of such outlier bins (AOBs) for each cell. Cells exhibiting significantly elevated AOB values, indicative of extensive genome-wide CNAs, are identified as DTCs / CTCs. The AOB values in immune cells typically follow a Gaussian distribution, and the threshold for DTC identification can be determined from the extreme right tail of this distribution, corresponding to a statistically significant low p-value (FIG.1C). To empirically ascertain the optimal cut-off p-value, a mixture of three distinct lung cancer cell lines—H1975, H1650, and A549—was spiked into a blood sample, emulating the heterogeneous DTC / CTC populations typically present in liquid biopsy samples, and subsequently subjected to DTC enrichment for scRNA-seq and DTC / CTC identification (FIG.1C and FIGS.4A-4D). Utilizing a selective panel of cell line-specific markers, we quantified the cell number of each line in the resulting scRNA-seq dataset (FIG> 4E). This quantification served as a ground truth for evaluating DTCFinder's discriminative efficacy. DTCFinder reached highest balance accuracy of 98.98% for the identification of spike-in tumor cells at the cut-off p value of 10-8, accompanied by an area under the curve (AUC) of 0.998 and a precision of 99.31% (FIG.1D). Under p=10-8, 150 genes per bin was determined to produce a near-perfect balanced accuracy and the highest precision for DTC / CTC identification (FIG.4F). To rigorously validate the accuracy of DTCFinder's inferred genome-wide CNAs, we employed low-coverage whole-genome sequencing (LC-WGS) as a benchmark. Remarkably, the CNAs inferred by DTCFinder were in high concordance with those experimentally measured through LC-WGS across all three lines (FIG.1E and FIGS.4G- 4H). Furthermore, we conducted a comparative analysis with InferCNV, an existing algorithm for CNA inference. DTCFinder exhibited consistent results with InferCNV, with enhanced precision in certain genomic regions (FIG.1E and FIGS.4G-4H). Importantly, unlike DTCFinder, InferCNV requires pre-specified reference cells for CNA inference and lacks the capability for autonomous DTC detection or TOTT tracing. To ascertain whether DTCFinder's analytical robustness would be compromised by different types of non-malignant cells, we conducted tests using normal lung tissues and various liquid biopsy samples, including blood and cerebrospinal fluid (CSF) from non- cancer patients. Notably, among these samples were those from patients with multiple sclerosis (MS) containing inflammatory immune cell types. Our findings revealed that normal epithelial cells in lung tissues exhibited AOB distributions that overlapped with CD45+ cells, resulting in no false-positive identifications (FIGS.5A-5B and Tables 5 and 6). Consistently, no DTC / CTC was identified except for one CSF sample from an MS patient, where 2 false-positive DTCs were found among ~8000 cells, indicating the high specificity of DTCFinder, even in analyzing complex liquid biopsy samples replete with atypical immune cell types (FIG.5C). We further extended our analysis by applying DTCFinder to a cohort of 15 liquid biopsy samples obtained from 10 lung cancer and 2 pancreatic cancer patients, 5 of whom had paired primary tumor tissues (Table 5). The algorithm discerned varying numbers of DTCs / CTCs across different patients. Importantly, the inferred CNAs of DTCs / CTCs closely matched those of paired tumor tissues measured by LC-WGS (FIG.2A and FIGS.6A and 6B). DTC-specific CNAs in some genomic regions were also observed. Moreover, the inferred CNAs showed marked pairwise correlations across individual DTCs / CTCs, a feature validated by LC-WGS using DTCs / CTCs collected from the same patients (FIG.2B and FIGS.6C-6D). Interestingly, in certain samples, such as DTC-S1 and DTC-S2, we identified distinct subgroups of DTCs / CTCs with unique CNA profiles that were highly correlated within the subgroup but less so with other DTCs / CTCs, suggesting the existence of multiple DTC / CTC clones possibly originating from different clonal regions of the primary tumor (FIG.2B and FIG.6C). In contrast to FDA-cleared CellSearch system or other similar methods, DTCFinder enables the detection of DTCs / CTCs devoid of epithelial markers. In six out of nine lung cancer patients in our cohort, DTCFinder detected a significant number of EPCAM-negative or PanCK-negative DTCs / CTCs in peripheral blood, undetectable by CellSearch, corroborating previous observations of PanCK-negative DTCs / CTCs in lung cancer patients (FIG.2C) (Yang et al., PNAS 118:e2012228118, 2021). Furthermore, given the intrinsic heterogeneity and extreme rarity of DTCs / CTCs, coupled with the presence of confounding cell types in liquid biopsy samples, conventional unsupervised clustering methods prove inadequate for precise DTC / CTC segregation. DTCs / CTCs often do not form an isolated cluster but rather intermingle with confounding cell types, dispersing across multiple clusters within liquid biopsy samples (FIGS.7A-7B). DTCFinder, capitalizing on hallmark genome- wide CNAs, effectively distinguishes DTCs / CTCs from non-malignant cells, even when they are interspersed within the same clusters. Indeed, for the peripheral blood samples from patient DTC-S1 and DTC-S2, non-CTC cells residing in CTC-enriched clusters were absent from characteristic CNAs present in the CTCs and paired tumor tissues, affirming their non- tumoral nature (FIG.7C). DTCFinder enables not just the identification but also the comprehensive characterization of DTCs / CTCs by extracting and analyzing their single-cell transcriptome profiles. Through comparison with their matched primary tumors, it delineates DTC / CTC- specific transcriptome signatures, including differentially expressed genes and enriched biological processes and pathways, crucial for unraveling the mechanisms underlying tumor metastasis and identifying targetable vulnerabilities for these metastatic cells. (FIGS.2D-2E and FIGS.8A-8E). For instance, in NSCLC patients DTC-N1 and DTC-N2, we observed elevated expression levels in genes encoding components of the large ribosomal subunit, such as RPL41 and RPL36A, as well as genes associated with a mesenchymal phenotype, like VIM, compared with tumor cells from primary lesions or normal lung tissues (FIG.2D). Overexpression of these genes in DTCs / CTCs has been implicated in promoting tumor invasion and metastasis (Yang et al., PNAS 118:e2012228118, 2021; Ebright et al., Science 367:1468-1473, 2020). Consistently, DTCs / CTCs from these patients exhibited elevated signature scores in metastasis and TGFβ signaling pathways while showing reduced scores in apoptosis and antigen-presentation processes, thereby implicating their role in both tumor metastasis and immune evasion (FIG.2E). Increased metastasis and invasion-related transcriptome signatures were also found in DTCs / CTCs from SCLC patients DTC-S1 and DTC-S2 compared with corresponding primary tumor cells (FIGS.8D-8E). Interestingly, we detected a marked elevation in both the expression levels and frequencies of many cancer- testis antigens (CTAs) in DTCs / CTCs compared to primary tumor cells or normal lung tissues, indicative of a more aggressive or metastatic phenotype (FIGS.2F-2G and FIGS.9A- 9B). Notably, KDM5B, a testis-selective CTA, exhibited increased expression levels and frequencies in DTCs / CTCs from patients with NSCLC and SCLC (FIG.2G and FIG.9B). The upregulation of KDM5B has been linked to the development of stem-like invasive phenotypes and increased drug tolerance in a variety of cancer types, prompting the active development of pharmacological inhibitors. Some of these CTAs are located in genomic segments with DTC / CTC-specific copy number amplification, which may contribute to the increased expression of these CTAs in the DTCs / CTCs (FIGS.9C-9D). DTCFinder further enables comparative analysis of DTCs / CTCs obtained from different liquid biopsy samples collected from the same patient. For example, in a lung cancer patient designated DTC-U2, both malignant pleural effusion (MPE) and peripheral blood (PB) samples were collected and analyzed. DTCFinder identified 631 DTCs from the MPE sample and 37 CTCs from the peripheral blood sample (Table 5). Projection of single- cell transcriptomic profiles using Uniform Manifold Approximation and Projection (UMAP) revealed that these DTCs / CTCs segregated into six distinct transcriptional clones (FIG.15A). Notably, the majority of CTCs from blood were assigned to Cluster 5 (C5) (FIG.15B). Analysis of lineage-specific marker expression showed that C0 to C3 predominantly comprised epithelial cells, characterized by high expression of EPCAM and CDH1, whereas C5 was enriched for cells of mesothelial lineage expressing classical mesothelial markers such as WT1 and PDPN (FIGs.15A and 15C). Cells in C5 also exhibited elevated expression of mesenchymal-associated genes, including CDH2, VIM, and COL1A2, indicating acquisition of mesenchymal features during metastatic dissemination. C4 contained tumor cells displaying a transitional phenotype, with co-expression of epithelial, mesothelial, and mesenchymal markers, thereby underscoring the plasticity of DTCs / CTCs during metastasis. Enrichment analysis utilizing the single-cell signature explorer demonstrated that C4 and C5 were highly enriched for epithelial-to-mesenchymal transition (EMT) and invasiveness gene signatures, suggesting increased motility and metastatic potential in these cells. Interestingly, metabolic pathway analysis revealed that invasive cells within C4 and C5 exhibited a greater reliance on fatty acid metabolism, whereas tumor cells with primarily epithelial characteristics (C0-C3) favored glucose metabolism (FIG.15D). This finding suggests a metabolic reprogramming that may present novel therapeutic vulnerabilities. Importantly, both MPE and PB samples contained metastatic tumor cells with both epithelial and mesothelial lineage signatures. However, the majority of CTCs detected in PB exhibited strong mesothelial and mesenchymal features, while only a minority of DTCs in MPE displayed such signatures. These findings suggest that tumor cells circulating in blood possess enhanced invasiveness and metastatic capacity relative to those residing in MPE. Furthermore, the single-cell transcriptome profiles extracted by DTCFinder enable computational assessment of differential drug sensitivities across these transcriptional clones. In silico drug screening of DTCs / CTCs using the PERCEPTION algorithm (Sinha et al., Nature Cancer 5:938-952, 2024) across 140 FDA-approved or investigational oncology drugs identified small molecule inhibitors—such as the MDM2 inhibitor Idasanutlin and the FGFR inhibitor Erdafitinib—as being particularly effective against the highly invasive C5 population. In contrast, agents such as 12-O-Tetradecanoylphorbol-13-acetate and the MEK inhibitor RO4987655 were less effective in C5 cells relative to other populations (FIG.15E). Tracing the TOTT is pivotal for leveraging DTCs / CTCs for diagnostic applications, notably in early disease stages where the primary tumor lesion are undetectable via conventional medical imaging or in the monitoring of tumor recurrence. DTCFinder employs its built-in Recursive Originomics Mapper (ROM) module for TOTT tracing using DTC / CTC transcriptome data (FIG.10A). The ROM module commences by formulating tumor type- specific gene signatures from TCGA transcriptome data of 31 solid tumors and develops individual tumor type predictors via logistic regression, optimizing classification accuracy through 10-fold cross validation (FIG.10B). Notably, when multiple positive predictions arise for a given sample, we introduce a recursive process to the model to progressively refine the subset until a singular positive TOTT prediction is achieved. This recursive process markedly improved the sensitivity and precision of TOTT predictions across all cancer types (FIG.10C). By applying this prediction module to the 13 liquid biopsy samples collected in this study, along with 15 additional DTC / CTC samples from published data, DTCFinder demonstrated impeccable prediction accuracy even for samples containing merely a couple of DTCs / CTCs (FIG.2H and Tables 5 and 6). The accurate TOTT tracing even in cases with only a couple of DTCs / CTCs confirms DTCFinder's capability to precisely pinpoint DTCs / CTCs within samples characterized by extreme DTC / CTC scarcity. We further challenged DTCFinder’s predictive capacity on single CTCs, CTC clusters, and CTC-white blood cell (WBC) clusters where the presence of WBCs may confound the resultant transcriptome data. Using data from a recently published study that collected all three forms of CTC samples from breast cancer patients (Diamantopoulou et al., Nature607:156-162, 2022), DTCFinder rendered 100% accurate predictions across all the samples, confirming the robust TOTT tracing performance of our algorithm (FIG.2I). We further assessed DTCFinder’s proficiency in contexts with low sequencing depths. We down-sampled the raw scRNA-seq data of three patients, characterized by relatively high number of DTCs / CTCs, from 34,000-116,000 reads / cell to 5000 reads / cell to quantitatively evaluate the influence of sequencing depth on the sensitivity and precision of DTC / CTC identification. Remarkably, even under conditions of markedly reduced sequencing depth, DTCFinder maintained its accuracy in DTC / CTC identification and relatively consistent inferred genome-wide CNA profiles, achieving an average precision of 94.53±2.24% and sensitivity of 90.20±4.47% across the three patients (FIG.2J and FIG. 11A). Notably, the identified DTCs / CTCs, even with low-depth transcriptome data, could be precisely traced back to their respective TOTT, underscoring the robust performance of DTCFinder on low-depth scRNA-seq data (FIG.11B). Isolated tumor cells (ITCs), a type of disseminated tumor cells (DTCs) in peritumoral tissues and lymph nodes (LNs), are critical in cancer metastasis. ITCs, comprising single cells or cell clusters with no cluster exceeding 0.2 mm, are DTCs detaching from primary tumor and spreading through direct invasion or intratumoral lymphatic vessels to peritumoral tissues or regional LNs. ITCs are increasingly recognized for their crucial role in cancer staging, prognosis, and clinical management. In early-stage breast cancer, ITC presence in regional LNs has been identified as an adverse prognostic factor, correlating with reduced disease-free survival. Patients with ITCs benefit from systemic adjuvant therapy, which improves disease-free survival. Similarly, in colon cancer, studies have shown that the presence of ITCs in LNs is associated with significantly worse disease-free and overall survival in stage I and II patients, suggesting ITCs as a potential high-risk factor influencing treatment decisions. Furthermore, in stage I non-small cell lung cancer (NSCLC), despite curative surgery, 30-35% of patients experience recurrence, often locally in the ipsilateral lung. This phenomenon implicates, at least in part, to the presence of undetected ITCs in peritumoral tissues or regional LNs during surgery contributing to the tumor recurrence. However, the detection of ITCs to date mostly relies on histopathological staining (H&E and IHC). These methods depend on morphological assessment, the expression of cancer-type specific markers, as well as the experience of the pathologists, which often leads to imprecise screening for minute ITCs scattered in peritumoral tissues and LNs. Moreover, these pathological assays do not allow for the retrieval of ITCs for comprehensive molecular profiling, leaving the omics-level understanding of these rare metastatic cells largely uncharted. To demonstrate the capability of DTCFinder in identifying these rare ITCs in peritumoral tissues and LNs, 100 cancer cells from a primary lung tumor were spiked into pathologically negative, distal peritumoral tissue located 5 cm from the tumor margin. The tissue sample was then processed to obtain RNA-seq data. Direct application of DTCFinder showed suboptimal ITC identification using parameters identical to those for liquid biopsy analysis. However, significantly improved results were observed when using three GMM components with CD45-negative cells as the baseline for outlier statistics and AOB calculations (FIGS.16A-16B). These results suggest that the diverse cellular composition in peritumoral tissues requires the use of a higher number of GMM components in the CNA inference step, as opposed to the dual-component GMM typically used for blood samples. In summary, DTCFinder offers a bench-to-bits toolkit for label-free, full-spectrum DTC / CTC discovery and analysis across various cancer types. It encompasses a holistic view of DTCs / CTCs, ranging from genome-wide CNAs and transcriptome signatures to TOTT tracing, demonstrating versatility across diverse types of liquid biopsy and compatibility with multiple commercial scRNA-seq platforms without a demand for deep sequencing, The adaptability of DTCFinder, positions it as a useful tool, readily adoptable by research community for fundamental and clinical studies on liquid biopsy-based cancer diagnosis and cancer metastasis. Example 3 CTC Enrichment Protocol This example describes a method for the negative enrichment of circulating tumor cells (CTCs) from the blood samples of cancer patients. CTCs, shed from primary or metastatic tumor lesions, are exceedingly rare in peripheral blood, typically present at a ratio of about one cell per billion. Enriching these scarce CTCs is a significant challenge. This method leverages a customized RosetteSep™ immunodensity cell separation kit from STEMCELL Technologies Inc. This approach involves crosslinking unwanted immune cells to red blood cells using specific tetrameric antibody complexes, followed by further purification of the target cells through density gradient centrifugation. The recovery rate using the standard protocol of this kit was between 40-50%, which is suboptimal for clinical studies. The modified protocol, achieving an ~85% CTC recovery rate, a substantial improvement over existing negative depletion-based CTC enrichment methods (FIG.12). Key modifications to the method included: (1) using a customized tetrameric antibody cocktail that targets only CD45 (a common leukocyte marker), CD61 (platelet marker), CD66b (granulocyte marker), and glycophorin A (red blood cell marker). This small antibody panel minimizes the unspecific depletion of CTC; (2) employing ultra-low cell binding tubes to reduce cell loss during processing; and (3) introducing new steps in the protocol to recover leftover CTCs that remain after standard protocol application, using an additional round of density gradient centrifugation. These innovations significantly improved the efficiency and efficacy of CTC enrichment, making our method more suitable for clinical applications. Validation experiment design: To compare the standard protocol and the modified protocol, a small number of pre-labeled cancer cells were spiked into peripheral blood samples from healthy donors to mimic the CTC-containing blood samples from cancer patients. Generally, 100 fluorophore-labeled melanoma cells were picked up with micromanipulation and spiked into 5 ml human healthy donor blood. Three replicates were processed with each protocol, the recovered cells were seeded into 24 well plate, imaged, and counted. The numbers of recovered cells were compared to validate the cell recovery rates. Materials: 2% fetal bovine serum (FBS) in phosphate buffered saline (PBS) solution Lymphoprep (STEMCELL Technologies, 07801) 1×lysis buffer (BD, 559759): Mix 1ml 10×Lysis buffer with 10ml sterile water Customized RosetteSep antibody cocktails (with CD45, CD61, CD66b, and glycophorin A antibodies to co-pellet and remove immune cells and platelets in with red blood cells, customized by STEMCELL Technologies) Equipment: STEMFULL™ Centrifuge Tube, 15 ml Low cell binding (Sbio, MS-90150Z). These ultra-hydrophilic polymer coated centrifuge tubes can greatly minimize loss of cells from non-specific sticking to the walls of the tubes, thus providing high recovery of the CTCs after enrichment. Eppendorf™ Protein LoBind™ Tubes (Eppendorf™ 0030122240). The low protein binding tube can reduce the cell loss during antibody labeling step and improve the final cell recovery. Procedure: Tumor cells isolation. Reducing cell attachment is very important to CTCs enrichment, the cell number would be 50k to 100k after enrichment, and only 10 to 100 of them are CTCs. To maximize the recovery rate, three optimizations are applied in this method: (1) low cell binding tubes are used in antibody labeling and cell collection to minimize the cell loss. (2) All the labeled blood is rinsed from the wall of specimen tubes and pipettes. (3) A modified protocol with two-time density gradient centrifuge. 1. Add 150 µl of antibody cocktails into 5ml blood (in specimen tube), incubate at room temperature for 20 minutes, gently flick the tube several times after 10 min incubation. 2. Transfer 15 ml 2% FBS PBS into a 50 ml tube (washing tube) and another 50ml tube for blood dilution (dilution tube). 3. Transfer the labeled blood to the dilution tube. 4. Pour ~ 5 ml 2% FBS PBS solution into the specimen tube to rinse the tube and transfer the solution to dilution tube. 5. Repeat step 4 two times to transfer all the labeled blood into dilution tube, the blood is diluted to 4x volume after this step. 6. Take a new 50 ml tube, add 15 ml Lymphoprep density gradient medium. 7. Gently mix the diluted blood and slowly transfer the blood to the tube containing the density gradient medium, do not disturb the layers. 8. Centrifuge the blood at 1200g for 20 min at room temperature, with brake off. 9. Discard the upper layer and transfer the cell layer into a low cell binding 15 ml tube to collect the cells (6-7 ml). 10. Top up the cell collection to 14 ml with 2% FBS PBS solution to wash the enriched cells. 11. Centrifuge at 600g 8 min to collect the enriched cells, with brake low (accelerate 6 decelerate 6). 12, Remove all the lymphoprep density gradient medium after step 8 (approximately 2 ml left), add 18 ml 2% FBS PBS solution to fully resuspend the blood cells (20ml in total). 13. Mix well, repeat steps 5-10, and combine all the collected tumor cells in one 15ml tube. This second density gradient centrifugation collects most of the residual unlabeled cell, further improving the recovery rate. 14. (Optional) If the collected cell pellet is still red, then additional red blood cell lysis is required. Lower the volume to 100 µl, adding 1 ml lysis buffer (up to the number of residual RBCs), incubate at room temperature for 5 min protecting from light. Centrifuge at 300g for 5 min to collect the enriched cells. Example 4 Example Implementation of Receiving RNA-Seq Data Any of the examples herein can include receiving a variety of data, such as RNA-seq data, for example, scRNA-seq data (for example, one or more datasets that include one or more datapoints). In practice, RNA-seq data or scRNA-seq can include data on genes or sets of genes. For example, a targeted set of genes or a genome-wide set of genes can be included. In practice, receiving RNA-seq data or scRNA-seq data can include data for at least one subject (such as a subject with a known tumor type, or a training subject; or a subject with an unknown tumor type, or a query subject) or at least one group of subjects (such a group of subjects with a common feature or characteristic, or a cohort). In specific, non- limiting examples, receiving RNA-seq data or scRNA-seq data can include data for at least two cohorts, such as cohorts with a different disease status or with different phenotypes. In examples, receiving RNA-seq data or scRNA-seq data can include transcriptome data for a subject or subjects with a common feature or characteristic, such as a disease (for example, cancer, or a malignant tumor characterized by abnormal or uncontrolled cell growth). In some examples, receiving RNA-seq data or scRNA-seq data can include transcriptome data from one or more DTCs or CTCs from a sample from a subject. In specific, non-limiting examples, receiving RNA-seq data or scRNA-seq data can include transcriptome data for single subjects or a group of subjects with a common disease (such as cancer, for example, a malignant tumor characterized by abnormal or uncontrolled cell growth). In practice, receiving RNA-seq data or scRNA-seq data can further include a variety of preprocessing or processing steps. In examples, preprocessing steps can include gene / cell filtering, normalization, and / or transformation. In some examples, gene / cell filtering includes distinguishing genuine cells from empty droplets based on cumulative distribution of total transcript counts. For example, cells exhibiting gene counts less than 200 or total transcript counts less than 500 may be filtered out in a preprocessing step. In addition, cells with mitochondrial transcripts exceeding one-third of their content may be excluded in the preprocessing step. In further examples, erythroid cells are also removed. In other examples, preprocessing includes data normalization and transformation. For example, each cell’s feature UMI counts may be normalized by dividing by its total UMI counts and scaling by a factor of 105. Genes with total normalized counts >10 summed over all cells are retained. The normalized UMI counts may be log2-transformed, with a pseudocount of 1. Example 5 Example Computing System FIG.13 illustrates a generalized example of a suitable computing system 1300 in which any of the described technologies may be implemented. The computing system 1300 is not intended to suggest any limitation as to scope of use or functionality, as the innovations may be implemented in diverse computing systems, including special-purpose computing systems. In practice, a computing system can comprise multiple networked instances of the illustrated computing system. With reference to FIG.13, the computing system 1300 includes one or more processing units 1310, 1315, and memory 1320, 1325. In FIG.13, this basic configuration 1330 is included within a dashed line. The processing units 1310, 1315 execute computer- executable instructions. A processing unit can be a central processing unit, processor in an application-specific integrated circuit (ASIC), or any other type of processor. In a multi- processing system, multiple processing units execute computer-executable instructions to increase processing power. For example, FIG.13 shows a central processing unit 1310 as well as a graphics processing unit or co-processing unit 1315. The tangible memory 1320, 1325 may be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two, accessible by the processing unit(s). The memory 1320, 1325 stores software 1380 implementing one or more innovations described herein, in the form of computer-executable instructions suitable for execution by the processing unit(s). A computing system may have additional features. For example, the computing system 1300 includes storage 1340, one or more input devices 1350, one or more output devices 1360, and one or more communication connections 1370. An interconnection mechanism (not shown) such as a bus, controller, or network interconnects the components of the computing system 1300. Typically, operating system software (not shown) provides an operating environment for other software executing in the computing system 1300, and coordinates activities of the components of the computing system 1300. The tangible storage 1340 may be removable or non-removable, and includes magnetic disks, magnetic tapes or cassettes, CD-ROMs, DVDs, or any other medium which can be used to store information in a non-transitory way and which can be accessed within a computing system. The storage 1340 stores instructions for the software 1380 implementing one or more innovations described herein. The input device(s) 1350 may be a touch input device such as a keyboard, mouse, pen, or trackball, a voice input device, a scanning device, or another device that provides input to the computing system 1300. For video encoding, the input device(s) 1350 may be a camera, video card, TV tuner card, or similar device that accepts video input in analog or digital form, or a CD-ROM or CD-RW that reads video samples into the computing system 1300. The output device(s) 1360 may be a display, printer, speaker, CD-writer, or another device that provides output from the computing system 1300. The communication connection(s) 1370 enable communication over a communication medium to another computing entity. The communication medium conveys information such as computer-executable instructions, audio or video input or output, or other data in a modulated data signal. A modulated data signal is a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media can use an electrical, optical, RF, or other carrier. The innovations can be described in the general context of computer-executable instructions, such as those included in program modules, being executed in a computing system on a target real or virtual processor (e.g., that is ultimately implemented on a hardware processor). Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or split between program modules as desired in various embodiments. Computer-executable instructions for program modules may be executed within a local or distributed computing system. For the sake of presentation, the detailed description uses terms like “determine” and “use” to describe computer operations in a computing system. These terms are high-level abstractions for operations performed by a computer and should not be confused with acts performed by a human being. The actual computer operations corresponding to these terms vary depending on implementation. Example 6 Example Cloud Computing Environment FIG.14 depicts an example cloud computing environment 1400 in which the described technologies can be implemented, including, e.g., the systems described herein. The cloud computing environment 1400 comprises cloud computing services 1410. The cloud computing services 1410 can comprise various types of cloud computing resources, such as computer servers, data storage repositories, networking resources, etc. The cloud computing services 1410 can be centrally located (e.g., provided by a data center of a business or organization) or distributed (e.g., provided by various computing resources located at different locations, such as different data centers and / or located in different cities or countries). The cloud computing services 1410 are utilized by various types of computing devices (e.g., client computing devices), such as computing devices 1420, 1422, and 1424. For example, the computing devices (e.g., 1420, 1422, and 1424) can be computers (e.g., desktop or laptop computers), mobile devices (e.g., tablet computers or smart phones), or other types of computing devices. For example, the computing devices (e.g., 1420, 1422, and 1424) can utilize the cloud computing services 1410 to perform computing operations (e.g., data processing, data storage, and the like). In practice, cloud-based, on-premises-based, or hybrid scenarios can be supported. Example 7 Example Computer-Readable Media Any of the computer-readable media herein can be non-transitory (e.g., volatile memory such as DRAM or SRAM, nonvolatile memory such as magnetic storage, optical storage, or the like) and / or tangible. Any of the storing actions described herein can be implemented by storing in one or more computer-readable media (e.g., computer-readable storage media or other tangible media). Any of the things (e.g., data created and used during implementation) described as stored can be stored in one or more computer-readable media (e.g., computer-readable storage media or other tangible media). Computer-readable media can be limited to implementations not consisting of a signal. Example 8 Example Implementations Although the operations of some of the disclosed methods are described in a particular, sequential order for convenient presentation, such manner of description encompasses rearrangement, unless a particular ordering is required by specific language set forth herein. For example, operations described sequentially can in some cases be rearranged or performed concurrently. Example 9 Example Computer-Executable Implementation Any of the methods described herein can be performed by computer-executable instructions (e.g., causing a computing system to perform the method, when executed) stored in one or more computer-readable media (e.g., storage or other tangible media) or stored in one or more computer-readable storage devices. Such methods can be performed in software, firmware, hardware, or combinations thereof. Such methods can be performed at least in part by a computing system (e.g., one or more computing devices). Such acts of the methods described herein can be implemented by computer- executable instructions in (e.g., stored on, encoded on, or the like) one or more computer- readable media (e.g., computer-readable storage media or other tangible media) or one or more computer-readable storage devices (e.g., memory, magnetic storage, optical storage, or the like). Such instructions can cause a computing device to perform the method. The technologies described herein can be implemented in a variety of programming languages. In any of the technologies described herein, the illustrated actions can be described from alternative perspectives while still implementing the technologies. For example, “receiving” can also be described as “sending” from a different perspective. It will be apparent that the precise details of the methods or compositions described may be varied or modified without departing from the spirit of the described aspects of the disclosure. We claim all such modifications and variations that fall within the scope and spirit of the claims below.
Claims
We claim:
1. A method for identifying disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) in a biological sample, comprising: a) enriching DTCs or CTCs from the biological sample; b) performing single cell RNA-sequencing (scRNA-seq) on the enriched DTCs or CTCs to obtain scRNA-seq data; and c) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells, thereby identifying DTCs or CTCs in the biological sample.
2. A method for identifying disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) in a biological sample, comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset for the sample; and b) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells, thereby identifying DTCs or CTCs in the biological sample.
3. The method of any one of claims 1 or 2, wherein identifying the genome-wide CNA signature in one or more individual cells comprises: i) segmenting the scRNA-seq data to a plurality of genomic bins each comprising a pre-defined number of genes from a genome; ii) determining a mean expression intensity for the genes in each bin; iii) determining genome-wide CNA inference from the expression intensity for each bin using Gaussian mixture modeling; iv) determining copy number deviation from a baseline for each bin; and v) identifying a number of outlier bins wherein copy number deviates from the baseline.
4. The method of claim 3, further comprising identifying the DTCs or CTCs when the number of outlier bins where copy number deviates from the baseline is greater than a threshold p-value.
5. The method of claim 4, wherein the threshold p-value is 10-8.
6. The method of any one of claims 1 or 2, wherein the scRNA-seq dataset is low-depth scRNA-seq data or high-depth scRNA-seq data.
7. The method of any one of claims 1 or 2, wherein the biological sample comprises a liquid biological sample.
8. The method of claim 7, wherein the liquid biological sample comprises whole blood, peripheral blood mononuclear cells, cerebrospinal fluid, pleural effusion, urine, or bile.
9. The method of claim 1, wherein the biological sample is a blood sample, and enriching the CTCs from the blood sample comprises: i) contacting the sample with a plurality of tetrameric antibodies, wherein each tetrameric antibody comprises a first portion that specifically binds to erythrocytes and a second portion that specifically binds to an antigen expressed by a selected cell type; ii) applying the sample contacted with the plurality of tetrameric antibodies to a first density gradient medium and performing a first density gradient centrifugation to obtain a first population of enriched cells; and iii) resuspending the first population of enriched cells, applying the resuspended cells to a second density gradient medium, and performing a second density gradient centrifugation to obtain a second population of enriched cells comprising DTCs or CTCs.
10. The method of any one of claims 1 or 2, wherein the biological sample comprises a solid biological sample.
11. The method of claim 10, wherein the solid biological sample comprises a peritumoral tissue biopsy, or a lymph node biopsy.
12. The method of claim 10, wherein the DTCs are isolated tumor cells (ITCs), and the Gaussian mixture model includes at least three components.
13. A method of identifying a tumor type or tissue of origin of a tumor in a biological sample from a subject, comprising:a) identifying disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) in the biological sample by the method of any one of claims 1 or 2; b) constructing signature gene sets for a plurality of solid tumor types for tumor type- specific score calculations; c) determining tumor type-specific gene signature scores of the DTCs or CTCs from the scRNA-seq data obtained from the sample; and d) identifying the tumor type or tissue of origin based on recursive logistic regression model prediction using tumor type specific gene signature scores as input.
14. A method of identifying a tumor type or tissue of origin of a tumor in a biological sample from a subject, comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset from a sample; b) determining a gene signature score of the DTCs or CTCs from the scRNA-seq data obtained from the sample; c) comparing the gene signature score of the DTCs or CTCs in the sample with one or more tumor type-specific signature scores; and d) identifying the tumor type or tissue of origin based on the gene signature score.
15. The method of claim 14, wherein identifying the tumor type or tissue of origin utilizes recursive logistic regression.
16. The method of any one of claims 13 or 14, wherein the scRNA-seq dataset is low- depth scRNA-seq data or high-depth scRNA-seq data.
17. The method of any one of claims 13 or 14, wherein the biological sample comprises a liquid biological sample or a solid biological sample.
18. The method of claim 17, wherein the liquid biological sample comprises whole blood, peripheral blood mononuclear cells, cerebrospinal fluid, pleural effusion, urine, or bile.
19. The method of claim 17, wherein the solid biological sample comprises a peritumoral tissue biopsy or a lymph node biopsy.
20. The method of any one of claims 13 or 14, wherein the tumor type or tissue of origin is lung, colorectal, liver, prostate, breast, or melanoma 21. A method of enriching disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) from a biological sample, comprising: a) contacting the sample with a plurality of tetrameric antibodies, wherein each tetrameric antibody comprises a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by a selected cell type; b) applying the sample contacted with the plurality of tetrameric antibodies to a first density gradient medium and performing a first density gradient centrifugation to obtain a first population of enriched DTCs or CTCs; c) resuspending the remaining cells following the first density gradient centrifugation, applying the resuspended cells to a second density gradient medium, and performing a second density gradient centrifugation to obtain a second population of enriched DTCs or CTCs; and d) combining the first and second populations of enriched DTCs or CTCs, thereby enriching the DTCs or CTCs.
22. The method of claim 21, wherein the second portion of the plurality of tetrameric antibodies specifically binds to an antigen expressed by leukocytes, an antigen expressed by platelets, or an antigen expressed by granulocytes.
23. The method of claim 22, wherein the plurality of tetrameric antibodies comprises each of a tetrameric antibody comprising a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by leukocytes, a tetrameric antibody comprising a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by platelets, and a tetrameric antibody comprising a first portion that specifically binds to an antigen expressed by erythrocytes and a second portion that specifically binds to an antigen expressed by granulocytes.
24. The method of claim 22 or claim 23, wherein the antigen expressed by leukocytes is CD45, the antigen expressed by platelets is CD61, and the antigen expressed by granulocytes is CD66b.
25. The method of claim 21, wherein the antigen expressed by erythrocytes is glycophorin A.
26. The method of claim 21, wherein the biological sample comprises a liquid biological sample.
27. The method of claim 26, wherein the liquid biological sample comprises whole blood or peripheral blood mononuclear cells.
28. The method of claim 21, wherein the contacting the sample with the plurality of tetrameric antibodies is carried out in low protein binding tubes or ultra-hydrophilic polymer coated tubes.
29. The method of claim 21, wherein the first and / or second density gradient centrifugation is carried out in ultra-hydrophilic polymer coated centrifuge tubes.
30. The method of claim 21, further comprising diluting the sample contacted with the plurality of tetrameric antibodies prior to applying to the first density gradient medium.
31. The method of claim 21, wherein the first and / or second density gradient centrifugation is at 1200 g for 20 minutes with no braking.
32. The method of claim 21, wherein the biological sample is from a subject with a tumor or suspected of having a tumor.
33. The method of claim 32, wherein the tumor is a solid tumor.
34. The method of claim 33, wherein the solid tumor is a lung tumor, a colorectal tumor, a liver tumor, a prostate tumor, a breast tumor, or a melanoma.
35. A disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) identification system comprising: one or more processors; andmemory coupled to the one or more processors, wherein the memory comprises computer-executable instructions causing the one or more processors to perform a process comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset for a sample; and b) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells.
36. The system of claim 35, wherein identifying the genome-wide CNA signature in one or more individual cells comprises: i) segmenting the scRNA-seq data to a plurality of genomic bins each comprising a pre-defined number of genes from a genome; ii) determining a mean expression intensity for the genes in each bin; iii) determining genome-wide CNA inference from the expression intensity for each bin using Gaussian mixture modeling; iv) determining copy number deviation from a baseline for each bin; and v) identifying a number of outlier bins wherein copy number deviates from the baseline.
37. One or more computer-readable media having encoded thereon computer-executable instructions that, when executed, cause a computing system to perform a disseminated tumor cells (DTCs) or circulating tumor cells (CTCs) identification method comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset for a sample; and b) identifying a genome-wide copy number alteration (CNA) signature in one or more individual cells, wherein the CNA signature distinguishes DTCs or CTCs from normal cells.
38. The computer-readable media of claim 37, wherein identifying the genome-wide CNA signature in one or more individual cells comprises: i) segmenting the scRNA-seq data to a plurality of genomic bins each comprising a pre-defined number of genes from a genome; ii) determining a mean expression intensity for the genes in each bin; iii) determining genome-wide CNA inference from the expression intensity for each bin using Gaussian mixture modeling;iv) determining copy number deviation from a baseline for each bin; and v) identifying a number of outlier bins wherein copy number deviates from the baseline.
39. A tumor type or tissue of origin of a tumor identification system comprising: one or more processors; and memory coupled to the one or more processors, wherein the memory comprises computer-executable instructions causing the one or more processors to perform a process comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset from a sample; b) constructing signature gene sets for a plurality of solid tumor types for tumor type- specific score calculations. c) determining tumor type-specific gene signature scores of the DTCs or CTCs from the scRNA-seq data obtained from the sample; d) identifying the tumor type or tissue of origin based on the recursive logistic regression model predictions using tumor type specific gene signature scores as input.
40. One or more computer-readable media having encoded thereon computer-executable instructions that, when executed, cause a computing system to perform a tumor type or tissue of origin of a tumor identification system method comprising: a) receiving a single cell RNA sequencing (scRNA-seq) dataset from a sample; b) determining a gene signature score of the DTCs or CTCs from the scRNA-seq data obtained from the sample; c) comparing the gene signature score of the DTCs or CTCs in the sample with one or more tumor type-specific signature scores; and d) identifying the tumor type or tissue of origin based on the gene signature score.
41. A method of identifying a compound effective against a drug target, comprising: a) receiving an RNA-seq dataset of disseminated tumor cells (DTCs) or circulating tumor cells (CTCs); b) determining transcriptional states of the DTCs or CTCs based on the RNA-seq dataset; c) computationally assessing the compound sensitivity based on the transcriptional states; andd) identifying the compound as having or not having increased predicted efficacy against the transcriptional states.
42. The method of claim 41, wherein the method comprises in b) sorting the transcription states into two or more clusters, in c) computationally assessing the compound sensitivity based on the transcription states of the two or more clusters, and in d) identifying the compound as having or not having increased predicted efficacy against the transcription state of at least one of the clusters.
43. The method of claim 41 or 42, further comprising selecting a subject with cancer for treatment with the compound, wherein the compound is selected for treating the subject if the compound is determined as having increased predicted efficacy against the transcriptional states, or against the transcription state of at least one of the clusters .
44. The method of claim 43, further comprising administering the selected compound to the subject.
45. The method of claim 41, wherein the determining and computationally assessing comprises applying a drug and drug target identification model comprising a functional module states (FM-States) framework or drug sensitivity prediction algorithm.
46. The method of claim 45, wherein the DTC or CTC RNA-seq data is aggregated single cell RNA-seq data.
47. The method of claim 41, wherein the DTCs or CTCs are identified according to the method of any one of claims 1 to 2.