A method for tracing primary tumors of unknown origin based on ngs and machine learning

CN122551892APending Publication Date: 2026-08-11ZHONGSHAN HOSPITAL FUDAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-10
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]机器学习(ML)算法:面对NGS技术产生的海量、高维基因表达数据,传统生物统计方法已难以有效挖掘其深层规律

Benefits of technology

[0052]本发明提供了一种整合实验与分析的系统性解决方案,相较于现有的原发灶不明肿瘤诊断技术,在预测准确性、临床适用性、技术整合度及诊疗指导价值方面取得了显著提升,具体体现在:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551892A_ABST
    Figure CN122551892A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on next-generation sequencing (NGS) and machine learning technology's primary focus unknown tumor tracing method, belong to biological medicine technical field.The present application provides a kind of based on NGS and machine learning's primary focus unknown tumor tracing method, obtains the expression data of at least 1000 genes by targeted RNA sequencing;After data pre-processing and standardization, key gene features are screened using SVM-RFE algorithm;Based on feature set, use SVM to build multi-classification prediction model, and the model is trained and independently verified;To realize the accurate tracing of tumor primary site.By the present application, a complete technical process from experiment to analysis is established, which provides a systematic solution for clinical auxiliary diagnosis of CUP.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for tracing the origin of tumors with unknown primary foci based on next-generation sequencing (NGS) and machine learning technologies, belonging to the field of biomedical technology. Background Technology

[0002] Cancer of unknown primary (CUP) is a metastatic malignant tumor in which the primary site cannot be clearly determined after evaluation of medical history, physical examination, imaging, serum biomarkers, and histopathology. 1 Currently, CUP accounts for approximately 3%–5% of new cancer cases worldwide. 2 With the development of diagnostic technology, its incidence rate is declining. Identifying the primary site is crucial for developing individualized treatment plans. Although approximately 20% of CUP patients can have their primary site located through diagnostic techniques and receive targeted treatment, 3 Site-specific therapy (SST) is used, but in about 80% of patients the primary tumor remains undiagnosed, requiring empirical chemotherapy (EC). The prognosis for EC is generally poor, with a median overall survival of 2-11 months and a 1-year survival rate of only 15%-20%. 2-6 .

[0003] Next-generation sequencing (NGS) technology: Diagnostic methods for CUP have gradually shifted from traditional histopathology and immunohistochemistry (IHC) to molecular detection. Although IHC is a commonly used primary screening tool, its accuracy in identifying the primary site is limited (≤66%), and it is constrained by issues such as sample depletion, subjective judgment, and overlapping antigen expression. 7-10 Comprehensive genome analysis (CGP) has poor specificity for tumor origin because many cancers share similar mutations. However, it can detect targetable gene variants (such as KRAS) and is primarily used to guide treatment rather than locate the primary tumor. 11-14 In contrast, NGS-based gene expression profiling (GEP) analysis technology demonstrates significant advantages. GEP can detect the mRNA expression of hundreds or even thousands of genes at once and compare them with a database, achieving an accuracy of 80%–90% in the diagnosis of primary site of CUP. 15-16It is no longer limited to individual genes, but can perform unbiased, high-throughput detection of the expression of thousands of genes at once, thereby capturing a more comprehensive range of tumor biological characteristics. Its high sensitivity allows it to be used not only for fresh tissue and exfoliated cells, but also for degraded samples such as FFPE (formalin-fixed paraffin-embedded tissue samples), effectively overcoming the limitations of traditional techniques in terms of sample volume and information content. More importantly, the NGS platform can simultaneously perform gene expression profiling (GEP) analysis and comprehensive genome sequencing (CGP), addressing both the localization of "tumor origin" and the guidance of "how to treat" in a single test, maximizing diagnostic efficiency.

[0004] Machine Learning (ML) Algorithms: Faced with the massive, high-dimensional gene expression data generated by NGS technology, traditional biostatistical methods have become increasingly ineffective in uncovering deep patterns. Machine learning algorithms, especially classification models such as Support Vector Machines (SVM) and Random Forests (RF), have demonstrated immense value in this field. 17-18 These algorithms can learn through training to automatically identify the most discriminative feature combinations from thousands of genes, constructing high-precision classification models to accurately predict the primary site of tumors. Their advantage lies in capturing the complex nonlinear relationships between genes and performing dimensionality reduction and noise reduction on the data, ultimately achieving high-sensitivity source tracing accuracy.

[0005] References

[0006] 1.Greco FA, Hainsworth JD: Cancer of unknown primary site, inDeVitaJr VT, Lawrence TS, Rosenberg S (eds): Cancer: Principles and practice of oncology. Philadelphia, PA, Wolters Kluwer, 2015, pp1720-1737

[0007] 2. Rassy E, Pavlidis N: The currently declining incidence of cancer of unknown primary. Cancer Epidemiol 2019, 61:139-144

[0008] 3.Hainsworth JD, Fizazi K: Treatment for patients with unknownprimary cancer and favorable prognostic factors. Semin Oncol 2009, 36:44-51

[0009] 4.Greco FA, Pavlidis N: Treatment for patients with unknown primarycarcinoma and unfavorable prognostic factors. Semin Oncol 2009, 36:65-74

[0010] 5.Moran S, Martinez-Cardús A, Boussios S, Esteller M: Precisionmedicine based on epigenomics: the paradigm of carcinoma of unknown primary.Nat Rev Clin Oncol 2017, 14:682e694

[0011] 6.Liu X, Zhang X, Jiang S, et al: Site-specific therapy guided by the90-gene expression assay versus empiric chemotherapy in patients with cancerof unknown primary. Lancet Oncol 2024, 25:1092-1098

[0012] 7.Greco FA, Lennington WJ, Spigel DR, et al: Molecular profilingdiagnosis in unknown primary cancer: Accuracy and ability to complementstandard pathology. J Natl Cancer Inst 2013, 105:782-790

[0013] 8. Anderson GG, Weiss LM: Determining tissue of origin for metastaticcancers: meta-analysis and literature review of immunohistochemistryperformance. Appl Immunohistochem Mol Morphol 2010, 18:3e8

[0014] 9.National Comprehensive Cancer Network: NCCN Clinical PracticeGuidelines in Oncology (NCCN Guidelines) Occult Primary Cancer of UnknownPrimary (CUP) V.2.2024. https: / / NCCN.org

[0015] 10. Selves J, Long-Mira E, Mathieu MC, et al: Immunohistochemistryfor diagnosis of metastatic carcinomas of unknown primary site. Cancers(Basel) 2018, 10:108-134

[0016] 11.Ross JAS, Wang K, Gay LJL, et al: Comprehensive genomic profilingof carcinoma of unknown primary site: New routes to targeted therapies. JAMAOncol 2015, 1:40-47

[0017] 12.Kato S, Krishnamurthy N, Banks KC, et al: Utility of genomicanalysis in circulating tumor DNA from patients with carcinoma of unknownprimary. Cancer Res 2017, 77:4238-4246

[0018] 13.Greco FA, Lennington WJ, Spigel DR, et al: Molecular profilingdiagnosis in unknown primary cancer: Accuracy and ability to complementstandard pathology. J Natl Cancer Inst 2013, 105:782-790

[0019] 14.Varghese AM, Arora A, Capanu M, et al: Clinical and molecularcharacterization of patients with cancer of unknown primary in the modernera. Ann Oncol 2017, 28:3015-3021

[0020] 15.Meiri E, Mueller WC, Rosenwald S, et al: A second-generationmicro-RNA based assay for diagnosing tumor tissue origin. Oncologist 2012,17:801-818,

[0021] 16.Pentheroudakis G, Pavlidis N, Foumtzilas G, et al: Novel microRNA-based array demonstrates 92% agreement with diagnosis based onclinicopathologic and management data in a cohort of patients with carcinomaof unknown primary. Mol Cancer 2013, (12):57-65

[0022] 17.Moran S, Martinez-Cardus A, Sayols S, et al: Epigenetic profiling to classify cancer of unknown primary: A multicenter, retrospective analysis. Lancet Oncol 2016, 17:1386-1394

[0023] 18.Sun W, Wu W, Wang Q, et al: Clinical validation of a 90-geneexpression test for tumor tissue of origin diagnosis; a large-scalemulticenter study of 1417 patients. J Transla Med 2022, 20:114-124 Summary of the Invention

[0024] The purpose of this invention is to solve the technical problem of tracing the origin of tumors with unknown primary lesions. This invention provides a method and product for tracing the origin of tumors with unknown primary lesions based on multi-gene expression profiling analysis. It utilizes machine learning technology to model and analyze the expression data of 2660 tumor-related genes to achieve high-precision prediction of the origin of the primary tissue, which can be used to assist in the diagnosis and individualized treatment decision-making of CUP patients.

[0025] To achieve the objectives of this invention, this invention provides a method for tracing the origin of tumors with unknown primary foci based on NGS and machine learning. The method involves obtaining expression data covering at least 1,000 genes through targeted RNA sequencing; after data preprocessing and standardization, key gene features are screened using the SVM-RFE algorithm; based on the feature set, a multi-classification prediction model is constructed using SVM, and the model is trained and independently validated.

[0026] Preferably, the method includes the following steps:

[0027] Step 1: Establish the sample queue and dataset; including two independent queues, namely the discovery queue for model training and the independent validation queue for model performance evaluation;

[0028] Step 2, Sample Library Construction and Sequencing: Collect FFPE tumor samples covering multiple cancer types, extract RNA and construct strand-specific sequencing libraries, capture 2660 target genes using a target probe panel, and perform paired-end sequencing on a high-throughput sequencing platform; RNA sequencing is performed on the discovery cohort and the independent validation cohort, and DNA sequencing is performed on the independent validation cohort.

[0029] Step 3: Construction of the CUP classification model;

[0030] Step 3.1, Data Preprocessing and Feature Generation: Quality control of sequencing data is performed, alignment to a reference genome is performed, and the TPM value of each gene is calculated; further, the TPM is normalized between samples using internal reference genes, and log2 transformation is performed to reduce technical bias; Gene set variation analysis (GSVA) ​​is performed on gene expression profiles to obtain the score of each functional gene set corresponding to the sample; finally, the standardized TPM of each gene and the score of each gene set are used as sample features.

[0031] Step 3.2, Feature Selection and Model Training: From the standardized expression data, the SVM-RFE method is used to select 250 gene feature subsets with the most discriminative power; based on this feature subset, a multi-classification model is trained using linear kernel SVM. During the training process, the SMOTE algorithm is introduced to handle class imbalance, and the model parameters are optimized through cross-validation.

[0032] Step 4, Model Validation and Application: Input the preprocessed independent validation set sample expression data into the trained model, output the predicted category of the primary cancer type, and evaluate the accuracy and generalization ability of the model by comparing it with the true label.

[0033] Preferably, all analyses in step 3.1 above are performed using Python and R; the gene expression data of the discovery cohort samples that have passed quality control are randomly divided, with the ratio of training set to test set being 7:3.

[0034] Preferably, the data preprocessing in step 3.1 above includes selecting internal reference genes with a coefficient of variation (CV) lower than 0.35 from the training set; normalizing the TPM value of each sample by dividing it by the average TPM value of the selected internal reference genes; and performing a log2 transformation on the normalized TPM value.

[0035] Preferably, in step 3.2 above, the feature selection and model training include: using an SVM with a linear kernel function; standardizing and scaling the features; using the SMOTE algorithm to solve the class imbalance problem; optimizing hyperparameters on the training set through grid search combined with 5-fold cross-validation; finally, the model is trained on all training set samples and saved for evaluation on the test set and independent validation queue; the evaluation metrics include accuracy, sensitivity, specificity, positive prediction value, and negative prediction value.

[0036] Preferably, the independent validation cohort in step 1 above also includes a Chinese genome dataset cBioPortal, ID:pan_origimed_2020 cohort; this cohort contains 11,553 Chinese patients, covering 25 different cancer types, and uses a CYCS NGS panel targeting 450 cancer-related genes for sequencing analysis.

[0037] Preferably, the model validation in step 4 includes the integration of DNA mutation information with the CUP classification model; it includes using the Chinese genome dataset to identify gene sets with significantly different mutation frequencies for each pair of tumor types; and then, in an independent validation cohort, integrating the DNA mutation data of the top three predicted tumor types for each sample to optimize the classification.

[0038] Preferably, the integration of the DNA mutation information with the CUP classification model includes calculating, for a sample, its adjusted probability for tumor type i according to the following formula:

[0039]

[0040] and

[0041]

[0042] in:

[0043] The adjusted probability of the i-th tumor type;

[0044] : Weight score of the i-th tumor type;

[0045] The probability of the i-th tumor type given by the CUP model;

[0046] The number of genes with significantly different mutation frequencies for each tumor type pair;

[0047] The frequency of the g-th differentially expressed gene in tumor type i relative to tumor type j;

[0048] Based on the probability value of each sample after integrating DNA mutation information, the tumor type with the highest probability value is selected as the prediction result for that sample.

[0049] This invention provides an application of the above-described method for tracing the origin of tumors of unknown primary location based on NGS and machine learning, the application including the preparation of a diagnostic kit for tracing the origin of tumors of unknown primary location.

[0050] This invention provides a source tracing system for tumors of unknown primary origin. The system includes a source tracing detection program that applies the aforementioned method for tracing tumors of unknown primary origin based on NGS and machine learning.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] This invention provides a systematic solution integrating experimentation and analysis, which significantly improves upon existing diagnostic techniques for tumors of unknown primary origin in terms of predictive accuracy, clinical applicability, technological integration, and diagnostic and treatment guidance value. Specifically, it includes:

[0053] 1. High accuracy and excellent predictive performance

[0054] By employing a machine learning process combining recursive feature elimination with support vector machines (SVMs), this invention can screen a small number of highly discriminative features (i.e., 250 key genes / gene sets) from high-dimensional gene expression data and construct an efficient multi-classification model. This model achieves an overall accuracy of no less than 85% on an independent validation set, with a prediction accuracy exceeding 80% for common cancers such as breast cancer, lung cancer, and colorectal cancer. This significantly outperforms traditional immunohistochemical methods that rely on single antibodies or limited biomarkers (whose accuracy is approximately 50-65%), providing more reliable diagnostic evidence for clinical practice.

[0055] 2. Strong clinical applicability

[0056] The classification model established in this invention can predict the tissue origin of most common solid tumors, solving the clinical challenge of tracing the origin of tumors of unknown primary origin, which are heterogeneous tumors. In particular, this method exhibits excellent compatibility with FFPE tissue clinical samples, requiring only nanograms of RNA for high-quality sequencing, greatly improving its feasibility in routine pathological testing. The superior performance of this invention on multiple validation sets demonstrates the robustness of the method and shows its potential to directly guide clinical practice.

[0057] 3. Innovation and efficiency of the technical process

[0058] This invention proposes a targeted RNA sequencing-based approach that can simultaneously acquire gene expression profile data for source tracing in a single test. Compared to whole transcriptome sequencing, this approach focuses on 2660 tumor-related genes, effectively reducing testing costs and data analysis complexity while ensuring information abundance and relevance. Combined with standardized data preprocessing and the class imbalance strategy handled by the SMOTE algorithm, a stable, reproducible, and computationally efficient technical workflow is formed, laying the foundation for large-scale clinical applications.

[0059] 4. Direct clinical benefits and treatment guidance value

[0060] By precisely tracing the primary site of the tumor, this invention enables patients to receive more precise "site-specific treatment." Clinical studies have demonstrated that, compared to traditional empirical chemotherapy, "site-specific treatment" based on molecular tracing results can significantly prolong the median progression-free survival of patients (from 6.6 months to 9.6 months). This indicates that this invention is not only a diagnostic tool, but can also be directly transformed into a treatment decision support system to improve patient prognosis, thus advancing the precision medicine process for tumors of unknown primary origin. Attached Figure Description

[0061] Figure 1 A comprehensive assessment of feature selection and its biological significance diagram;

[0062] Figure 2 The classification diagrams for primary tumors (Figure A) and metastatic tumors (Figure B) in the internal validation set of the CUP model, and the confusion matrix diagram in the external independent cohort (Figure C);

[0063] Figure 3 Integrating DNA mutation information and a GEP-based CUP model can improve the accuracy of tracing the origin of ambiguous samples;

[0064] Figure 4 This is a flowchart of the method for tracing the origin of tumors of unknown primary foci based on NGS and machine learning according to the present invention.

[0065] Figure 5 This is the imaging and pathological manifestation of the first discovery of a space-occupying lesion in the patient of Example 2;

[0066] Figure 6 Genetic testing of the patient in Example 2 indicated that the cervical lymph node metastatic squamous cell carcinoma originated from the lungs;

[0067] Figure 7 In Example 2, the detection and calculation methods provided by the present invention indicated that the cervical lymph node metastatic squamous cell carcinoma in the patient originated from the lung.

[0068] Figure 8The imaging and pathological manifestations of the lung lesion in the patient in Example 2 after chemotherapy;

[0069] Figure 9 The imaging and pathological features of the neck lesion in the patient in Example 2 after chemotherapy;

[0070] Figure 10 This is a PET / CT image of the patient in Example 3 who was the first to have a liver lesion detected.

[0071] Figure 11 The pathological features of the mass in the left lateral lobe of the liver in patient Example 3;

[0072] Figure 12 In Example 3, the detection and calculation methods provided by the present invention indicated that the mass in the left lateral lobe of the liver was a malignant mesothelioma. Detailed Implementation

[0073] To make the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings:

[0074] Example 1

[0075] I. Materials and Methods

[0076] 1. Sample queue and dataset

[0077] This invention relates to two separate cohorts for developing and validating a gene expression profile (GEP)-based classification model for primary unexplained metastatic cancer (CUP).

[0078] Discovery Cohort (for Model Training): This cohort collected 1469 histologically confirmed formalin-fixed paraffin-embedded (FFPE) tumor tissue samples from Zhongshan Hospital Affiliated to Fudan University (FDUZH) that fell within the 36 tumor types predicted by the model. Samples were obtained through surgical resection or biopsy between January 2020 and December 2024, and the tumor content was confirmed to be no less than 60% by hematoxylin and eosin (H&E) staining. The RNA portion of this cohort samples was sequenced using the AmoyDx® Master Panel (covering 571 DNA mutation genes and 2660 RNA expression genes) to generate GEP data for model construction.

[0079] Independent Validation Cohort (for Model Performance Evaluation): A retrospective clinical FFPE sample collection from FDUZH for routine molecular diagnostics during the same period was used as an independent validation cohort. This cohort represents a real-world clinical scenario and underwent simultaneous DNA and RNA sequencing using the AmoyDx® Master Panel to assess model performance and explore the added value of DNA mutation data. Sample acquisition for both cohorts strictly adhered to the institution's informed consent policy, and the study protocol was approved by the FDUZH Institutional Review Board.

[0080] To further enhance the RNA-based CUP classification model, this invention obtained differentially expressed genetic information across cancer types from a publicly available Chinese genome dataset (cBioPortal, ID: pan_origimed_2020). This cohort includes 11,553 Chinese patients covering 25 different cancer types, and sequencing analysis was performed using CYCS NGSpanel targeting 450 cancer-related genes.

[0081] 2. Sample preparation and targeted sequencing

[0082] 2.1 RNA sequencing (discovery cohort and independent validation cohort)

[0083] Total RNA was extracted from FFPE tumor samples. RNA concentration was quantified using a Quantus™ fluorometer, and RNA fragment size and integrity were assessed using an Agilent 2100 Bioanalyzer. RNA samples were fragmented at 95°C for 0–15 minutes based on DV200 values. Subsequently, reverse transcription, cDNA synthesis, and strand-specific library construction were performed using the NEBNext® Ultra™ II Directional RNA Library Prep Kit for Illumina®. The libraries were hybridized and captured using the AmoyDx® MasterPanel (targeting 2660 genes for RNA expression analysis). After amplification, quantification, and fragment size assessment, the captured libraries were sequenced at 150 bp paired ends on an Illumina NovaSeq 6000 system.

[0084] 2.2 DNA sequencing (independent validation cohort)

[0085] DNA concentration was quantified using a Quantus™ fluorometer, and DNA fragment size distribution was assessed using an Agilent 2100 Bioanalyzer. Genomic DNA was fragmented to 200–250 bp using the Covaris LE220 system. NGS libraries were constructed using the NEBNext® Ultra™ II DNA Library Prep Kit and then targeted enriched using the AmoyDx® Master Panel (covering 571 genes for comprehensive analysis of SNVs, Indels, fusions, CNVs, MSIs, and TMBs). The captured DNA libraries were then processed and sequenced at 150 bp paired ends on an Illumina NovaSeq 6000 platform at a sequencing depth of 500x.

[0086] 3. Construction method of CUP classification model

[0087] 3.1 Data Preprocessing and Feature Generation

[0088] All analyses were performed using Python and R. Gene expression data from 1469 discovery cohort samples that passed quality control were randomly split, with a training set to test set ratio of 7:3.

[0089] Specifically, the gene expression data are standardized as follows:

[0090] • Select housekeeping genes from the training set with a coefficient variation (CV) of less than 0.35.

[0091] • Normalize the TPM value of each sample by dividing it by the average TPM value of the selected internal reference genes;

[0092] • Perform a log2 transformation on the normalized TPM value (add 0.001 before transformation to prevent undefined values).

[0093] In addition, gene set variation analysis (GSVA) ​​was performed on the gene expression profiles to obtain scores for each functional gene set corresponding to the sample. Finally, the TPM of each gene and the scores of each gene set were used as sample features after standardization.

[0094] 3.2 Feature Selection and Model Training

[0095] Feature selection was performed using Support Vector Machine-Recursive Feature Elimination (SVM-RFE), iteratively removing features to retain the 250 most informative genes for CUP model construction.

[0096] Specifically, the following steps are used to train the Support Vector Machine (SVM) model:

[0097] • SVM using a linear kernel function;

[0098] • Standardize and scale the features;

[0099] • Use the SMOTE algorithm to solve the class imbalance problem;

[0100] • Optimize hyperparameters on the training set by combining grid search with 5-fold cross-validation.

[0101] The final model was trained on all training set samples and saved for evaluation on the test set and independent validation queue. Evaluation metrics included accuracy, sensitivity, specificity, positive predictive value, and negative predictive value.

[0102] 4. Integration of DNA mutation information with the CUP classification model

[0103] To evaluate whether integrated genomic analysis at the DNA level can further improve the classification accuracy of the GEP model, this invention employs the following multi-step method:

[0104] First, using the aforementioned Chinese pan-cancer dataset, gene sets with significantly different mutation frequencies were identified for each pair of tumor types (p < 0.05). Then, in an independent validation cohort, DNA mutation data for the top three predicted tumor types of each sample were integrated to optimize classification. Specifically, for a sample, its adjusted probability for tumor type i was calculated using the following formula:

[0105]

[0106] and

[0107]

[0108] in:

[0109] The adjusted probability of the i-th tumor type;

[0110] : Weight score of the i-th tumor type;

[0111] The probability of the i-th tumor type given by the CUP model;

[0112] The number of genes with significantly different mutation frequencies for each tumor type pair;

[0113] : The frequency of the g-th differentially expressed gene in tumor type i (relative to tumor type j).

[0114] Based on the probability value of each sample after integrating DNA mutation information, the tumor type with the highest probability value is selected as the prediction result for that sample.

[0115] result:

[0116] 1. Feature Filtering:

[0117] The discovery cohort used for model training, after experimental and bioinformatics quality control, comprised 1469 clinical samples covering primary and metastatic tumors. Gene expression data and functional gene set scores from the training set (n=1025) were used for feature selection and model training. Feature evaluation determined that 250 features were sufficient to achieve optimal classification accuracy (e.g., ...). Figure 1 A). A support vector machine recursive feature elimination method was used to iteratively remove features, ultimately retaining 250 of the most relevant features for CUP tumor type classification. Gene set enrichment analysis showed that the selected genes were mainly enriched in cancer-related pathways, cytokine-cytokine receptor interactions, abnormal transcriptional regulation in cancer, and Ras signaling pathways (such as...). Figure 1 B). Furthermore, analysis of the higher biological functions associated with enrichment pathways revealed that these pathways are involved in fundamental processes of cellular life (B). Figure 1 C), including development, cell differentiation, proliferation, and metabolism.

[0118] like Figure 1 As shown, Figure 1 A comprehensive evaluation of feature selection and its biological significance is presented. Figure A shows the change in accuracy of the CUP classification model with the number of input features. The solid line represents the average performance, and the shaded area represents the 95% confidence interval. Figure B shows the enriched KEGG pathways identified from the selected feature genes, sorted by their comprehensive scores. Figure C analyzes the high-level biological functions of the enriched pathways.

[0119] 2. Model performance:

[0120] The model's performance was further validated in the internal validation set. The CUP classification model demonstrated robust performance in an internal validation cohort containing 444 samples (362 primary tumors and 82 metastatic tumors). The overall accuracy reached 92.8%, with a precision of 94.2% for primary tumors. Figure 2 A) Significantly higher than that of metastatic tumor samples (86.6%, such as...) Figure 2 B).

[0121] In the primary tumor subgroup ( Figure 2(A) Of the 36 tumor types, 22 achieved 100% perfect classification, including breast cancer, cervical squamous cell carcinoma, colorectal cancer, esophageal squamous cell carcinoma, hepatocellular carcinoma, lung adenocarcinoma, and prostate adenocarcinoma (see Table 1 for details). These types collectively accounted for 59.12% (214 / 362) of the primary samples. However, classification of certain specific tumor types remained challenging: the accuracy was 76.9% (10 / 13) for bile duct carcinoma, 73.3% (11 / 15) for pancreatic adenocarcinoma, and 83.3% (5 / 6) for cervical adenocarcinoma, duodenal adenocarcinoma, thymic squamous cell carcinoma, and pituitary adenocarcinoma.

[0122] In an external independent validation set containing 249 tumor samples, the model achieved an overall CUP classification accuracy of 87.9%, with accuracy exceeding 80% for 18 of the 23 tumor types (see [link to validation dataset]). Figure 2 C). All samples of adrenocortical carcinoma (ACC), duodenal adenocarcinoma (DAC), intrahepatic cholangiocarcinoma (ICC), squamous cell carcinoma of the lung (LUSC), adenocarcinoma of the lung (LUAD), melanoma, prostate cancer (PRAD), small cell lung cancer (SCLC), and thyroid cancer were correctly identified. Misclassifications exhibited similar patterns to those in the primary tumor subset, primarily including subtype confusion within the same organ (e.g., 1 out of 10 hepatocellular carcinoma (LICA) cases was misclassified as intrahepatic cholangiocarcinoma (ICC), and 1 as gallbladder adenocarcinoma (GC)), and misclassifications of similar tissue origin (e.g., 1 out of 19 pancreatic cancer (PAAD) cases was misclassified as duodenal adenocarcinoma (DAC), 1 as gastric cancer (GAS), and 1 as hepatocellular carcinoma (LICA)).

[0123] like Figure 2 As shown, Figure 2 The CUP model is classified into primary tumors (Figure A) and metastatic tumors (Figure B) in the internal validation set, and the confusion matrix is ​​shown in the external independent cohort (Figure C). The numerical values ​​in the matrix represent the proportion of tumor types predicted by the model, and the values ​​on the diagonal represent the prediction accuracy for each tumor type.

[0124] Table 1. Sensitivity, specificity, positive predictive value, and negative predictive value of the CUP model on the internal validation set for each cancer type.

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138] 3. Integrating DNA mutation information to improve traceability performance

[0139] By comparing the performance of the GEP-based CUP model in this study with that in previous studies, we found that regardless of the method used, the classification accuracy for certain tumor types remained relatively low, including urinary tract cancer, pancreatic cancer, melanoma, endometrial cancer, and head and neck cancers (such as...). Figure 3 A). Furthermore, analysis of the model prediction probability values ​​in the independent validation queue shows that when the model makes an incorrect prediction, the probability difference between the first two predicted labels is greater than the difference when the prediction is correct (e.g., ...). Figure 3 B). These findings suggest that the classification accuracy for specific tumor types remains consistently low across different platforms, with small probability differences when correctly predicted, potentially reflecting inherent limitations of the CUP classification method based on expression data.

[0140] Therefore, we introduced DNA mutation information into an external validation cohort to assess its potential to improve classification performance for difficult-to-distinguish samples. We first screened for differentially expressed genes in a publicly available Asian population cohort by comparing tumor type pairs in the external cohort; then, we optimized the CUP model's predictions using DNA mutation information. A total of 350 genes differing between different tumor type pairs were identified, including both common and specific differentially expressed genes, ranging in number from 1 to 294. By integrating DNA mutation information into the CUP model predictions, 4 out of 34 samples misclassified by CUP were correctly reclassified, a rate of 11.8%. Figure 3 C). After introducing DNA mutation information, no cases occurred where samples correctly predicted by CUP were misclassified.

[0141] like Figure 3The diagram illustrates how integrating DNA mutation information with the GEP-based CUP model can improve the accuracy of source tracing for ambiguous samples. Figure A compares the prediction accuracy of different gene expression models across various tumor types. Figure B shows that the difference in the Top 2 prediction probabilities is smaller for misclassified samples compared to correctly classified samples. Figure C shows the prediction results after integrating DNA information for samples misclassified by the CUP model.

[0142] Example 2

[0143] CUP Diagnosis and Treatment Guided by Source Tracing Model: A Successful Case of Neck Metastatic Squamous Cell Lung Carcinoma with Metachronic Second Primary Lung Adenocarcinoma

[0144] The patient, a 68-year-old male, accidentally palpated a mass behind his right ear in late 2023, about the size of a pigeon egg. He did not pay much attention to it at the time, but it later progressively enlarged, becoming firm with relatively clear borders. The patient reported mild pain and tenderness, and a slight foreign body sensation when opening his mouth, without local skin redness, swelling, or ulceration. He then visited our outpatient clinic. On January 4, 2024, electronic laryngoscopy showed no obvious mass in the larynx, nasal cavity, or sinuses. On January 5, 2024, PET / CT images showed a mass with abnormally high glucose metabolism in the right upper neck, lobulated margins, approximately 3.12 × 2.56 cm in size, with an average CT value of approximately 40.0 HU and a maximum SUV value of approximately 28.5. Possible diagnoses include multiple metastatic lesions (MT) in the right neck; a few chronic inflammations and old lesions in the right upper lobe of the lung; a thyroid nodule in the right lobe (must be examined with ultrasound); liver and bilateral renal cysts; and bilateral renal stones. On January 9, 2024, a neck CT scan showed a nodular soft tissue density lesion, approximately 3.2 × 2.3 cm in size, posterior to the right submandibular gland. The lesion had relatively clear borders and showed heterogeneous enhancement after contrast administration. Multiple slightly larger lymph nodes were observed adjacent to the lesion. Right cervical MT was suspected, with slight localized thickening of the soft tissue on the right oropharyngeal wall. Endoscopic examination was recommended. An ultrasound-guided biopsy was performed. Pathology revealed poorly differentiated carcinoma (right cervical lesion) with focal squamous epithelial differentiation, indicating non-keratinizing squamous cell carcinoma. No normal tissue structure was observed. Morphology and immunohistochemical staining could not determine the tissue origin; further examination was recommended. Immunohistochemical staining results: Tumor cells positively expressed P40, P63, and GATA3, but not TTF1 or Napsin A. The Ki67 positivity index was approximately 40%. PD-1 was present (less than 1% of lymphocytes in the stroma were positive). PDL1 clone number 22C3 (TPS=6, CPS=10) (e.g., ...). Figure 5 ). CUP tissue tracing testing was then performed (genetic testing by Kebang Gene (Hangzhou Kebang Gene Technology Co., Ltd.) indicated lung origin, and this model also indicated lung origin). Figure 6 , Figure 7 ).

[0145] like Figure 5The images show the imaging and pathological findings of a newly discovered cervical mass. Image A is a PET / CT image showing a mass with abnormally high glucose metabolism in the right upper neck, lobulated margins, and a size of approximately 3.12 × 2.56 cm. Image B is the postoperative pathology report showing the mass in the right upper neck as poorly differentiated carcinoma with focal squamous epithelial differentiation, indicating non-keratinizing squamous cell carcinoma. HE staining is shown. Image C shows the immunohistochemical results, indicating tumor cells expressing P40. Image D shows the immunohistochemical results, indicating tumor cells expressing P63.

[0146] like Figure 6 As shown, genetic testing indicates that the cervical lymph node metastatic squamous cell carcinoma originated from the lungs.

[0147] like Figure 7 As shown, the detection and calculation methods provided by the technical solution of this invention indicate that the metastatic squamous cell carcinoma of the cervical lymph nodes originates from the lungs.

[0148] Following a multidisciplinary clinical discussion, the patient underwent C1D1 chemotherapy (albumin-bound paclitaxel 200mg + cisplatin 100mg) on ​​January 12, 2024. On January 20, 2024, the patient received C1D8 chemotherapy (albumin-bound paclitaxel 200mg). On February 2, 2024, the patient received C2D1 chemotherapy (albumin-bound paclitaxel 200mg + cisplatin 100mg), during which albumin levels were normal. On February 10, 2024, the patient received C2D8 chemotherapy (albumin-bound paclitaxel 200mg). During chemotherapy, occasional leukopenia was observed, which was treated with white blood cell boosting injections, after which the white blood cell count returned to normal. On February 27, 2024, the patient received C3D1 chemotherapy (albumin-bound paclitaxel 200mg + cisplatin 100mg). On March 5, 2024, the patient received C3D8 chemotherapy (albumin-bound paclitaxel 200mg). On March 19, 2024, the patient underwent chemotherapy with the C4D1 regimen (200 mg of albumin-bound paclitaxel + 100 mg of cisplatin). On March 26, 2024, the patient underwent chemotherapy with the C4D8 regimen (200 mg of albumin-bound paclitaxel). On April 18, 2024, a PET / CT scan showed that the mass in the right upper neck with abnormally high glucose metabolism had significantly decreased in size compared to the previous scan, measuring approximately 1.47 × 1.37 cm, with an average CT value of approximately 50.0 HU and a maximum SUV value of approximately 12.4. The boundary with adjacent muscles was not clearly defined. A mixed ground-glass nodule with slightly increased glucose metabolism was seen in the apical segment of the right upper lobe, measuring 1.00 × 0.8 cm. Compared with the PET / CT image taken at this hospital on January 5, 2024, the lesions in the right neck showed varying degrees of reduction in size and decreased glucose metabolism; a peripheral metastatic tumor (MT) in the apical segment of the right upper lobe was possible. On April 28, 2024, a CT-guided biopsy of the lesion in the right upper lung was performed. Pathology revealed lung adenocarcinoma, predominantly of the acinar type. Immunohistochemical staining results: tumor cells positively expressed TTF1 and Napsin A, but did not express ALK (D5F3), P40, P63, ROS1, C-met (80%++), Her2 (40%+), Ki67 positivity index approximately 10%, PD-1 (1% interstitial lymphocytes positive), PDL1 clone number 22C3 (TPS=0, CPS=1) Figure 8 , Figure 9 This lung adenocarcinoma is considered to be a second primary malignant tumor.

[0149] On April 30, 2024, a functional neck lymph node dissection was performed. Postoperative pathology revealed 12 lymph nodes (right neck), 8 of which showed squamous cell carcinoma metastasis. Two additional cancerous nodules with extranodal invasion were also detected. Immunohistochemical staining results showed that tumor cells expressed P40 but did not express TTF1, Napsin A, or ALK (D5F3), and the Ki67 positivity rate was approximately 20%. Figure 9The patient underwent radiotherapy on May 14, 2024, and continued until July 9, 2024. The patient had a history of diabetes for over ten years, regularly taking metformin once daily (qd) with good control. They also had a history of hypertension for several years, regularly taking amphetamine to lower blood pressure, with good control. They had a history of pulmonary tuberculosis 40 years prior. Outpatient follow-up continued until June 10, 2025, with regular checkups and no recurrence / metastasis.

[0150] like Figure 8 The images show the imaging and pathological findings of a lung lesion discovered for the first time after chemotherapy. Image A, a PET / CT image, shows a mixed ground-glass nodule in the apical segment of the right upper lobe, with slightly increased glucose metabolism, measuring 1.00 × 0.8 cm. Image B shows a biopsy of the lesion in the right upper lobe; pathology revealed lung adenocarcinoma, predominantly acinar type. HE staining was performed. Image C shows immunohistochemical results indicating Napsin A expression in the tumor cells. Image D shows immunohistochemical results indicating TTF1 expression in the tumor cells.

[0151] like Figure 9 The images show the imaging and pathological findings of the cervical lesion after chemotherapy. Image A is a PET / CT scan, showing that the right-sided cervical lesion has shrunk to varying degrees compared to before, and glucose metabolism is reduced (approximately 1.47 × 1.37 cm in size). Image B shows a functional neck lymph node dissection, with postoperative pathology indicating squamous cell carcinoma with lymph node metastasis.

[0152] In this case, no primary squamous cell carcinoma was found in the lung on whole-body imaging at the initial diagnosis, which may be due to occult primary tumor. We speculate that the reasons may be: (1) the primary lesion is small in size or grows in an occult manner, below the resolution detection limit of PET / CT; (2) a more reasonable explanation is that the subsequent systemic chemotherapy for squamous cell carcinoma of the lung, while effectively treating the neck metastases, may have also cleared any small primary lesions in the lung that were not identified by imaging. Secondly, a solitary ground-glass nodule was subsequently discovered in the upper lobe of the right lung, which was confirmed by biopsy to be lung adenocarcinoma. Its histological morphology and immunophenotype (TTF1+ / NapsinA+ / P40-) were completely different from the poorly differentiated squamous cell carcinoma of the neck (P40+ / TTF1-) at the initial diagnosis. Therefore, this lung adenocarcinoma should be diagnosed as a second primary malignant tumor, rather than the primary lesion in the lung of the initially diagnosed neck metastases. This is not uncommon in elderly male patients with a long history of smoking.

[0153] Therefore, the development of a second primary lung adenocarcinoma during treatment in this case demonstrates that the source tracing model based on multi-gene expression profiling can accurately identify the dominant primary lesion leading to the current metastatic disease through complex clinical presentations. This also provides a more complex application scenario for evaluating the value of source tracing models. Firstly, at initial diagnosis, in the absence of a clear primary lesion in the lungs, this model, consistent with independent commercial reagent kits, traced the poorly differentiated squamous cell carcinoma metastasis in the neck to the lungs. This prediction was subsequently validated by the significant efficacy of chemotherapy regimens targeting lung squamous cell carcinoma, confirming the accuracy of this source tracing result. The subsequent newly developed lung adenocarcinoma, histologically distinct from the neck metastasis, should be diagnosed as an independent second primary cancer. This complex situation precisely highlights the clinical practicality of this model: in elderly patients with the potential for multiple primary cancers, this model can accurately pinpoint the primary lesion causing the current metastatic tumor for precise treatment, thereby avoiding confusion caused by the discovery of new lesions.

[0154] Example 3

[0155] The patient, a 67-year-old male, presented with a liver lesion discovered during a routine check-up at another hospital in February 2025. MRI suggested possible angiosarcoma or hepatocellular carcinoma. He then sought further diagnosis and treatment at our hospital. On February 26, 2025, a PET / CT scan at our hospital revealed a large, slightly low-density mass with heterogeneous elevated glucose metabolism protruding into the abdominal cavity from the left lateral lobe of the liver. The mass measured approximately 11.4 × 6.44 cm, with an average CT value of approximately 33.0 HU and a maximum SUV value of approximately 14.5. The diagnosis was considered to be a large hepatic metastatic lesion (MT) with daughter lesions in the left lobe, peritoneal seeding metastasis, and right scapular metastasis with soft tissue mass formation. On the same day, tests at our hospital showed CA125: 795 U / mL, Cyfra211: 3.9 ng / mL, SCC: 6.2 ng / mL, while AFP, CEA, CA199, CA153, and PSA were normal. On March 5, 2025, a liver biopsy at our hospital suggested adenocarcinoma of bile duct origin. Immunohistochemical staining results: Tumor cells positively expressed CK19 and Galectin-3, with a small amount of CD56 expression, and did not express TTF1, BRAF-V600E, HBME-1, TPO, PAX8, TG, Calcitonin, SYN, CgA, CDX2, SATB2, HNF-1β, ARG-1, AR, NKX3.1, SALL4, and the Ki67 positivity index was approximately 5%. On March 12, 2025, the patient underwent GEMOX + lenvatinib + PD-1 therapy at our hospital for 6C. A bone scan on July 15, 2025, showed a slight reduction in the size of the metastatic lesion in the right scapula, a slight increase in bone density, and an increased degree of bone mineral metabolism. On the same day, MRI results showed that after MT treatment in the left lobe of the liver, most of the tumor remained viable, roughly the same size as before, with a tumor thrombus in the left branch of the portal vein. On July 31, 2025, a laparoscopic exploration was performed at our hospital. The pathological results showed that there were multiple nodules (liver nodules), (abdominal wall nodules), and (diaphragmatic roof nodules) under the microscope. The tumor tissue showed glandular structure and some papillary structure, with an "adenocarcinoma"-like morphology. Combined with the immunohistochemical results, it was suspected to be an epithelioid malignant mesothelioma (tubular papillary type). Immunohistochemical staining results: Tumor cells positively expressed Mesothelin, MTAP, WT-1, D2-40, CK5 / 6, N-cad, Muc-1, CK19, CK7, MLH1, and MSH2, partially expressed CD56, and did not express HNF-1β, CRP, TTF1, S-100P, MUC5AC, Muc-2, Muc-6, GATA3, Ber-EP4, EPCAM, Calretinin, CEA, HER-2, or Claudin18.2. On September 4, 2025, the patient underwent a special liver segment resection and intestinal adhesion lysis at our hospital. Intraoperative exploration revealed a small amount of ascites in the abdominal cavity. Multiple nodular masses were found on the omentum and abdominal surface. Postoperative pathology showed an epithelioid cell tumor (left lateral lobe of the liver), consistent with epithelioid malignant mesothelioma (clear cell subtype) based on immunohistochemistry.Immunohistochemical staining results: Tumor cells positively expressed D2-40, WT-1, Calretinin, CK5 / 6, Vimentin, CA9, ARID1α, CD56, CK7, and CK19, but did not express ARG-1, AFP, Hepa, GPC-3, GS, α-inhibin, CD34, S-100P, or PAX-8. Simultaneous CUP tissue tracing analysis showed that the model's prediction results were consistent with the HE and immunohistochemical staining results, all indicating a malignant mesothelioma origin. These three findings corroborate each other, indirectly confirming the reliability and accuracy of the CUP prediction model.

[0156] like Figure 10 The image shown is a PET / CT image of the first discovery of a liver mass. A large, slightly low-density mass with unevenly increased glucose metabolism is seen protruding into the abdominal cavity in the left lateral lobe of the liver, measuring approximately 11.4 × 6.44 cm.

[0157] Figure 11 The images show the pathological features of a mass in the left lateral lobe of the liver. Image A shows the surgical resection of the mass, pathologically revealing lung adenocarcinoma and epithelioid malignant mesothelioma (clear cell subtype). HE staining was performed. Image B shows the immunohistochemical results, indicating tumor cells expressing WT-1. Image C shows the immunohistochemical results, indicating tumor cells expressing D2-40. Image D shows the immunohistochemical results, indicating tumor cells expressing CK5 / 6. Image E shows the immunohistochemical results, indicating tumor cells expressing Calretinin. Image F shows the immunohistochemical results, indicating tumor cells not expressing Hepa. Image G shows the immunohistochemical results, indicating tumor cells not expressing ARG-1. Image H shows the immunohistochemical results, indicating tumor cells not expressing GPC-3.

[0158] Figure 12 The detection and calculation methods provided by the technical solution of this invention also indicate that the mass in the left lateral lobe of the liver is a malignant mesothelioma.

[0159] This invention proposes a targeted RNA sequencing scheme based on NGS, which can simultaneously acquire high-quality GEP expression data for CUP tracing in a single test.

[0160] This invention employs a machine learning process combining SVM-RFE and SVM to automatically select a small, efficient set of feature genes from high-dimensional gene expression data and construct a high-precision classification model.

[0161] This invention establishes a complete technical process from experimentation to analysis, providing a systematic solution for the clinical auxiliary diagnosis of CUP.

[0162] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any form or substance. It should be noted that those skilled in the art can make various improvements and additions without departing from the present invention, and these improvements and additions should also be considered within the scope of protection of the present invention. Any modifications, alterations, and equivalent changes made by those skilled in the art based on the above-disclosed technical content without departing from the spirit and scope of the present invention are equivalent embodiments of the present invention. Furthermore, any modifications, alterations, and evolutions made to the above embodiments based on the essential technology of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A method for NGS and machine learning based de novo primary tumor provenance, characterized in that, Expression data covering at least 1,000 genes were obtained through targeted RNA sequencing; after data preprocessing and standardization, key gene features were screened out using the SVM-RFE algorithm. Based on the feature set, a multi-class prediction model is constructed using SVM, and the model is trained and independently validated.

2. The method of tracing primary tumors of unknown origin based on NGS and machine learning according to claim 1, characterized in that, The method includes the following steps: Step 1: Establish the sample queue and dataset; including two independent queues, namely the discovery queue for model training and the independent validation queue for model performance evaluation; Step 2, Sample Library Construction and Sequencing: Collect FFPE tumor samples covering multiple cancer types, extract RNA and construct strand-specific sequencing libraries, capture 2660 target genes using a target probe panel, and perform paired-end sequencing on a high-throughput sequencing platform; RNA sequencing is performed on the discovery cohort and the independent validation cohort, and DNA sequencing is performed on the independent validation cohort. Step 3: Construction of the CUP classification model; Step 3.1, Data Preprocessing and Feature Generation: Quality control of sequencing data is performed, alignment to a reference genome is performed, and the TPM value of each gene is calculated; further, the TPM is normalized between samples using internal reference genes, and log2 transformation is performed to reduce technical bias; Gene set variation analysis (GSVA) ​​is performed on gene expression profiles to obtain the score of each functional gene set corresponding to the sample; finally, the standardized TPM of each gene and the score of each gene set are used as sample features. Step 3.2, Feature Selection and Model Training: From the standardized expression data, the SVM-RFE method is used to select 250 gene feature subsets with the most discriminative power; based on this feature subset, a multi-classification model is trained using linear kernel SVM. During the training process, the SMOTE algorithm is introduced to handle class imbalance, and the model parameters are optimized through cross-validation. Step 4, Model Validation and Application: Input the preprocessed independent validation set sample expression data into the trained model, output the predicted category of the primary cancer type, and evaluate the accuracy and generalization ability of the model by comparing it with the true label.

3. The method for tracing the origin of tumors of unknown primary foci based on NGS and machine learning according to claim 2, characterized in that, All analyses in step 3.1 were performed using Python and R; the gene expression data of the discovery cohort samples that passed quality control were randomly divided into training and test sets with a ratio of 7:

3.

4. The method of tracing primary tumors of unknown origin based on NGS and machine learning according to claim 2, characterized in that, The data preprocessing in step 3.1 includes selecting internal reference genes with a coefficient of variation (CV) lower than 0.35 from the training set; normalizing the TPM value of each sample by dividing it by the average TPM value of the selected internal reference genes; and performing a log2 transformation on the normalized TPM value.

5. The method of tracing primary site of origin of a tumor of unknown primary based on NGS and machine learning according to claim 2, wherein, In step 3.2, feature selection and model training, the model training includes using an SVM with a linear kernel function; standardizing and scaling the features; using the SMOTE algorithm to solve the class imbalance problem; optimizing hyperparameters on the training set through grid search combined with 5-fold cross-validation; finally, the model is trained on all training set samples and saved for evaluation on the test set and independent validation queue; the evaluation metrics include accuracy, sensitivity, specificity, positive prediction value, and negative prediction value.

6. The method of tracing primary site of origin of a tumor of unknown primary based on NGS and machine learning according to claim 2, wherein, The independent validation cohort in step 1 also includes a Chinese genome dataset cBioPortal, ID:pan_origimed_2020 cohort; this cohort contains 11,553 Chinese patients covering 25 different cancer types, and uses a CYCS NGS panel targeting 450 cancer-related genes for sequencing analysis.

7. The method of tracing primary tumors of origin based on NGS and machine learning according to claim 6, characterized in that, The model validation in step 4 includes the integration of DNA mutation information with the CUP classification model; it includes using the Chinese genome dataset to identify gene sets with significantly different mutation frequencies for each pair of tumor types; and then, in an independent validation cohort, integrating the DNA mutation data of the top three predicted tumor types for each sample to optimize the classification.

8. The method of tracing primary tumors of origin based on NGS and machine learning according to claim 7, characterized in that, The integration of the DNA mutation information with the CUP classification model includes calculating the adjusted probability for tumor type i for a given sample using the following formula: and in: adjusted probability of the i-th tumor type; : weight score for the i-th tumor type; The probability of the i-th tumor type given by the CUP model; : Number of genes with significantly different mutation frequencies for pairs of tumor types; : frequency of the gth differentially expressed gene in tumor type i relative to tumor type j; Based on the probability value of each sample after integrating DNA mutation information, the tumor type with the highest probability value is selected as the prediction result for that sample.

9. Use of a method for the tracing of primary tumors of unknown origin based on NGS and machine learning according to any one of claims 1 to 8, characterized in that, The applications include the development of diagnostic kits for tracing the origin of tumors with unknown primary foci.

10. A system for tracing a primary tumor unknown, characterized in that, The system includes a source tracing detection program that applies the method for tracing the source of tumors of unknown primary origin based on NGS and machine learning, as described in any one of claims 1-8.