Application and detection system of markers for tracing tumor tissue of unknown primary lesions
Through liquid biopsy technology, a copy number variation map analysis was carried out on plasma free DNA at the whole genome level, and a classification model was constructed, which solved the accuracy of tumor tissue traceability in unknown primary foci, and achieved efficient and accurate tumor tissue traceability detection, especially the accuracy of 17 tumor types reached 85.8%.
Patent Information
- Application Number
- CN202410730754.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-06-06
AI Technical Summary
It is difficult for the prior art to accurately trace tumor tissues with unknown primary foci, traditional treatment effect is poor, patients have poor prognosis, traditional detection methods have insufficient samples or inconsistencies, and ctDNA detection sensitivity is limited.
Through liquid biopsy technology, a genome-wide copy number variation map was analyzed for plasma free DNA, a classification model was constructed, and a CNV marker in 140 gene regions was used for tumor tissue traceability, and a H2O AutoML algorithm was used for model training and prediction.
It has achieved efficient and accurate traceability of 17 tumor types, with large number of sample verifications, many types of cancers covered, high detection performance, and accuracy of more than 85.8%.
Smart Images

Figure CN118308490B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for tracing and detecting the origin of tumor tissue of unknown primary focus, belonging to the technical field of tumor gene detection. Background Art
[0002] Cancers of unknown primary (CUP) are pathologically diagnosed metastatic malignancies, but the primary site remains unidentified based on clinical history, a complete physical examination, routine laboratory tests, imaging and radio-metabolic techniques, and careful examination of histological specimens. According to statistics, CUP accounts for approximately 2%-10% of all cancer cases, ranking eighth in incidence and fourth in mortality among common malignant tumors. The incidence of CUP increases significantly with age, becoming uncommon in patients under 40 years of age and peaking around 80 years of age. Studies have reported that CUP occurs most frequently in the respiratory and digestive organs, with the liver being the most commonly documented single metastatic site. A significant number of cases still lack a specific site of unknown primary in cancer registries. The most common histological subtype of CUP is adenocarcinoma, accounting for 42% to 50%, followed by poorly differentiated carcinoma and squamous cell carcinoma. Traditional treatment for CUP primarily relies on broad-spectrum chemotherapy; however, the overall prognosis for patients with CUP is relatively poor. The median survival after chemotherapy is 4.5 months, with a 1-year survival rate of 20% and a 5-year survival rate of only 4.7%. Therefore, accurately inferring the tissue origin of CUP has important clinical significance and can provide CUP patients with efficient diagnosis and guided treatment methods earlier.
[0003] Possible reasons for an unknown primary site include: inadequate testing methods, insufficient tissue sample collection, the primary site being removed, too small, or regressing spontaneously, or widespread metastasis making identification of the primary site difficult. Current diagnostic techniques primarily include comprehensive assessment, imaging, pathological analysis, immunohistochemistry, and genetic testing. Histopathology is the gold standard for CUP diagnosis, while genetic testing is considered an adjunct diagnostic tool when tissue biopsy is inconclusive. Many cancer cells retain characteristics of their tissue of origin during metastasis; in other words, the genetic profile of metastatic cancer should be consistent with that of the original tissue. Studies have found that the genetic profile of metastatic tumors differs from that of the tissue at the metastatic site but is more similar to that of the primary site. Genetic information from the tissue of origin is consistently retained throughout tumor development, progression, and metastasis. Copy number variations are significantly correlated with individual cancers and can be distinguished by detecting copy number variations in cancer-related genes or specific genomic regions. Based on this theory, researchers used second-generation sequencing technology to develop a series of molecular markers derived from gene copy number variations (CNVs) to trace the tissue origin of tumors. They further used bioinformatics methods to extract gene characteristic molecules as specific "molecular fingerprints" to determine the primary lesion, and established corresponding models to predict and determine the primary site of CUP.
[0004] Liquid biopsy technology is an in vitro diagnostic technique that uses samples of bodily fluids such as blood and urine to analyze and diagnose tumors. In recent years, liquid biopsy has become a research hotspot in cancer diagnosis and treatment. Circulating free DNA (cfDNA) in peripheral blood contains circulating tumor DNA (ctDNA) derived from tumors. Therefore, the various characteristics carried by cfDNA can be used for cancer identification and tissue tracing. However, in some cases, the genetic information in ctDNA may be inconsistent or unequal to that in tumor tissue. This is primarily because ctDNA is derived from DNA fragments released into the circulation during apoptosis or necrosis by tumor cells. These characteristics differ in some respects from the genetic information detected directly in tumor tissue. ctDNA primarily originates from necrotic and apoptotic tumor cells, circulating tumor cells, and exosomes secreted by tumor cells. These DNA fragments carry tumor genomic signatures, including mutations, deletions, insertions, rearrangements, copy number abnormalities, and methylation. Because ctDNA is DNA fragments released into the bloodstream by tumor tissue, its content is typically low and is affected by multiple factors, such as tumor stage, tumor burden, tumor type, and individual patient differences. Tumor tissue is heterogeneous, meaning that tumor cells in different regions may harbor distinct genetic variations and expression patterns. Therefore, even if a tumor tissue sample is obtained through a biopsy, it may not fully reflect the overall genetic profile of the tumor due to improper sample selection or insufficient sample size. However, ctDNA may more comprehensively reflect the overall genetic variation of the tumor because it is released from multiple tumor cells. Furthermore, during tumor progression, tumor cells may undergo clonal evolution, meaning that some tumor cells may acquire new genetic mutations that may not be localized in a specific region of the tumor tissue but can be detected by ctDNA. Furthermore, the low content of ctDNA may limit its sensitivity. Consequently, in some cases, ctDNA may fail to detect certain low-frequency mutations or variants present in tumor tissue. Therefore, diagnosing primary tumors based on ctDNA detection information also requires addressing relevant technical challenges. Summary of the Invention
[0005] The purpose of this invention is to perform genome-wide copy number variation profiling of tumors with unknown primary sites through liquid biopsy technology, thereby determining the primary source of the tumor (tissue tracing).
[0006] The dataset used in this patent is derived from 17 tumor types, including lung cancer, colorectal cancer, hepatobiliary cancer, gastric cancer, head and neck cancer, kidney cancer, breast cancer, pancreatic cancer, ovarian cancer, prostate cancer, thyroid cancer, cervical cancer, urothelial cancer, esophageal cancer, endometrial cancer, sarcoma, and melanoma. A total of 4,169 samples were included. The composition of the entire dataset refers to the distribution of different cancer types in CUP reported in current research. Among them, there were 2,212 males, accounting for 53.06%; and 1,957 females, accounting for 46.94%.
[0007] The use of a reagent for detecting gene copy number variation derived from plasma cell-free DNA in the preparation of a reagent for tracing the origins of 17 tumor types, wherein the genes include 140 gene regions, and the locations of the gene regions on the genome are as follows:
[0008] ;
[0009] ;
[0010] ;
[0011] ;
[0012] ;
[0013] ;
[0014] .
[0015] The application further comprises the following steps:
[0016] S1: Obtain plasma samples from tumor patients, extract plasma free DNA, construct libraries to obtain library products, and perform whole-genome sequencing;
[0017] S2: Sequencing data is compared to the reference genome to obtain sequencing data results on the CNV marker region;
[0018] S3: Obtain the CNV value of each CNV marker region;
[0019] S4: Take the CNV value of each CNV marker region as the independent variable and whether it can be traced back to the corresponding tumor type as the dependent variable, build a classifier, train the model, and obtain the classification model; then predict the tissue traceability source of the sample to be tested based on the classification model.
[0020] In step S3, the CNV value of each sample to be tested is calculated by using the depth after correction and normalization of the sample to be tested / the depth after correction and normalization of the population baseline (2*2^log2Ratio) in the region of the CNV value marker to obtain the CNV value of each sample to be tested.
[0021] The reference genome is version hg19.
[0022] The sequencing depth was 4-6 times.
[0023] The classification model uses the probability of tissue tracing back to a certain tumor as the output value.
[0024] A device for tracing and detecting the origin of tumor tissue of unknown primary focus, comprising:
[0025] A sequencing module is used to extract cell-free DNA from plasma samples and perform whole-genome sequencing to obtain sequencing data results for CNV gene marker regions;
[0026] The alignment module is used to align the sequencing data results to the reference genome and obtain the CNV value of each CNV marker region;
[0027] The judgment module is used to take the CNV value of each CNV marker region as the independent variable and whether it can be traced back to a certain cancer as the dependent variable, build a classifier, train the model, and obtain a classification model; then, based on the classification model, predict whether the sample to be tested has tissue traceability.
[0028] In the comparison module, the CNV value of each sample to be tested is calculated by using the depth after correction and normalization of the sample to be tested / the depth after correction and normalization of the population baseline (2*2^log2Ratio) in the region of the CNV value marker.
[0029] The classifier is the H2O AutoML algorithm classifier.
[0030] A computer-readable medium having recorded thereon a computer program capable of performing tissue tracing and classification of 17 types of tumors, including lung cancer, intestinal cancer, hepatobiliary cancer, gastric cancer, head and neck cancer, kidney cancer, breast cancer, pancreatic cancer, ovarian cancer, prostate cancer, thyroid cancer, cervical cancer, urothelial cancer, esophageal cancer, endometrial cancer, sarcoma, and melanoma; the computer program comprising the following steps:
[0031] Compare the sequencing data results of the CNV marker region to the reference genome to obtain the CNV value of each CNV marker region;
[0032] The CNV values in each CNV marker region were used as independent variables, and whether they could be traced to the corresponding tumor type was used as the dependent variable. An H2O AutoML algorithm classifier was constructed and trained to obtain a classification model. The classification model was then used to predict the tissue traceability source of the test sample. Beneficial effects
[0033] The present invention provides for the first time a liquid biopsy gene copy number variation biomarker diagnostic model for tracing the tissue origin of tumors of unknown primary foci from 17 types of tumors, including lung cancer, intestinal cancer, hepatobiliary cancer, gastric cancer, head and neck cancer, kidney cancer, breast cancer, pancreatic cancer, ovarian cancer, prostate cancer, thyroid cancer, cervical cancer, urothelial cancer, esophageal cancer, endometrial cancer, sarcoma and melanoma. The model can trace the tissue origin of tumors of unknown primary foci and has the advantages of a large number of sample verifications, coverage of multiple cancer types and high detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Shown is a flow chart of this patent;
[0035] Figure 2 The sample types and quantities of the 17 tumor datasets used in this patent are shown;
[0036] Figure 3 The box plots of the best 140 CNV markers in tumor plasma samples from different tissue sources are shown;
[0037] Figure 4 The accuracy of the model trained with different algorithms for 140 CNV markers is shown.
[0038] Figure 5 The accuracy of 140 CNV markers in predicting the top 1 and top 2 of different tumor types in the training set is shown;
[0039] Figure 6 The accuracy of 140 CNV markers in predicting the top 1 and top 2 of different tumor types in the validation set is shown;
[0040] Figure 7 The data distribution of the accuracy of 140 CNV markers in predicting the top 1 of different tumor types in the training set is shown.
[0041] Figure 8 The data distribution of the accuracy of 140 CNV markers in predicting the top 2 of different tumor types in the training set is shown.
[0042] Figure 9The data distribution of the accuracy of 140 CNV markers in predicting the top 1 of different tumor types in the validation set is shown.
[0043] Figure 10 The data distribution of the accuracy of 140 CNV markers in predicting the top 2 of different tumor types in the validation set is shown. DETAILED DESCRIPTION Example
[0044] This patent first selects the clinical samples used to form the data set source from the actual samples submitted by the cooperative hospital to the patent applicant and the samples actually tested by the patent applicant. The data set in this patent includes a total of 4,169 clinical plasma samples from patients with 17 types of tumors, including lung cancer, intestinal cancer, hepatobiliary cancer, gastric cancer, head and neck cancer, kidney cancer, breast cancer, pancreatic cancer, ovarian cancer, prostate cancer, thyroid cancer, cervical cancer, urothelial cancer, esophageal cancer, endometrial cancer, sarcoma and melanoma. The specific number of samples and their proportions are: 782 cases (18.76%) ), 383 cases (9.19%), 376 cases (9.02%), 333 cases (7.99%), 328 cases (7.87%), 272 cases (6.52%), 264 cases (6.33%), 248 cases (5.95%), 219 cases (5.25%) ), 172 cases (4.13%), 154 cases (3.69%), 142 cases (3.41%), 119 cases (2.85%), 106 cases (2.54%), 104 cases (2.49%), 101 cases (2.42%), 66 cases (1.58%).
[0045] Among the above samples, all are preoperative plasma samples of 17 types of tumors whose primary tissue origin was confirmed by clinical postoperative pathology. After cfDNA extraction and testing, the tissue information of the primary lesion carried in their ctDNA can be traced; because the information carried by metastatic tumors is the same as that of the primary tumor and can be released into the blood and carried by ctDNA, the plasma free DNA of samples of unknown primary lesion origin is detected to identify the primary lesion tissue information they carry, thereby tracing the tumor type.
[0046] The specific age and gender of the data set used in this invention are as follows:
[0047]
[0048] The specific TNM staging of the dataset used in this invention is as follows:
[0049]
[0050] The detection method of this patent first requires the extraction, library construction, and sequencing of free DNA from plasma samples. A magnetic bead extraction kit is used for extraction, and a whole genome library construction kit is used for library construction. The sequencing process can use existing sequencing technologies to obtain base information of the free DNA.
[0051] The purpose of extracting cell-free DNA from plasma samples is to obtain cell-free DNA from plasma. The extraction used in the present invention mainly includes the following steps:
[0052] Centrifuge 1-5 mL of plasma sample to obtain the supernatant containing free DNA and remove the precipitate.
[0053] Lysis and binding solution, proteinase K and magnetic beads are added to the centrifuged plasma sample for lysis and binding treatment, and the lysis is fully incubated to obtain magnetic beads and DNA binding products.
[0054] The magnetic beads and the DNA binding product were washed in sequence using washing solution 1, washing solution 2, and anhydrous ethanol, respectively, to recover and purify the DNA, thereby obtaining washed magnetic beads and the DNA binding product.
[0055] The DNA was eluted from the magnetic beads using nuclease-free water to obtain purified free DNA product.
[0056] After obtaining the free DNA, the extracted DNA sample is subjected to whole-genome sequencing library construction. Library construction is the process of adding adapters to the sequencing fragments. The library construction used in the present invention mainly includes the following steps:
[0057] End repair and addition of base A: T4 polymerase and Klenow E. coli polymerase were used to fill in the 5' protruding sticky ends of the cfDNA and to blunt the 3' protruding sticky ends to produce blunt ends.
[0058] While performing end repair, base A is added to the 3' end of the cfDNA fragment to obtain cfDNA fragments with sticky end A for subsequent adapter ligation.
[0059] Adapter ligation: Use T4 ligase to connect the adapters to the two ends of the cfDNA fragments.
[0060] Library purification: The library with adapters added is purified by magnetic beads. The purified library can be used for sequencing.
[0061] The plasma free DNA library obtained above was subjected to whole genome sequencing (WGS) with a sequencing depth of approximately 5 times, in order to obtain the base information of the free DNA. Example
[0062] After cell-free DNA was extracted from plasma samples and whole-genome sequencing was completed, bcl2fastq was used to generate fastq files. Data quality control was performed using FastQC software, and adapters and low-quality sequences were removed using Trimmomatic software. The resulting clean data was then aligned to the hs37d5 genome using bwa software to obtain the specific location of each DNA fragment on the genome. Picard software was then used to remove redundant data. Finally, samtools software was used to remove paired-end reads with low alignment quality and sort them by alignment position for further analysis.
[0063] The CNV marker screening process of this patent is as follows:
[0064] Genome-wide gene locus region acquisition: Download the GRCh37 version annotation information released by the GENCODE database (https: / / www.gencodegenes.org / ) and extract its gene regions.
[0065] Initial screening of CNV markers in public databases: This patent uses CNV-related genes from the cBioPortal (https: / / www.cbioportal.org / ) database, which integrates research data from the Cancer Genome Atlas (TCGA, http: / / cancergenome.nih.gov / ) and the International Cancer Genome Consortium (ICGC). CNV genes that frequently appear in 17 cancer types were collected and used as tissue-tracing markers for each cancer type. The requirement was that the frequency of occurrence in a particular cancer type be greater than 5%.
[0066] Analysis of marker gene copy number in WGS test samples: The high-frequency CNV-related genes of cancer collected from the cBioPortal database were summarized. Data preprocessing was performed on the dataset used in this patent, including filtering low-quality reads of Bam files (MAPQ threshold 15), GC correction to eliminate the impact of GC content bias on the number of reads, and PCA noise reduction to find the true CNV difference signal. 200 healthy people were selected as the baseline, and the baseline sample coverage depth of each gene region and the coverage depth of each sample to be tested were calculated. Finally, 2*2^log2Ratio (depth after correction and normalization of the sample to be tested / depth after correction and normalization of the population baseline) was used as the CNV value of each sample to be tested.
[0067] Assessment of marker gene stability in healthy individuals: 200 healthy individuals were randomly selected as a baseline, and the copy number range of all CNV-related genes was calculated. Genes with large fluctuations in baseline sample copy number and severe deviations from the normal copy number (normal 2 copies) were removed.
[0068] Further screening of cancer-specific CNV marker genes: The data set in Example 1 was randomly split into a training set and a validation set in a ratio of 6:4. 50 cases were selected from 17 types of tumor patient samples in the training set, including lung cancer, intestinal cancer, hepatobiliary cancer, gastric cancer, head and neck cancer, kidney cancer, breast cancer, pancreatic cancer, ovarian cancer, prostate cancer, thyroid cancer, cervical cancer, urothelial cancer, esophageal cancer, endometrial cancer, sarcoma and melanoma. The above CNV genes were classified using the random forest multi-classification algorithm of machine learning, and the genes with the top 85% of feature weights were retained as candidate genes. The Wilcoxon rank sum test was then used to screen out the high-frequency CNV-related genes of each cancer type and the genes that were different from other cancer types (p<0.01). The top genes with the most significant differences were selected for each cancer type and the combined set was 140. The genomic positions of the optimal CNVs and the corresponding tumor types are shown in the table below:
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079] Example 3
[0080] According to the split training set and validation set in Example 2, the 140 genes screened above were used as input features. H2O AutoML was used to create a classifier for model training. During the training process, the H2O algorithm used different machine learning algorithms to train the features multiple times. Each training session would adjust the relevant parameters based on the results of the previous training session until the model performance reached the preset target or the number of training sessions reached a threshold. Ultimately, each CNV marker feature would have a different number of models trained by different algorithms. Model training algorithms include: random forest (DRF), hyperrandom forest (XRT), generalized linear model (GLM), extreme gradient boosting model (XGBoost), gradient boosting model (GBM), deep learning model (DeepLearning), and an integrated model (Stacked Ensemble) that integrates all models. The model with the highest accuracy in the training set was selected as the optimal model, and its performance was statistically analyzed in the validation set.
[0081] The 140 CNVs identified were used as predictive markers for tracing the origin of unknown primary tumor tissues. The H2O AutoML method was used to establish an optimal prediction model. After model training, the validation samples were used to predict classification results, determining the probability of tracing the origin of different cancer types. The top two cancer types (first and second classifications) were selected for accuracy verification. The prediction results for the validation set showed an accuracy of 85.8%, a sensitivity of 85.79%, and a specificity of 99.07% for the first classification; the accuracy of the second classification was 88.8%, a sensitivity of 88.31%, and a specificity of 99.23%. For clinical application, the prediction should be based on other actual clinical indicators.
Claims
1. Use of a reagent for detecting gene copy number variation derived from plasma free DNA in the preparation of a reagent for tracing the origin of 17 tumor types, characterized in that: Tumor types include lung cancer, intestinal cancer, hepatobiliary cancer, gastric cancer, head and neck cancer, kidney cancer, breast cancer, pancreatic cancer, ovarian cancer, prostate cancer, thyroid cancer, cervical cancer, urothelial cancer, esophageal cancer, endometrial cancer, sarcoma, and melanoma; the genes are composed of the following 140 gene regions, and the locations of the gene regions on the genome are as follows: ; ; ; ; ; 。
Citation Information
Patent Citations
Prediction model for simultaneously detecting multiple tumors and performing tissue traceability and training method and application thereof
CN116665771A
DNA copy number variation-based prediction method for kind of cancer
WO2019066421A2