Scale carcinoma tissue tracing method and system based on methylation chip and machine learning

Through a method based on methylation chip and machine learning, a classification prediction model for squamous cell carcinoma and urothelial carcinoma tissues was constructed, which solved the problem of difficult to accurately diagnose squamous cell carcinoma and urothelial carcinoma in the prior art, and achieved high-accurate tissue origin prediction.

CN120199334APending Publication Date: 2025-06-24FUDAN UNIV SHANGHAI CANCER CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510148388.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately diagnose the tissue origin of squamous cell carcinoma and urothelial carcinoma with unknown primary foci, especially in cases with similar morphology and immunohistochemistry expression patterns.

Method used

Using a method based on methylation chips and machine learning, a classification prediction model for squamous cell carcinoma and urothelial carcinoma tissue was constructed by screening specific methylation sites. This model uses the CatBoost algorithm and combines the big data in the public database to establish a methylation traceability model with superior diagnostic performance.

Benefits of technology

It has achieved accurate prediction of the origin of squamous cell carcinoma and urothelial carcinoma, improved the accuracy and reliability of diagnosis, and has important clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199334A_ABST
    Figure CN120199334A_ABST
Patent Text Reader

Abstract

The invention relates to a squamous cell carcinoma tissue tracing method and system based on a methylation chip and machine learning, and belongs to the technical field of biological medicine. Comprising the following steps: acquiring DNA methylation data of squamous cell carcinoma and urothelial carcinoma samples from a big database, and dividing the data into a training set, a test set and a verification set; the DNA methylation data is preprocessed, and then a specific methylation site is screened out; constructing a model for classifying and predicting squamous cell carcinoma and urothelium carcinoma tissues through a machine learning algorithm; the specific methylation sites comprise 106 methylation sites. A tissue origin prediction model constructed by combining methylation sites with a CatBoost machine learning algorithm can effectively predict primary lesions of squamous cell carcinoma and urothelial carcinoma. The prediction accuracy in an actual clinical sample reaches 83.5%, and clinical diagnosis and treatment requirements can be effectively met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for tracing squamous cell carcinoma tissues based on methylation chips and machine learning, and belongs to the field of biomedical technologies. Background Art

[0002] Cancer of unknown primary (CUP) is a type of malignant tumor that is histopathologically diagnosed as metastatic, but the primary site still cannot be determined after exhaustive detection and evaluation. Although its incidence is relatively low, about 6 / 100,000 - 16 / 100,000, its malignancy is high and the prognosis of patients is extremely poor. The results of multiple clinical trials show that site-specific treatment targeting the predicted primary site of the tumor based on information such as gene expression or variation shows better therapeutic value. [1-3] For true CUP cases where the primary site cannot be detected by routine examinations including blood and imaging, histopathological examination is the basis of clinical examination. Pathological examination includes morphological and immunohistochemical examinations. According to tissue morphological characteristics, CUP can usually be preliminarily classified into: ① highly and moderately differentiated adenocarcinoma (50%), ② poorly differentiated carcinoma / adenocarcinoma (30%), ③ squamous cell carcinoma (15%), ④ undifferentiated tumor (5%). [4] Although the most common histological type is adenocarcinoma, the morphologies of adenocarcinomas in different sites (such as breast cancer, colorectal adenocarcinoma, gastric adenocarcinoma, lung adenocarcinoma, endometrial adenocarcinoma, etc.) are different, and there are still many organ-specific immunohistochemical markers that can be used (such as GATA3, mammaglobin, TRPS1 in breast cancer; TTF1, NapsinA in lung cancer; CDX2, CK20 in colorectal adenocarcinoma; CK7, STAT3 in gastric cancer, etc.) to assist in diagnosing the primary site. However, the morphologies of squamous cell carcinomas in different sites are basically the same, and there are no effective organ-specific immunohistochemical markers; in actual clinical work, the diagnosis and differential diagnosis of the tissue origin of metastatic squamous cell carcinoma are the difficulties and blind spots of diagnosis. Previous research results also show that the proportion of squamous cell carcinoma in CUP cases in China is higher than that reported in the West. [5] It can reach 20% - 30%. [6,7] Moreover, previous literature on tissue origin models based on RNA expression profiles and broad methylation levels also shows that the accuracy of squamous cell carcinoma tracing is relatively low. [8-10] In addition, considering that urothelial carcinoma in the bladder / ureter shows a high similarity in morphology and immunohistochemical expression patterns with squamous cell carcinoma during invasion and metastasis, therefore, establishing a tissue origin prediction model based on common squamous cell carcinomas (lung, esophagus, cervix, head and neck) and urothelial carcinoma cases in China has important and urgent clinical significance, which is of great help to the timely diagnosis, treatment and prognosis of patients.

[0003] In recent years, with the rapid development of technologies in the field of molecular biology, prediction models related to tumor types have been established at different molecular genetic levels through the combination of biochips and high-throughput sequencing technologies with big data analysis methods. The detection of DNA methylation status has been increasingly widely applied and highly regarded in clinical work, covering multiple aspects such as early screening, diagnosis, treatment plan selection, and prognosis of tumors. In terms of tumor diagnosis and classification, in 2018, the team at the Heidelberg Children's Cancer Research Center established 91 molecular classifications based on methylation levels for central nervous system (CNS) tumors, and the compliance of the classification was confirmed to be as high as 88% in the pathological verification of more than 1,100 tumor samples.

[11] Subsequently, in 2021, the team also explored a methylation classification model for 62 types of soft tissue and bone sarcomas. After comparative analysis with histological diagnosis, the diagnoses of 29 cases (29 / 428, 7%) of samples were revised, indicating the important diagnostic value of this methylation classification model and providing an online shared analysis platform.

[12] The results of recent studies also showed that for the model established using 28 data classifiers based on methylation detection data of paraffin tissues, the prediction accuracy of metastatic tumors was 81%-93%, and the primary sites of 81%-93% of cases were successfully predicted in 68 patients with carcinoma of unknown primary (CUP).

[13] It also includes that methylation profiles can effectively distinguish the tissue origins of neuroendocrine tumors in different locations.

[14] The above research results all showed the important clinical value of DNA methylation characteristics in tumor classification.

[15] It is suggested that DNA methylation profiles with superior tissue origin prediction performance can be tried to solve the problem of the primary sites of squamous cell carcinomas in different locations with difficult differential diagnosis.

[0004] Existing public databases, including The Cancer Genome Atlas (TCGA), Gene Expression Omnibus (GEO), EBI ArrayExpress / Stanford Microarray Database (SMD), and Yale Microarray Database (YMD), etc., contain various information on the clinicopathology and molecular genetics of more than 30 types of tumors and tens of thousands of patients. Methylation data is an important part of them, and how to analyze and process the data is extremely crucial. There are various machine learning algorithms for feature extraction from big data. In the ensemble learning model, it mainly includes the random forest model based on the Bagging algorithm and the XGBoost model based on the Boosting algorithm, which have good system diversity and classification performance. In addition, compared with classical machine learning algorithms, the classification model established based on the convolutional neural network (CNN) can perform multiple processing on the data, balance the quantity differences between different groups, reduce the requirements for the data volume, and extract deep-level high-level features for more accurate classification. [16-18] 。

[0005] Therefore, the present invention innovatively addresses the difficult problem of differentiating and diagnosing squamous cell carcinomas at different sites in clinical work. For the first time, a tissue origin classification model with superior predictive performance is established based on the DNA methylation level, and it includes as few methylation sites as possible, which is expected to be successfully translated into clinical applications. By using more practical NGS or PCR methods, the clinical problem of tracing the origin of squamous cell carcinoma with unknown primary site can be ultimately solved.

[0006] References

[0007] [1] Olivier T, Fernandez E, Labidi-Galy I, et al. Redefining cancer of unknown primary: Is precision medicine really shifting the paradigm?[J]. Cancer Treatment Reviews, 2021, 97: 102204.

[0008] [2] Varadhachary G R, Raber M N. Cancer of unknown primary site[J]. The New England journal of medicine, 2014, 371(8): 757 - 65.

[0009] [3] Cai Ming, Nian Weiqi, Wang Hongwei, et al. Chinese Expert Consensus on the Clinical Diagnosis and Treatment Practice of Tumors of Unknown Primary Based on Molecular Guidance (2023 Edition) [J]. Chinese Journal of Cancer Prevention and Treatment, 2023, 15(03): 252-62. [4] Rassy E, Pavlidis N. The currently declining incidence of cancer of unknown primary [J]. Cancer Epidemiol, 2019, 61: 139-41.

[0010] [5] Mnatsakanyan E, Tung WC, Caine B, et al. Cancer of unknown primary: time trends in incidence, United States [J]. Cancer Causes Control, 2014, 25(6): 747-57.

[0011] [6] Urban D, Rao A, Bressel M, et al. Cancer of unknown primary: a population-based analysis of temporal change and socioeconomic disparities [J]. British Journal of Cancer, 2013, 109(5): 1318-24.

[0012] [7] Pentheroudakis G, Golfinopoulos V, Pavlidis N. Switching benchmarks in cancer of unknown primary: From autopsy to microarray [J]. European Journal of Cancer, 2007, 43(14): 2026-36.

[0013] [8] Pavlidis N, Fizazi K. Carcinoma of unknown primary (CUP) [J]. Critical Reviews in Oncology / Hematology, 2009, 69(3): 271-8.

[0014] [9]Rassy E, Pavlidis N. Progress in refining the clinical management of cancer of unknown primary in the molecular era[J]. Nature reviews. Clinical oncology, 2020, 17(9): 541-54.

[0015]

[10] Kamposioras K, Pentheroudakis G, Pavlidis N. Exploring the biology of cancer of unknown primary: breakthroughs and drawbacks[J]. European Journal of Clinical Investigation, 2013, 43(5): 491-500.

[0016]

[11] Chebly A, YamMINe T, Boussios S, et al. Chromosomal instability in cancers of unknown primary[J]. European Journal of Cancer, 2022, 172: 323-5.

[0017]

[12] López-Lázaro M. The migration ability of stem cells can explain the existence of cancer of unknown primary site. Rethinking metastasis[J]. Oncoscience, 2015, 2(5): 467-75.

[0018]

[13] Penson A, Camacho N, Zheng Y, et al. Development of Genome-Derived Tumor Type Prediction to Inform Clinical Cancer Care[J]. JAMA Oncology, 2020, 6(1): 84-91.

[0019]

[14] Capper D, Jones D T W, Sill M, et al. DNA methylation-based classification of central nervous system tumours[J]. Nature, 2018, 555(7697): 469-74.

[0020]

[15] Koelsche C, Schrimpf D, Stichel D, et al. Sarcoma classification by DNA methylation profiling[J]. Nature Communications, 2021, 12(1).

[0021]

[16] Yuan Y, Shi Y, Li C, et al. DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutations. Bmc Bioinformatics. 2016;17.

[0022]

[17] Zhao Y, Pan Z, Namburi S, et al. CUP-AI-Dx: A tool for inferring cancer tissue of origin and molecular subtype using RNA gene-expression data and artificial intelligence. Ebiomedicine. 2020;61: 103030.

[0023]

[18] Yuan Y, Shi Y, Su X, et al. Cancer type prediction based on copy number aberration and chromatin 3D structure with convolutional neural networks. Bmc Genomics. 2018;19: 565。 Summary of the Invention

[0024] The object of the present invention is to solve the technical problem of how to perform tissue tracing on squamous cell carcinoma and urothelial carcinoma with unknown primary sites.

[0025] To achieve the purpose of solving the above problems, the present invention provides a method for tracing squamous cell carcinoma tissues based on methylation chips and machine learning, including obtaining DNA methylation data of squamous cell carcinoma and urothelial carcinoma samples, and dividing the data into a training set, a test set, and a validation set; preprocessing the DNA methylation data and screening out specific methylation sites; and then constructing a model for classifying and predicting squamous cell carcinoma and urothelial carcinoma tissues through a machine learning algorithm; the specific methylation sites include:

[0026] cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg17583946, cg07049592, cg00892567, cg13752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352, cg04271791, cg00646731, cg01561259, cg09646197, cg11335133, cg10089145, cg05343480, cg23830290, cg04863892, cg07145664, cg10532384, cg17153122, cg13661519, cg03763508, cg02720618, cg18585988, cg10191210, cg18953784, cg06786372, cg13544006, cg01612366, cg26333652, cg19480724, cg19468528, cg11246938, cg09263516, cg08557649, cg08198862, cg13905586, cg02776128, cg13030332, cg14434755, cg13356117, cg06395028, cg11423130, cg05787952, cg26292895, cg25228995, cg21341817, cg25408237, cg22149516, cg13717446, cg20425384, cg26347197, cg16561266, cg02531516, cg09143801, cg03113878, cg06328338, cg20562447, cg08382235, cg04003500, cg22289434, cg02900995, cg14327531, cg21929771, cg22216738, cg25541958, cg02500392, cg02374996, cg17923947, cg17749454, cg01610488, cg27039662, cg08545287, cg04075191, cg16831085, cg21981270, cg23900712, cg04143876, cg04194494, cg02506353, cg17853216, cg02125316, cg19221545,cg08922090, cg11272874, cg20749769, cg10725316, cg25123470, cg13551227, cg19702194, cg21574675, cg02928916, cg10767350, cg10982664, cg10096177, cg22060611, cg24051554, cg15153684, a total of 106 methylation sites.

[0027] Preferably, the machine learning algorithm is the CatBoost algorithm.

[0028] The present invention provides a squamous cell carcinoma tissue traceability system based on a methylation chip and machine learning. The system includes a model for classifying and predicting squamous cell carcinoma and urothelial carcinoma tissues. The model obtains DNA methylation data of squamous cell carcinoma and urothelial carcinoma samples, and divides the data into a training set, a test set, and a validation set; after preprocessing the DNA methylation data, specific methylation sites are screened out; and then the model is constructed through a machine learning algorithm. The specific methylation sites include:

[0029] cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg17583946, cg07049592, cg00892567, cg13752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352, cg04271791, cg00646731, cg01561259, cg09646197, cg11335133, cg10089145, cg05343480, cg23830290, cg04863892, cg07145664, cg10532384, cg17153122, cg13661519, cg03763508, cg02720618, cg18585988, cg10191210, cg18953784, cg06786372, cg13544006, cg01612366, cg26333652, cg19480724, cg19468528, cg11246938, cg09263516, cg08557649, cg08198862, cg13905586, cg02776128, cg13030332, cg14434755, cg13356117, cg06395028, cg11423130, cg05787952, cg26292895, cg25228995, cg21341817, cg25408237, cg22149516, cg13717446, cg20425384, cg26347197, cg16561266, cg02531516, cg09143801, cg03113878, cg06328338, cg20562447, cg08382235, cg04003500, cg22289434, cg02900995, cg14327531, cg21929771, cg22216738, cg25541958, cg02500392, cg02374996, cg17923947, cg17749454, cg01610488, cg27039662, cg08545287, cg04075191, cg16831085, cg21981270, cg23900712, cg04143876, cg04194494, cg02506353, cg17853216, cg02125316, cg19221545,cg08922090, cg11272874, cg20749769, cg10725316, cg25123470, cg13551227, cg19702194, cg21574675, cg02928916, cg10767350, cg10982664, cg10096177, cg22060611, cg24051554, cg15153684, a total of 106 methylation sites.

[0030] The present invention provides an application of a methylation marker in the preparation of a diagnostic reagent for the primary focus of squamous cell carcinoma and urothelial carcinoma tissues. The methylation marker is a specific methylation site, including: cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg17583946, cg07049592, cg00892567, cg13752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352, cg04271791, cg00646731, cg01561259, cg09646197, cg11335133, cg10089145, cg05343480, cg23830290, cg04863892, cg07145664, cg10532384, cg17153122, cg13661519, cg03763508, cg02720618, cg18585988, cg10191210, cg18953784, cg06786372, cg13544006, cg01612366, cg26333652, cg19480724, cg19468528, cg11246938, cg09263516, cg08557649, cg08198862, cg13905586, cg02776128, cg13030332, cg14434755, cg13356117, cg06395028, cg11423130, cg05787952, cg26292895, cg25228995, cg21341817, cg25408237, cg22149516, cg13717446, cg20425384, cg26347197, cg16561266, cg02531516, cg09143801, cg03113878, cg06328338, cg20562447, cg08382235, cg04003500, cg22289434, cg02900995, cg14327531, cg21929771, cg22216738, cg25541958, cg02500392, cg02374996, cg17923947, cg17749454, cg01610488, cg27039662, cg08545287, cg04075191, cg16831085, cg21981270, cg23900712,cg04143876, cg04194494, cg02506353, cg17853216, cg02125316, cg19221545, cg08922090, cg11272874, cg20749769, cg10725316, cg25123470, cg13551227, cg19702194, cg21574675, cg02928916, cg10767350, cg10982664, cg10096177, cg22060611, cg24051554, cg15153684, a total of 106 methylation sites.,

[0031] Preferably, the application includes preparing a diagnostic kit for diagnosing the primary focus of squamous cell carcinoma or urothelial carcinoma tissues.

[0032] Preferably, the application includes preparing a diagnostic kit for differential diagnosis of the primary focus of squamous cell carcinoma or urothelial carcinoma tissues.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The squamous cell carcinoma traceability system provided by the present invention combines methylation chip data and machine learning technology. Based on a large amount of data in a public database, a variety of machine learning algorithms are used to establish a methylation traceability model with excellent diagnostic performance and a small number of sites, and further verified in the methylation chip detection data of a large number of Chinese patients.

[0035] When a patient has multiple squamous cell carcinomas at different sites simultaneously or metachronously, it is necessary to clarify whether it is a second primary or metastatic tumor, which is very crucial for the treatment and staging prognosis evaluation of the patient's disease. Or when metastatic squamous cell carcinoma occurs, but it is difficult to clarify the location of the primary focus through clinical physical examination, hematology and imaging examinations, the present invention can predict the location of the primary focus, assist in identifying the tissue origin, contribute to the determination of the treatment plan, and improve the treatment effect and prognosis of the patient. Or when a patient has a history of urothelial carcinoma and presents a lesion with the morphology of squamous cell carcinoma at other sites, since urothelial carcinoma often undergoes squamous metaplasia during metastasis and is difficult to effectively distinguish from squamous cell carcinoma, the present invention can assist in differentiating metastatic urothelial carcinoma or second primary squamous cell carcinoma. Therefore, in summary, the present invention has high clinical application value in the identification of the tissue origin of squamous cell carcinoma and urothelial carcinoma.

[0036] Based on the Illumina methylation chip technology commonly used in clinical work, the present invention obtains more than about 900,000 methylation sites that widely cover the whole genome. The detection method is mature, with high clinical accessibility and inclusiveness, applicable to paraffin-embedded samples stored within 5 years, and the data analysis method is relatively simple and publicly available. By using a machine learning algorithm with excellent classification efficiency, it can accurately and effectively screen site-specific methylation sites and construct a squamous cell carcinoma traceability model with few methylation sites and excellent performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flowchart for constructing the model of the present invention.

[0038] Figure 2 It is a comparison chart of the performance of prediction models by different machine learning algorithms;

[0039] Among them, the CatBoost algorithm has the highest prediction accuracy.

[0040] Figure 3 It is a screening chart of the number of methylation sites included in the methylation prediction model;

[0041] Figure 4 It is a weight chart of 106 specific methylation sites included in the methylation prediction model;

[0042] Figure 5 It is a chart showing significant differences in the expression of 106 methylation sites included in the classification model in 5 types of tumors;

[0043] Among them, Figure A is the PanCanAtlas test dataset; Figure B is the external public database test set; Figure C is the external validation set of FUSCC (Fudan University Shanghai Cancer Center).

[0044] Figure 6 It is a chart of the model performance in the training set;

[0045] Among them, Figure A is the ROC curve; Figure B is the confusion matrix; Figure C is the model recall rate curve;

[0046] Figure 7 It is a chart of the model performance in the external public database test set;

[0047] Among them, Figure A is the ROC curve; Figure B is the confusion matrix; Figure C is the model recall rate curve;

[0048] Figure 8 It is a chart of the model performance in the FUSCC validation set of this center;

[0049] Among them, Figure A is the ROC curve; Figure B is the confusion matrix; Figure C is the model recall rate curve. DETAILED DESCRIPTION OF THE INVENTION

[0050] To make the present invention more obvious and understandable, the following is a detailed description with preferred embodiments and accompanying drawings:

[0051] Example (1) Case screening and data collection:

[0052] Retrospectively collect samples of primary and metastatic lung, head and neck, esophageal, and cervical squamous cell carcinomas and urothelial carcinoma samples diagnosed in the Pathology Department of FUSCC (Fudan University Shanghai Cancer Center) from May 2016 to December 2023. All samples were diagnosed by at least two professional pathologists, and it was evaluated that there were sufficient FFPE (Formalin-Fixed Paraffin-Embedded) samples to extract DNA for subsequent methylation chip detection.

[0053] (2) Methylation detection

[0054] 1. DNA extraction and determination: Use the Qiagen QIAsymphony DSP DNA mini Kit to extract DNA from FFPE samples. The Qubit dsDNA assay method was used to quantify the extracted DNA. Evaluate the total amount of DNA in the test samples. The minimum input amount for methylation detection of FFPE samples is 500 ng.

[0055] 2. DNA sample quality control: Use the QIAxcel DNA High Resolution Cartridge fragment analysis kit to detect the size of sample DNA fragments in the QIAxcel nucleic acid analyzer (QIAGEN). The running mode is OM500-AM10s, and the Alignment marker range used is 15 bp - 3000 bp. The classification criteria are as follows: Grade A - the genomic main peak is obvious at about 10 Kbp; Grade B - the main peak range is in the 2K - 10Kbp interval, and there is no obvious peak less than 200 bp; Grade C - the peaks are mainly concentrated in 200 - 2000 bp, and the main peak is above 1000 bp; Grade D - the main peak range is 200 - 500 bp. Among them, samples of Grade A and Grade B can be used for subsequent methylation detection. Samples of Grade C need to be focused on during subsequent experiments, and the experimental operation can be stopped if necessary. Samples of Grade D with unqualified quality control are not included in the subsequent experimental content.

[0056] 3. Bisulfite conversion: Use the EZ DNA Methylation-Gold TM Kit (D5005) to perform bisulfite conversion. The total DNA input requirement is: not less than 500 ng for FFPE samples. If the FFPE samples are stored for more than 3 years, it is recommended that the input amount is not less than 1000 ng. The specific experimental operation steps are as follows:

[0057] 1) Reagent preparation: Add 900 μl of water, 300 μl of M-Dilution Buffer, and 50 μl of M-Dissolving Buffer to the CT Conversion Reagent tube to prepare the CT conversion reagent, and mix well by shaking at room temperature for 10 minutes. Add 24 ml of absolute ethanol to 6 ml of M-Wash buffer and mix well for later use.

[0058] 2) PCR reaction: Add 130 μl of the CT conversion reagent to 20 μl of the DNA sample. If the DNA sample is less than 20 μl, make it up with water. When the total amount of DNA is low, up to 40 μl of the DNA sample and 110 μl of the CT conversion reagent can be added. Invert and mix well, then centrifuge. The PCR reaction system is 98 °C for 10 minutes - 64 °C for 2.5 hours (150 minutes) - store at 4 °C for up to 20 hours.

[0059] 3) Elution: Add the sample after PCR reaction to the purification column containing 600 μl of M-binding buffer, invert and mix well, centrifuge at the maximum speed of the centrifuge (>10000 g) for 30 s, and discard the waste liquid; add 100 μl of M-Wash Buffer to the purification column, centrifuge at the maximum speed of the centrifuge (>10000 g) for 30 s, and discard the waste liquid; add 200 μl of M-Desulphonation Buffer to the purification column, let it stand at room temperature (20 - 30 °C) for 15 - 20 minutes, centrifuge at the maximum speed of the centrifuge (>10000 g) for 30 s, and discard the waste liquid; add 200 μl of M-Wash Buffer to the purification column, centrifuge at the maximum speed of the centrifuge (>10000 g) for 30 s, and discard the waste liquid, repeat this step once more; place the purification column into a new 1.5 ml centrifuge tube, add 10 μl of M-Elution Buffer to the purification column, centrifuge at the maximum speed of the centrifuge (>10000 g) for 30 s to obtain the eluted DNA for later use.

[0060] 4. Isothermal amplification: Add 4 μl of the sample obtained from the above transformation to 20 μl of MA1, then add 4 μl of 0.1 N NaOH, mix well by shaking at 1600 rpm for 1 minute, centrifuge briefly, and incubate at room temperature for 10 minutes. Add 68 μl of RPM and 75 μl of MSM to the above sample tube, mix well by shaking for 1 minute and centrifuge briefly. Incubate in a hybridization oven at 37 °C for 20 - 24 hours.

[0061] 5. DNA Fragmentation and Resuspension: After taking out the isothermal amplification product from the hybridization oven and performing an instantaneous centrifugation, add 50 μl of FMS reagent and incubate in the hybridization oven at 37°C for 1 hour. Add 100 μl of PM1 to each sample tube and incubate in the 37°C hybridization oven for 5 minutes. Add 300 μl of 100% isopropanol to each sample, incubate in a 4°C refrigerator for 30 minutes. Centrifuge at 3000 g for 20 minutes at 4°C to ensure that blue crystals appear at the bottom of each sample tube. Take out the sample tubes as soon as the centrifugation is completed, quickly invert them, and slowly pour out the supernatant. Keep the sample tubes upside down and air-dry at room temperature in a fume hood for 1 hour, ensuring that the blue crystals remain at the bottom of the tubes. Add 46 μl of RA1 reagent to each sample tube, place it in a preheated hybridization oven, incubate at 48°C for 1 hour, shake and mix evenly at 1800 rpm for 1 minute, perform an instantaneous centrifugation, and repeat the above steps until the precipitate dissolves.

[0062] 6. Chip Hybridization: After resuspension, denature the sample on a 95°C heater for 20 minutes, cool at room temperature for more than 30 minutes, and perform an instantaneous centrifugation. Assemble the hybridization chambers, and add 400 μl of PB2 to the humidity chamber of the hybridization chambers. Take out the chip from the packaging box, taking care not to touch the surface of the chip with your hands. Use a pipette to aspirate 26 μl of the DNA sample and add it to the corresponding chip wells A1 - H1 in sequence, and record the chip number and sample position. Before injecting the sample, ensure that the pipette tip is placed at the sample injection port of the chip, taking care to avoid generating air bubbles. After adding the sample, cover and lock the lid. Incubate in a 48°C hybridization oven for at least 16 hours and no more than 24 hours.

[0063] 7. Chip Washing: Take out the hybridization chambers from the hybridization oven and cool them at room temperature on a horizontal table for 30 minutes. Prepare 2 wash boxes containing 200 ml of PB1 washing solution, and prepare a multi-sample beadchip alignment fixture and pour 150 ml of PB1 washing solution into it. Take out the chip and remove the sealing film, taking care not to touch the hybridization area of the chip with your hands. Quickly and carefully place the chip in the chip washing rack and immerse it in the wash box containing 200 ml of PB1 washing solution, and lift it up and down to wash the chip. Assemble the flow-through chambers in the multi-sample beadchip alignment fixture, and clip two metal buckles at about 5 mm from the upper edge and the reagent trough respectively.

[0064] 8. Chip Extension and Staining: Assemble and set up the Chamber Rack, ensure that the water circulator is at the appropriate water level, turn on the switch, set the temperature to 44 °C until the temperature stabilizes, and place the chip washing chamber. Add RA1, XC1, XC2, TEM, 95% formamide / 1 nM EDTA (preparation method: 9.5 ml formamide + 20 μl 0.5 M EDTA (pH = 8) + 480 μl H2O), and XC3 into the reagent trough in batches for incubation. Set the water circulator to 32 °C, and add STM, XC3, ATM, XC3, STM, XC3, ATM, XC3, STM in sequence for incubation for the corresponding time.

[0065] 9. Chip Washing and Encapsulation: Prepare two washing trays and pour PB1 and XC4 into them respectively. Disassemble the chip washing chamber and take out the chip, taking care not to touch the front of the chip. Place the chip on the staining rack and first wash and soak it in the PB1 reagent, and then soak it in XC4. Quickly take out the staining rack with the chip in XC4 and place it on the storage rack or storage box, ensuring that the front of the chip is facing up. Dry it in a vacuum dryer or a biosafety cabinet for 50 - 55 minutes.

[0066] 10. Chip Scanning: Pre-download the dmap file corresponding to each chip, and perform chip scanning and collect data on the Illumina 550Dx instrument.

[0067] The above steps are operated according to the operation steps and reagents provided by the commercially available kit.

[0068] (III) Establishment of Methylation Classification Model (Establishment of Methylation Prediction Model for Four Types of Squamous Cell Carcinomas and Urothelial Carcinoma):

[0069] 1. Data Source:

[0070] The Pan-Cancer Atlas (PanCanAtlas) dataset integrates molecular changes at the DNA, RNA, proteome, and epigenetic levels of 33 human tumors in TCGA. In this study, DNA methylation data of five types of tumor samples were downloaded from PanCanAtlas as the training set and internal test set, totaling 1,651 cases, including 195 cases of cervical squamous cell carcinoma, 364 cases of lung squamous cell carcinoma, 154 cases of esophageal squamous cell carcinoma, 523 cases of head and neck squamous cell carcinoma, and 415 cases of urothelial carcinoma; the methylation data were all based on the Illumina HumanMethylation450 BeadChip platform. The external test set cohort was from public databases such as GEO, TCGA, and ArrayExpress, including 107 cases of cervical squamous cell carcinoma (CGCI-HTMCP-CC), 199 cases of lung squamous cell carcinoma (CPTAC-3), 39 cases of esophageal squamous cell carcinoma (GSE178212, 24 cases; GSE121930, 15 cases), 120 cases of head and neck squamous cell carcinoma (GSE178216, 15 cases; GSE178218, 20 cases; GSE178219, 32 cases; E-MTAB-10576, 53 cases), and 17 cases of urothelial carcinoma (GSE222933). The specific data sources of the external test set and the methylation chip platforms used are shown in Table 1. The external validation set cohort was from 5 types of tumor samples diagnosed by FUSCC in our center, totaling 107 cases, including 21 cases of cervical squamous cell carcinoma, 20 cases of lung squamous cell carcinoma, 24 cases of esophageal squamous cell carcinoma, 21 cases of head and neck squamous cell carcinoma, and 21 cases of urothelial carcinoma. The 107 cases of data in the external validation set cohort were from FFPE samples of 5 types of tumors diagnosed by FUSCC in our center. The specific case information and methylation detection chip types are shown in Table 2.

[0071] Table 1 Data sources of the external test set cohort

[0072]

[0073]

[0074] Table 2 Case information of the external validation set cohort of Fudan University Shanghai Cancer Center

[0075]

[0076] 2. Data preprocessing: The PanCanAtlas, GSE178212, GSE121930, GSE178216, GSE178218, GSE178219, and E-MTAB-10576 data cohorts used in the study were all based on the Illumina HumanMethylation450BeadChip platform. The 45 samples in the GSE222933, CGCI-HTMCP-CC, CPTAC-3, and external validation set cohorts were all based on the Illumina Infinium MethylationEPIC BeadChip platform, and the remaining 62 samples in the external validation set cohort were based on the Illumina Infinium MethylationEPIC v2.0 BeadChip platform. All data were analyzed using Beta signal values. For the PanCanAtlas cohort, the champ.filter method in the R package ChAMP (v2.29.1) was used for preprocessing to filter low-quality probes, and the exclusion criteria included: (1) probes with missing Beta signal values; (2) probes at non-CpG sites; (3) probes related to single nucleotide polymorphisms (SNPs); (4) probes mapped to multiple locations; and (5) probes on sex chromosomes. The champ.norm method in the R package ChAMP (v2.29.1) was used to correct and standardize type I and type II probes for the data through BMIQ (Beta Mixture Quantile dilation). For external test set cohorts such as CGCI-HTMCP-CC, CPTAC-3, GSE121930, E-MTAB-10576, GSE222933, and the external validation set cohort, the openSesame process in the R package sesame (v1.20.0) was used for preprocessing.

[0077] 3. Specific methylation site screening: Five tumor types in the PanCanAtlas cohort were divided into 15 groups for further analysis according to two comparison methods: 1 vs. 1 (i.e., pairwise comparison of tumor types, a total of 10 groups) and 1 vs. all (i.e., comparison of a single tumor type with other tumor types, a total of 5 groups). Limma differential analysis and ROC (Receiver Operating Characteristic) curve analysis were performed on these 15 groups of cohorts respectively. In these two analyses, specific methylation sites that met two conditions were screened out: (1) sites with FDR (False Discovery Rate) less than 0.05 in the differential analysis; (2) the top 20 sites with the highest Area Under the Curve (AUC) of the ROC curve. Since the methylation sites included in different datasets are different, only the sites common to the three chip platforms were included during screening. Limma differential analysis was implemented using the champ.DMP method in the R package ChAMP (v2.29.1), and ROC analysis was completed through the ROC method in the R package pROC (v1.18.5).

[0078] 4. Machine learning model construction: To streamline the model, for the specific methylation sites that met the conditions, 50% of the feature sites with the highest contribution to the model were retained. Feature selection was performed using the SelectFromModel method in the Python library scikit-learn (v1.2.2), and the LightGBM (Light Gradient Boosting Machine) classifier was used to estimate the feature weights. Multiple machine learning algorithm model parameters were tried. CatBoost (Categorical Boosting) is a machine learning algorithm based on gradient-boosted decision trees, and its significant advantage lies in its ability to precisely handle the interactions between features while minimizing overfitting to ensure excellent prediction performance. In this study, Catboost was used to build a classification prediction model for five tumors in the PanCanAtlas cohort. Model performance was evaluated using metrics such as Accuracy, AUC, Recall, Precision, F1 score, Cohen's Kappa coefficient, and Matthews correlation coefficient (MCC). The PanCanAtlas cohort was divided into a training set of 1155 cases and a test set of 496 cases at a ratio of 7:3. The create_model method and default parameters in the Python library PyCaret (v3.0.4) were used to build the model, and 10-fold cross-validation was used for model training and evaluation. The overall data analysis and processing flow is asFigure 1 as shown

[0079] 4.1. Comparison of model performance established by different machine learning algorithms and the number of methylation sites:

[0080] Models were constructed using methylation sites common to different types of chips. To construct the best-performing methylation model, this part of the study compared and analyzed the prediction performance of various machine learning algorithms commonly used in scientific research. The results are shown in the appendix Figure 2As shown, the CatBoost algorithm has the highest prediction accuracy. Therefore, the classification prediction model constructed by this method is finally selected. After screening 20 loci with FDR < 0.05 and the highest AUC, a total of 212 specific methylation sites are obtained, including: cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg24948406, cg17583946, cg07049592, cg00892567, cg20585676, cg13752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352, cg15402210, cg11849086, cg04271791, cg00646731, cg01561259, cg09929238, cg09646197, cg11335133, cg10089145, cg05343480, cg01853561, cg23286646, cg15316660, cg12713583, cg18495682, cg23830290, cg04863892, cg07145664, cg10532384, cg17153122, cg01698448, cg15085447, cg13661519, cg02803819, cg01781725, cg03763508, cg03239178, cg02720618, cg18585988, cg18030833, cg10191210, cg18953784, cg06786372, cg27230697, cg15661473, cg20518572, cg13544006, cg25309888, cg07920525, cg01612366, cg26333652, cg24869815, cg00290994, cg19480724, cg02238549, cg19468528, cg11246938, cg06099014, cg09263516, cg08557649, cg26728709, cg08198862, cg04933208, cg13905586, cg16277169, cg02776128, cg13030332, cg09053536, cg14434755, cg13356117, cg06395028, cg11423130, cg00549566, cg24281697, cg05787952,cg26292895, cg25228995, cg21341817, cg25408237, cg20390045, cg04255416, cg04499325, cg00159100, cg22149516, cg13717446, cg11021995, cg22111038, cg06159352, cg08453021, cg20425384, cg26347197, cg16561266, cg18116902, cg02531516, cg07807165, cg09143801, cg03113878, cg03346415, cg00237391, cg06328338, cg20180843, cg20562447, cg08382235, cg12292134, cg01697719, cg04003500, cg22289434, cg02900995, cg07002201, cg13941474, cg10069493, cg09806262, cg25114913, cg14327531, cg15202738, cg21929771, cg01489686, cg13409449, cg26884027, cg22216738, cg25541958, cg02500392, cg21411490, cg03088791, cg02774542, cg02374996, cg17923947, cg17749454, cg20631750, cg02014003, cg26033586, cg01610488, cg27039662, cg08545287, cg16266667, cg14294758, cg01610605, cg15245095, cg01588748, cg06664258, cg09571420, cg07559273, cg01718116, cg07028533, cg16910830, cg04075191, cg19939065, cg16831085, cg21981270, cg23900712, cg18104645, cg27239243, cg04143876, cg13525067, cg20754708, cg04194494, cg00675569, cg02506353, cg14517217, cg17853216, cg05305327, cg02125316, cg19221545, cg08922090, cg11272874, cg20749769,cg10725316, cg25123470, cg13551227, cg20564892, cg23346516, cg12769138, cg10397440, cg19702194, cg21574675, cg02928916, cg07833420, cg22782271, cg04278702, cg14774364, cg03241649, cg23395449, cg13851334, cg09607548, cg10135483, cg07739205, cg10767350, cg10982664, cg10096177, cg16488737, cg02983451, cg22060611, cg20309121, cg03882242, cg20227806, cg24051554, cg11743448, cg27265630, cg15153684, cg03493300, cg14817783, cg08144172, cg24596898, cg01692968.,

[0081] And further, the number and accuracy of 212 specific sites were evaluated. Finally, it was shown that the model containing 106 methylation sites had the highest diagnostic accuracy and a relatively small number of sites (such as Figure 3 , Figure 3Screening for the number of methylation sites included in the model: Among the total of 212 specific sites, the analysis showed that when 50% (i.e., 106 sites) were included, the model had high accuracy and a small number of sites. The 106 sites include: cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg17583946, cg07049592, cg00892567, cg13752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352, cg04271791, cg00646731, cg01561259, cg09646197, cg11335133, cg10089145, cg05343480, cg23830290, cg04863892, cg07145664, cg10532384, cg17153122, cg13661519, cg03763508, cg02720618, cg18585988, cg10191210, cg18953784, cg06786372, cg13544006, cg01612366, cg26333652, cg19480724, cg19468528, cg11246938, cg09263516, cg08557649, cg08198862, cg13905586, cg02776128, cg13030332, cg14434755, cg13356117, cg06395028, cg11423130, cg05787952, cg26292895, cg25228995, cg21341817, cg25408237, cg22149516, cg13717446, cg20425384, cg26347197, cg16561266, cg02531516, cg09143801, cg03113878, cg06328338, cg20562447, cg08382235, cg04003500, cg22289434, cg02900995, cg14327531, cg21929771, cg22216738, cg25541958, cg02500392, cg02374996, cg17923947, cg17749454, cg01610488, cg27039662, cg08545287, cg04075191, cg16831085, cg21981270,cg23900712, cg04143876, cg04194494, cg02506353, cg17853216, cg02125316, cg19221545, cg08922090, cg11272874, cg20749769, cg10725316, cg25123470, cg13551227, cg19702194, cg21574675, cg02928916, cg10767350, cg10982664, cg10096177, cg22060611, cg24051554, cg15153684.,

[0082] There are significant differences in the weights of different loci in the model (such as Figure 4 ), and there are obvious differences in the methylation levels of these 106 methylation sites in 5 types of tumors (such as Figure 5 ), and the biological functions of specifically highly or lowly methylated sites in different tumor types can be further explored.

[0083] 4.2. Performance verification of the methylation model

[0084] 4.2.1 Performance of the methylation model in the training set

[0085] The prediction model established by the CatBoost algorithm has excellent prediction performance in the PanCanAtlas training set. The AUCs are 0.995 for BLCA; 0.998 for CESC; 0.969 for ESCC; 0.991 for HNSC; 0.991 for LUSC. The confusion matrix shows that among 496 samples, only 6 samples have inconsistent tumor types and pathological results predicted by the model, and the overall prediction accuracy reaches 98.79% (490 / 496) (such as Figure 6 ).

[0086] 2.2 Performance of the methylation model in the validation set

[0087] The performance of the model was further verified with data from the external test set public database. The results show that the prediction accuracy in urothelial carcinoma is 94.12% (16 / 17), in cervical squamous cell carcinoma is 92.52% (99 / 107), in esophageal squamous cell carcinoma is 89.74% (35 / 39), in head and neck squamous cell carcinoma is 73.33% (88 / 120), and in lung squamous cell carcinoma is 94.44% (102 / 108). The overall accuracy reaches 86.96% (340 / 391) (such as Figure 7 ).

[0088] In addition, among the samples detected at the FUSCC external validation set of our center, methylation data of 4 samples, including 2 cases of HNSCC, 1 case of LUSC, and 1 case of CESC, did not pass the quality control and were not included in the subsequent analysis. Based on histological diagnosis, the overall validation accuracy of the cases in our center was 83.50% (86 / 103). Among them, the prediction accuracy in urothelial carcinoma was 95.24% (20 / 21), in cervical squamous cell carcinoma was 70.00% (14 / 20), in esophageal squamous cell carcinoma was 91.67% (22 / 24), in head and neck squamous cell carcinoma was 78.95% (15 / 19), and in lung squamous cell carcinoma was 78.95% (15 / 19) (as shown in Table 3) (as Figure 8 ). The prediction accuracy in 82 primary samples was 89.02% (73 / 82), and in 21 metastatic samples was 61.9% (13 / 21). The accuracy of metastatic samples was significantly lower than that of primary samples. The reason for the analysis was that among the 8 metastatic samples with incorrect predictions, the tumor cell content of 4 samples was ≤ 30%. In addition, there were 2 cases of lung squamous cell carcinoma with a history of esophageal cancer. The histological diagnosis of the lung lesions tended to be metastatic esophageal cancer, while this methylation prediction model diagnosed it as lung squamous cell carcinoma. Considering that the lung lesions in these two cases were solitary, the possibility of second primary lung squamous cell carcinoma could not be excluded and further verification was needed. By comparing and analyzing the prediction scores of correctly and incorrectly predicted samples, the results showed that the average score of correctly predicted samples was 0.80, which was significantly higher than the average score of incorrectly predicted samples, which was 0.51, suggesting that samples with low prediction scores should be interpreted with caution and comprehensive evaluation should be combined with clinical, pathological, and other information.

[0089] Table 3 Performance verification results of the methylation model in the test set

[0090]

[0091] Finally, the tissue origin prediction model constructed by combining 106 methylation sites screened by the present invention with the CatBoost machine learning algorithm can effectively predict the origin of squamous cell carcinoma and urothelial carcinoma. The prediction accuracy in actual clinical samples reached 83.5%, which can effectively meet the clinical diagnosis and treatment needs. The methylation detection technology adopted by the present invention is mature and stable, and the origin tracing model of the present invention is mutually verified by using public databases and local patient data in our country.

[0092] As mentioned above, it is only the preferred embodiment of the present invention, and there is no restriction on the present invention in any formal or substantial form. It should be pointed out that for those of ordinary skill in the art, without departing from the present invention, several improvements and supplements can still be made, and these improvements and supplements should also be regarded as the protection scope of the present invention. Those skilled in the art, without departing from the spirit and scope of the present invention, when making some equivalent changes such as minor modifications, decorations and evolutions by using the technical content disclosed above, are all equivalent embodiments of the present invention; at the same time, any equivalent changes, modifications and evolutions made to the above embodiments based on the substantial technology of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A squamous cell carcinoma tissue tracing method based on methylation chip and machine learning, characterized in that: The method includes obtaining DNA methylation data of squamous cell carcinoma and urothelial carcinoma samples, and dividing the data into a training set, a test set, and a validation set; preprocessing the DNA methylation data to screen out specific methylation sites; A model for classifying and predicting squamous cell carcinoma and urothelial carcinoma tissues was then constructed using a machine learning algorithm; the specific methylation sites included: cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg17583946, cg07049592, cg00892567, cg13752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352, cg04271791, cg00646731, cg01 561259, cg09646197, cg11335133, cg10089145, cg05343480, cg23830290, cg04863892, cg07145664, cg10532384, cg17153122, cg13661519, cg03763 508, cg02720618, cg18585988, cg10191210, cg18953784, cg06786372, cg1 3544006, cg01612366, cg26333652, cg19480724, cg19468528, cg11246938 , cg09263516, cg08557649, cg08198862, cg13905586, cg02776128, cg1303 0332, cg14434755, cg13356117, cg06395028, cg11423130, cg05787952, cg 26292895, cg25228995, cg21341817, cg25408237, cg22149516, cg1371744 6. cg20425384, cg26347197, cg16561266, cg02531516, cg09143801, cg031 13878, cg06328338, cg20562447, cg08382235, cg04003500, cg22289434, c g02900995, cg14327531, cg21929771, cg22216738, cg25541958, cg025003 92, cg02374996, cg17923947, cg17749454, cg01610488, cg27039662, cg08 545287, cg04075191, cg16831085, cg21981270, cg23900712, cg04143876,cg04194494, cg02506353, cg17853216, cg02125316, cg19221545, cg08922090, cg11272874, cg20749769, cg10725316, cg25123470, cg13551227, cg19702194, cg21574675, cg02928916, cg10767350, cg10982664, cg10096177, cg22060611, cg24051554, cg15153684, a total of 106 methylation sites. , 2. A squamous cell carcinoma tissue tracing method based on methylation chip and machine learning according to claim 1, characterized in that: The machine learning algorithm is the CatBoost algorithm.

3. A squamous cell carcinoma tissue tracing system based on methylation chip and machine learning, characterized in that: The system includes a model for classifying and predicting squamous cell carcinoma and urothelial carcinoma tissues. The model obtains DNA methylation data of squamous cell carcinoma and urothelial carcinoma samples and divides the data into a training set, a test set and a validation set; pre-processes the DNA methylation data to screen out specific methylation sites; and then constructs the model through a machine learning algorithm; the specific methylation sites include: cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg17583946, cg07049592, cg00892567, cg1 3752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352 , cg04271791, cg00646731, cg01561259, cg09646197, cg11335133, cg1008 9145, cg05343480, cg23830290, cg04863892, cg07145664, cg10532384, cg 17153122, cg13661519, cg03763508, cg02720618, cg18585988, cg1019121 0,cg18953784,cg06786372,cg13544006,cg01612366,cg26333652,cg194 80724, cg19468528, cg11246938, cg09263516, cg08557649, cg08198862, c g13905586, cg02776128, cg13030332, cg14434755, cg13356117, cg063950 28, cg11423130, cg05787952, cg26292895, cg25228995, cg21341817, cg25 408237, cg22149516, cg13717446, cg20425384, cg26347197, cg16561266, cg02531516, cg09143801, cg03113878, cg06328338, cg20562447, cg08382 235, cg04003500, cg22289434, cg02900995, cg14327531, cg21929771, cg2 2216738, cg25541958, cg02500392, cg02374996, cg17923947, cg17749454,cg01610488, cg27039662, cg08545287, cg04075191, cg16831085, cg21981270, cg23900712, cg 04143876, cg04194494, cg02506353, cg17853216, cg02125316, cg19221545, cg08922090, cg112 72874, cg20749769, cg10725316, cg25123470, cg13551227, cg19702194, cg21574675, cg02928916, cg10767350, cg10982664, cg10096177, cg22060611, cg24051554, cg15153684, a total of 106 methylation sites. , 4. A squamous cell carcinoma tissue tracing system based on methylation chip and machine learning according to claim 3, characterized in that: The machine learning algorithm is the CatBoost algorithm.

5. Use of a methylation marker in the preparation of a diagnostic reagent for primary lesions of squamous cell carcinoma and urothelial carcinoma tissue, characterized in that: The methylation markers are specific methylation sites, including: cg01084693, cg27313941, cg26687072, cg14849140, cg09343092, cg07823492, cg04904318, cg17583946, cg07049592, cg00892567, cg13752649, cg26634219, cg07485775, cg07852402, cg22660933, cg10936352, cg04271791, cg00646731, cg01561259, cg09646197, cg113351 33, cg10089145, cg05343480, cg23830290, cg04863892, cg07145664, cg10 532384, cg17153122, cg13661519, cg03763508, cg02720618, cg18585988, c g10191210, cg18953784, cg06786372, cg13544006, cg01612366, cg263336 52, cg19480724, cg19468528, cg11246938, cg09263516, cg08557649, cg081 98862, cg13905586, cg02776128, cg13030332, cg14434755, cg13356117, c g06395028, cg11423130, cg05787952, cg26292895, cg25228995, cg2134181 7, cg25408237, cg22149516, cg13717446, cg20425384, cg26347197, cg165 61266, cg02531516, cg09143801, cg03113878, cg06328338, cg20562447, cg 08382235, cg04003500, cg22289434, cg02900995, cg14327531, cg2192977 1. cg22216738, cg25541958, cg02500392, cg02374996, cg17923947, cg1774 9454, cg01610488, cg27039662, cg08545287, cg04075191, cg16831085, cg2 1981270, cg23900712, cg04143876, cg04194494, cg02506353, cg17853216,cg02125316, cg19221545, cg08922090, cg11272874, cg20749769, cg10725316, cg25123470, cg13551227, cg19702194, cg21574675, cg02928916, cg10767350, cg10982664, cg10096177, cg22060611, cg24051554, cg15153684, a total of 106 methylation sites. , 6. The use according to claim 5, characterized in that: The application includes preparing a diagnostic kit for diagnosing primary lesions of squamous cell carcinoma or urothelial carcinoma tissue.

7. The use according to claim 5, characterized in that: The application includes preparing a diagnostic kit for differential diagnosis of primary lesions of squamous cell carcinoma or urothelial carcinoma tissue.