A protein combination, related kit and system for prognosis stratification of patients with medullary thyroid carcinoma

By detecting specific protein combinations in patients with medullary thyroid carcinoma and using a random forest classifier model, we have addressed the issue of prognostic factors that existing assessment tools have failed to fully consider, enabling more accurate prediction of postoperative recurrence risk and personalized treatment guidance for medullary thyroid carcinoma (MTC).

CN120559239BActive Publication Date: 2026-03-27WESTLAKE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing prognostic assessment tools for MTC rely on the TNM staging system, but fail to fully incorporate important prognostic factors such as age, sex, genetics, and postoperative biochemical indicators, resulting in insufficient objectivity and accuracy in assessment, especially with uncertain universality in Asian populations.

Method used

Using a protein combinatorial and machine learning model, we constructed a kit and system for predicting the prognosis of medullary thyroid carcinoma by detecting the expression levels of 18 proteins, including SRI, PTPRM, and HSPB7, and combining them with a random forest classifier, and performed risk stratification.

Benefits of technology

It achieves a more objective and accurate prediction of the risk of recurrence after MTC surgery, and can be effectively applied in patient cohorts in multiple hospitals in China to guide personalized follow-up and treatment strategies and reduce the risk of recurrence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120559239B_ABST
    Figure CN120559239B_ABST
Patent Text Reader

Abstract

The present application relates to a protein combination for prognosis stratification of patients with medullary thyroid carcinoma, a related kit and a prediction system. The prediction system is based on a combination of 18 proteins and 2 clinical information, or based on a combination of 29 proteins. The prognosis stratification method of the present application is more objective than the grading method of the prior art, has stronger prediction ability, and has been proven effective in patient cohorts of multiple hospitals in China.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical diagnosis, in particular to a protein combination for prognosis stratification of medullary thyroid carcinoma patients, a related kit and a prediction system. BACKGROUND

[0002] Medullary thyroid carcinoma (MTC) is a rare neuroendocrine tumor caused by parafollicular C cells. The incidence of MTC accounts for only 1-5% of all thyroid cancers, but it accounts for about 13% of thyroid cancer-related deaths. MTC has the characteristics of invasiveness and high metastatic potential, with a median disease-specific survival of 8.6 years. Resistance to radioiodine therapy limits the choice of treatment options. The main treatment for MTC is surgical intervention, especially total thyroidectomy and bilateral central neck dissection. However, postoperative recurrence remains a major problem, with a reported reoperation rate of 16.3% and a median time to reoperation of 6.4 months. Disease recurrence can severely affect disease-free survival and quality of life, so effective risk stratification is necessary to predict prognosis and optimize long-term follow-up strategies to improve treatment efficacy.

[0003] Currently, the evaluation of postoperative prognosis of MTC mainly relies on the TNM staging system, which evaluates the maximum diameter of the primary tumor, extrathyroidal invasion, lymph node metastasis and distant metastasis. However, this system does not include other important prognostic factors such as age, gender, genetics, and postoperative calcitonin and carcinoembryonic antigen levels. Therefore, a more comprehensive tool that integrates these different factors is needed to accurately assess the prognosis risk of MTC patients. Xu et al. recently proposed the International MTC Grading System (IMTCGS), which includes indicators of proliferative activity, including mitotic index and / or Ki67 proliferation index, and tumor necrosis factor [1] . This grading system divides MTC into high-grade and low-grade, with higher-grade tumors having lower disease-specific survival and higher rates of local and distant recurrence. Although this system has been validated in multiple cohorts in Europe, the United States and Australia, its generalizability in Asian populations remains uncertain, and the system only uses pathological indicators, which are subjective and may vary in accuracy among clinicians of different levels of experience due to subjective experience.

[0004] Genomic and transcriptomic technologies have been widely used to study the prognosis of MTC. Among them, RET gene mutations play a key role in both hereditary and sporadic MTC. Hereditary MTC involves germline mutations in the RET gene, with different RET mutation sites corresponding to different disease invasiveness and risk; while for sporadic MTC, somatic RET M918T mutations are associated with poor prognosis. Notably, 10-20% of sporadic MTC cases lack known driver mutations, so other techniques are needed as a supplement.

[0005] Therefore, there is an urgent need to develop a more practical, convenient, and objective system and method for predicting the prognosis of medullary thyroid carcinoma. Summary of the Invention

[0006] To address the above technical problems, in one aspect, the present invention provides the use of a protein combination in the preparation of a kit for predicting the prognosis of medullary thyroid carcinoma, said protein combination comprising the following:

[0007] SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK,

[0008] The kit contains reagents for detecting the expression levels of the protein combination.

[0009] In a specific embodiment, the expression levels of the protein combination are detected by liquid chromatography and mass spectrometry.

[0010] On the other hand, the present invention provides a kit for predicting the prognosis of medullary thyroid carcinoma, the kit containing reagents for detecting the expression levels of a protein combination comprising the following:

[0011] SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK.

[0012] On the other hand, the present invention provides a system for prognostic stratification of medullary thyroid carcinoma, the system comprising:

[0013] A subsystem for determining the expression levels of a protein combination, which comprises the following:

[0014] SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK;

[0015] A data storage module is used to store information on the expression levels of the protein combination, lymph node metastasis, and maximum tumor diameter.

[0016] The data processing module processes the information from the data storage module, and uses a random forest classifier to evaluate the fitness of feature subsets as the main performance indicator.

[0017] The output module outputs the data processing results from the data processing module as the model's predicted recurrence probability, ranging from 0 to 1, where a value greater than 0.5 indicates a high-risk group and a value less than or equal to 0.5 indicates a low-risk group.

[0018] In another aspect, the present invention provides the use of a protein combination in the preparation of a kit for predicting the prognosis of medullary thyroid carcinoma, said protein combination comprising the following:

[0019] HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1,

[0020] The kit contains reagents for detecting the expression levels of the protein combination.

[0021] In a specific embodiment, the expression levels of the protein combination are detected by liquid chromatography and mass spectrometry.

[0022] On the other hand, the present invention provides a kit for predicting the prognosis of medullary thyroid carcinoma, the kit containing reagents for detecting the expression levels of a protein combination comprising the following:

[0023] HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1.

[0024] On the other hand, the present invention provides a system for prognostic stratification of medullary thyroid carcinoma, the system comprising:

[0025] A subsystem for detecting the expression levels of the protein combination, which comprises the following:

[0026] HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1;

[0027] A data storage module for storing information on the expression levels of the protein combination;

[0028] The data processing module processes the information from the data storage module, and uses a random forest classifier to evaluate the fitness of feature subsets as the main performance indicator.

[0029] The output module outputs the data processing results from the data processing module as the model's predicted recurrence probability, ranging from 0 to 1, where a value greater than 0.5 indicates a high-risk group and a value less than or equal to 0.5 indicates a low-risk group.

[0030] Beneficial effects

[0031] 1. This application proposes a multimodal classifier, which includes a combination of 18 proteins and 2 clinical information, or a combination of 29 proteins, capable of predicting the risk of recurrence after surgery for medullary thyroid carcinoma, and classifying patients based on their risk of postoperative recurrence. Compared with previous stratification methods based on postoperative pathology, this method is more objective and has stronger predictive ability, and its effectiveness has been demonstrated in patient cohorts in multiple hospitals in China.

[0032] 2. In the queue design of this application, the six hospitals in the discovery set and the four hospitals in the test set are independent of each other. The classifier obtained from the discovery set remains effective in the independent test set, demonstrating the generality of the classifier.

[0033] 3. This invention allows patients with medullary thyroid carcinoma to be categorized into different risk levels based on their future recurrence risk after initial surgery. This effectively guides postoperative follow-up strategies and clinical medication, facilitating the development of personalized follow-up and treatment strategies by physicians. It should be noted that the recurrence mentioned in this invention only considers structural recurrence, i.e., the reappearance of MTC histological or radiological evidence after radical surgery.

[0034] 4. The two clinical indicators in the classifier of this application are simple and readily available, and are routine postoperative pathological information. Furthermore, protein can be detected using only a small amount of tissue sample (1 mg) removed during surgery for pathological analysis, without causing trauma to the patient beyond surgical treatment, making it convenient for practical application and promotion. Attached Figure Description

[0035] Figure 1 The flowchart of the model development for this application is shown.

[0036] Figure 2 The ROC curves of the random forest model are shown on the independent test set and the FUSCC dataset.

[0037] Figure 3 Kaplan-Meier survival curves for relapse-free survival (RFS) based on a comprehensive model with 20 features are shown in the independent test set and the FUSCC dataset.

[0038] Figure 4 This paper shows the ranking of the importance of 20 features in the integrated model based on Shapley additive explanations (SHAP) values.

[0039] Figure 5 The importance of model features is shown in the ranking. (A) Clinical model features; (B) Genetic model features; and (C) Proteomic model features are ranked by their SHAP values. Detailed Implementation

[0040] The modeling and verification processes of this application are described in detail below through specific embodiments to enable those skilled in the art to better understand this application. However, it should be understood that these specific implementation methods are not intended to limit the scope of this application.

[0041] the term

[0042] In this paper, "classifier" refers to the specific risk prediction model proposed in this application for the prognosis of medullary thyroid carcinoma.

[0043] The protein combinations involved in this application (combination 1 includes 18 proteins; combination 2 includes 29 proteins) have their numbers, search names, and English names in the Uniprot database as shown in Tables 1 and 2 below.

[0044] Table 1 18 protein combinations

[0045]

[0046]

[0047] Table 2 29 protein combinations

[0048]

[0049]

[0050]

[0051] Information on the biomarkers in Table 1-2 above can be found in the Uniprot database (https: / / www.uniprot.org).

[0052] Example 1:

[0053] Sample collection:

[0054] The discovery cohort contained 376 medullary thyroid carcinoma (MTC) samples from 6 independent hospitals. The testing cohort contained 105 MTC samples from 4 independent hospitals that did not overlap with the discovery cohort. All samples were evaluated by at least two pathologists and identified as medullary thyroid carcinoma.

[0055] Gene sequencing:

[0056] The sequencing library was developed using a next-generation gene assay suite from RigenBio to detect mutations in 28 genes associated with thyroid cancer (see Table 3 below for a gene list). DNA was extracted from 476 formalin-fixed and paraffin-embedded (FFPE) thyroid samples using a DNA extraction kit (RigenBio, China), excluding the 5 samples in the discovery dataset. The DNA extraction and sequencing protocols have been described in previously published articles. [2] This will not be elaborated upon here. In short, the extracted DNA undergoes multiplex amplification of the target region, followed by PCR amplification to incorporate unique dual-index adapters and Illumina sequencing adapters. After purification using magnetic beads, the indexed library is quantified using a Thermo Fisher Qubit fluorescence spectrometer and sequenced on an Illumina NovaSeq 6000 system, producing 150 bp paired-end reads.

[0057] The quality of the raw sequencing data was assessed using FastQC (version 0.11.9). Raw reads were preprocessed using Cutadapt (version 1.18) to remove aptamers and low-quality bases. The processed reads were then aligned to the hg19 human reference genome using Burrows-WheelerAligner software (version 0.7.17). Single nucleotide variants (SNVs) and insertions / deletions (InDels) were identified using VarScan2 (version 2.4.4), and variant annotation was performed using the Ensembl variant effect predictor (VEP) to assess potential impact.

[0058] Table 3. Detailed information on 28 gene testing combinations

[0059]

[0060]

[0061] Proteome sample preparation

[0062] The preparation method of FFPE tissue is as described in previously published articles. [3,4] In short, FFPE sections were dewaxed, rehydrated, and decrosslinked sequentially using heptane, three different concentrations (100%, 90%, and 75%) of ethanol, 100% water, and 100 mM Tris-HCl solution (pH 10.0). Then, with the aid of pressure cycling technology (PCT), the samples were lysed in a buffer containing 6 M urea, 2 M thiourea, 10 mM tris(2-carboxyethyl)phosphine, and 40 mM iodoacetamide. Trypsin and lysC protease were mixed and used for PCT digestion. Finally, the digested peptides were quenched with trifluoroacetic acid and desalted using a C18 column (Thermo Fisher Scientific).

[0063] DIA-MS mass spectrometry data acquisition and analysis

[0064] Inject the peptide sample with... The system (Brook Dalton GmbH, Germany) uses a self-made C18 separation column (size: 15cm × 75μm × 1.9μm). The sample was then separated by a 60-minute liquid chromatography (LC) gradient, with the concentration changing from 5% buffer B to 27% buffer B over 50 minutes, and then increasing to 40% buffer B over 10 minutes. Buffer A was an aqueous solution containing 0.1% formic acid, and buffer B was an acetonitrile solution containing 0.1% formic acid.

[0065] Peptide samples separated by liquid chromatography were analyzed using a trapped ion mobility spectrometry quadrupole time-of-flight mass spectrometry system (timsTOF Pro, Brook-Dalton, Germany). This instrument was equipped with a CaptiveSpray nanoparticle electrospray ion source, enabling secondary separation of peptides via ion mobility separation technology during the experiment. Parallel cumulative-serial fragmentation (PASEF) was performed in data-independent acquisition (DIA) mode. The dual TIMS analyzer had a cumulative and ramp time of 100 ms and a total cycle time of 1.17 seconds, comprising 14 PASEF scans, each with four two-dimensional separation windows of ion mobility-m / z. Ion mobility scans ranged from 0.6 to 1.6 Vs / cm². MS1 and MS2 acquisitions were performed in the m / z range of 100 to 1700 Th. Single-charged precursor ions were excluded.

[0066] The original DIA file was created by DIA-NN. [5,6] (Version 1.8.1) Reference thyroid-specific spectrogram library [7] Analysis revealed that the specific spectral library contained 12,000 proteins and 215,000 precursor ions. Variable modifications included methionine oxidation and N-terminal acetylation, while fixed modifications included cysteine ​​carbamoyl methylation. The peptide length range, precursor ion m / z range, and fragment ion m / z range were set to 6–30, 300–1800, and 200–1800, respectively. The false detection rate for both precursor ions and proteins was set to 1%. The "Irrelevant Run" and "Use Isotopes" options were selected. Protein inference was set to "Off." All other parameters remained at their default values.

[0067] Proteomics data quality control and preprocessing

[0068] To minimize potential biases during sample preparation and mass spectrometry acquisition, relapsed and non-relapsed samples were randomly assigned. Each batch included 15 tissue samples and one thyroid polypeptide mixture (as a quality control). Different samples from the same patient were used as bioreplicas. One sample from each batch was randomly selected as a technical replicate and acquired twice on the mass spectrometer to assess the stability of mass spectrometry quantification.

[0069] Missing values ​​in the protein matrix were determined using ridge regression. [8] and NAguideR[9] Software estimation. Using the R package sva

[10] The empirical Bayesian framework Combat in version 3.48.0 performs batch effect correction on the resulting protein matrix. Batch effects are corrected for based on different clinical centers and sample batches. Each pair of technical replicates is merged into a single sample by calculating the average protein abundance.

[0070] Example 2: Dataset Partitioning and Cross-Validation in Machine Learning

[0071] The dataset consists of a discovery set and two test sets. The two test sets are independent test sets from four independent medical centers (n=105) and test sets from published literature.

[11] The FUSCC dataset (n=93) was used. Of the 93 patients in the FUSCC dataset, 29 were already included in the discovery set. To avoid data duplication, overlapping patients were removed, resulting in a final FUSCC test set containing 64 patients. To optimize model performance, the discovery dataset was divided into five equal subsets for cross-validation. In each iteration, four subsets were selected for model training, and the remaining subset was used for model validation.

[0072] Feature selection and model building

[0073] First, proteins with missing values ​​exceeding 90% were removed, retaining 9380 proteins. Then, proteins significantly associated with prognosis were screened from the discovery dataset based on differential protein analysis (DEP) and coefficient of variation (CV) values. Differentially expressed proteins were defined as a |log2 fold change (FC)| > 0.25 between the SR (structural recurrence) and NR (recurrence-free) groups or the DSM (disease-specific death) and S (survival) groups, with a BH-corrected P-value < 0.05. DEPs were combined with the top 200 proteins by CV value, ultimately yielding 610 proteins with the most significant prognostic changes. This study collected 12 clinicopathological variables, including age, sex, genetics, tumor IMTCG grade, presence of Hashimoto's thyroiditis, multifocality, bilaterality, maximum tumor diameter, presence of papillary thyroid carcinoma, extrathyroidal invasion, lymph node metastasis, and vascular invasion. Furthermore, 15 gene mutation sites were detected, of which 9 mutation sites with a mutation rate > 1% were retained.

[0074] Subsequently, a genetic algorithm (GA) was used for feature selection; for details, please refer to previous studies. [12,13]In brief, the genetic algorithm is implemented using the `eaSimple` function in the Python package DEAP (version 1.4.1), with the following main parameter settings: crossover probability (cxpb) of 0.5, mutation probability (mutpb) of 0.2, and total number of iterations (ngen) of 400. Through iterative optimization, a subset of features that contribute significantly to model performance is retained.

[0075] Based on the optimal feature subset selected by GA, this application constructs three random forest (RF) models, using clinical, genomic, and proteomic features, respectively. Furthermore, a genetic algorithm is applied again to the 38 features across the three feature types to further reduce the number of features, ultimately resulting in a comprehensive model that includes 18 proteins and 2 clinical indicators.

[0076] The Random Forest (RF) model aggregates the predictions from all decision trees using majority voting to determine the final classification result. In each iteration, a Random Forest classifier (random_state=42) is used to evaluate the fitness of the feature subset as the primary performance metric. The classifier categorizes samples into high-risk and low-risk groups based on the contribution of the feature subset to the classification task. Finally, the model is optimized using 5-fold cross-validation (training to validation set ratio of 4:1) on the discovery dataset.

[0077] Establish and validate a machine learning model for predicting postoperative recurrence.

[0078] To predict the prognostic risk of MTC and develop individualized treatment and follow-up plans for patients, this application developed four machine learning models to predict the probability of structural recurrence after the initial surgery. Three of these models were built using clinical indicators, gene mutation information, and proteomic data, respectively, while the third model was a comprehensive model integrating the three different types of data.

[0079] Model construction consists of three stages: feature selection, model training and cross-validation, and model prediction (see above for the construction process). Figure 1 ).

[0080] Example 3: Validation on independent test datasets and FUSCC datasets

[0081] The aforementioned model was further validated on the independent test dataset (n=105) and the FUSCC dataset (n=64) to evaluate its generalization ability. Figure 2In the independent test set, the integrated model and the proteomics model achieved similar areas under the curve (AUCs) of 0.87 and 0.85, respectively, which were higher than the clinical model (AUC = 0.76), the gene model (AUC = 0.47), and the IMTCGS grading system (AUC = 0.51). In the FUSCC dataset, the proteomics model had the highest AUC (AUC = 0.78), followed by the integrated model (AUC = 0.77), the clinical model (AUC = 0.76), and the gene model (AUC = 0.53). It is noteworthy that the samples in the FUSCC dataset are fresh frozen samples, unlike the FFPE samples in the discovery dataset used to train the model. The similar performance of the models further demonstrates the robustness of the proposed model. Furthermore, the integrated model has 31% fewer features than the protein model; therefore, given similar AUCs, the integrated model is considered to perform better.

[0082] Based on protein expression and clinical indicators in the comprehensive model, patients in the two test sets were divided into high-risk and low-risk groups, and the difference in recurrence risk was as follows: Figure 3 As shown (P = 1e-4 and P = 0.046).

[0083] Figure 4 This diagram illustrates the importance of 20 features, including 18 proteins and 2 clinical indicators, in the comprehensive model selected as the optimal biomarker combination by this invention. The horizontal axis represents the importance of the features, and the vertical axis represents each feature, with importance decreasing from top to bottom. Orange represents clinical features, and green represents proteins (Uniprot database protein ID + Uniprot database search name).

[0084] Figure 5 The diagrams show the importance of features in clinical models, gene models, and protein models. The horizontal axis represents the importance of features, and the vertical axis represents each feature, with importance decreasing from top to bottom.

[0085] References:

[0086] [1 ]

[0087] [2]LI L,ZHANG L,JIANG W,et al.Mitochondrial Proteome DefinedMolecular Pathological Characteristics of Oncocytic Thyroid Tumors[J].EndocrPathol,2024,35(4):442-52.

[0088] [3]CAI X,XUE Z,WU C,et al.High-throughput proteomic samplepreparation using pressure cycling technology[J].Nat Protoc,2022,17(10):2307-25.

[0089] [4]ZHU Y,WEISS T,ZHANG Q,et al.High-throughput proteomic analysis ofFFPE tissue samples facilitates tumor stratification[J].Mol Oncol,2019,13(11):2305-28.

[0090] [5]DEMICHEV V,MESSNER C B,VERNARDIS S I,et al.DIA-NN:neural networksand interference correction enable deep proteome coverage in high throughput[J].Nat Methods,2020,17(1):41-4.

[0091] [6]DEMICHEV V,SZYRWIEL L,YU F,et al.dia-PASEF data analysis usingFragPipe and DIA-NN for deep proteomics of low sample amounts[J].Nat Commun,2022,13(1):3944.

[0092] [7]LI L,JIANG W,WEI W,et al.Comprehensive Mass Spectral Libraries ofHuman Thyroid Tissues and Cells[J].Sci Data,2024,11(1):1448.

[0093] [8]WEI R,WANG J,JIA E,et al.GSimp:A Gibbs sampler based left-censoredmissing value imputation approach for metabolomics studies[J].PLoS ComputBiol,2018,14(1):e1005973.

[0094] [9]WANG S,LI W,HU L,et al.NAguideR:performing and prioritizingmissing value imputations for consistent bottom-up proteomic analyses[J].Nucleic Acids Res,2020,48(14):e83.

[0095]

[10] LEEK J T,JOHNSON W E,PARKER H S,et al.The sva package forremoving batch effects and other unwanted variation in high-throughputexperiments[J].Bioinformatics,2012,28(6):882-3.

[0096]

[11] SHI X,SUN Y,SHEN C,et al.Integrated proteogenomiccharacterization of medullary thyroid carcinoma[J].Cell Discov,2022,8(1):120.

[0097]

[12] SUN Y,SELVARAJAN S,ZANG Z,et al.Artificial intelligence definesprotein-based classification of thyroid nodules[J].Cell Discov,2022,8(1):85.

[0098]

[13] ZANG Z,CHENG S,XIAH,et al.DMT-EV:An Explainable Deep Network forDimension Reduction[J].IEEE Trans Vis Comput Graph,2024,30(3):1710-27.

Claims

1. A system for prognostic stratification of medullary thyroid carcinoma, the system comprising: A subsystem for determining the expression levels of a protein combination, which comprises the following: SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK; A data storage module is used to store information on the expression levels of the protein combination, lymph node metastasis, and maximum tumor diameter. The data processing module processes the information from the data storage module, and uses a random forest classifier to evaluate the fitness of feature subsets as the main performance indicator. The output module outputs the data processing results from the data processing module as the model's predicted recurrence probability, ranging from 0 to 1, where a value greater than 0.5 indicates a high-risk group and a value less than or equal to 0.5 indicates a low-risk group.