Protein combination for prognosis stratification of thyroid myeloid cancer patient, related kit and system

By detecting the expression of 18 proteins in patients with medullary thyroid carcinoma and combining machine learning models, the subjectivity and incompleteness of existing evaluation methods are solved, and more accurate prognostic risk assessment and personalized treatment guidance are achieved.

CN120559239AActive Publication Date: 2025-08-29WESTLAKE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510700437.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-29
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing prognostic evaluation methods for medullary thyroid carcinoma rely on the TNM staging system and fail to fully include important prognostic factors, such as age, gender, genetics and postoperative biochemical indicators, resulting in inadequate assessment and subjectiveness, especially in uncertain generality among Asian populations.

Method used

A protein combination and machine learning model was used to detect the expression of 18 proteins such as SRI, PTPRM, and HSPB7, and combine it with a random forest classifier to construct a kit and system to predict the prognosis of medullary thyroid carcinoma for risk stratification.

Benefits of technology

A more objective and accurate prediction of postoperative recurrence risk of medullary thyroid cancer can be effectively applied in different populations, guide personalized follow-up and treatment strategies, and reduce the risk of recurrence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120559239A_ABST
    Figure CN120559239A_ABST
Patent Text Reader

Abstract

The invention relates to a protein combination for prognosis stratification of a thyroid myeloid cancer patient, a related kit and a prediction system. The prediction system is based on a combination of 18 proteins and 2 clinical information, or based on a combination of 29 proteins. Compared with a grading mode in the prior art, the prognosis layering mode is more objective, the prediction ability is higher, and effectiveness is proved in patient queues of multiple hospitals in China.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical diagnosis, and specifically to a protein combination, related kits, and prediction system for prognostic stratification of medullary thyroid cancer patients. Background Art

[0002] Medullary thyroid carcinoma (MTC) is a rare neuroendocrine tumor arising from parafollicular C cells. The incidence of MTC accounts for only 1–5% of all thyroid cancers but is responsible for approximately 13% of thyroid cancer-related deaths. MTC is characterized by aggressiveness and high metastatic potential, with a median disease-specific survival of 8.6 years. Resistance to radioiodine therapy limits treatment options. The primary treatment for MTC is surgical intervention, specifically total thyroidectomy and bilateral central neck lymph node dissection. However, postoperative recurrence remains a significant problem, with a reported reoperation rate of 16.3% and a median time to reoperation of 6.4 months. Disease recurrence can significantly impact disease-free survival and quality of life, necessitating effective risk stratification to predict prognosis and optimize long-term follow-up strategies to improve efficacy.

[0003] The current assessment of the prognosis of MTC after surgery mainly relies on the TNM staging system, which evaluates the maximum diameter of the primary tumor, extrathyroidal invasion, lymph node metastasis, and distant metastasis. However, this system does not incorporate other important prognostic factors such as age, gender, genetics, and postoperative calcitonin and carcinoembryonic antigen levels. Therefore, a more comprehensive tool that can integrate these different factors is needed to accurately assess the prognostic risk of MTC patients. Xu et al. recently proposed the International MTC Grading System (IMTCGS), which incorporates proliferative activity indicators, including mitotic index and / or Ki67 proliferation index, as well as tumor necrosis. [1] This grading system divides MTC into high-grade and low-grade tumors, with high-grade tumors associated with lower disease-specific survival and higher rates of local and distant recurrence. Although the system has been validated in multiple cohorts in Europe, the United States, and Australia, its universality in Asian populations remains uncertain. Furthermore, the system uses only pathological indicators and is somewhat subjective, and the accuracy of assessments may vary among clinicians of different experience levels due to their subjective experience.

[0004] Genomic and transcriptomic technologies have been widely used to study the prognosis of MTC. RET gene mutations play a key role in both hereditary and sporadic MTC. Hereditary MTC involves germline mutations in the RET gene, with different RET mutation sites corresponding to different disease aggressiveness and risks. In sporadic MTC, the somatic RET M918T mutation is associated with a poor prognosis. Notably, 10-20% of sporadic MTC cases lack known driver mutations, necessitating additional technologies to supplement this.

[0005] Therefore, there is an urgent need to develop a more practical, simple, and objective system and method for predicting the prognosis of medullary thyroid cancer. Summary of the Invention

[0006] In order to solve the above technical problems, on the one hand, the present invention provides a use of a protein combination in preparing a kit for predicting the prognosis of medullary thyroid cancer, wherein the protein combination is composed of:

[0007] SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK,

[0008] Wherein, the kit contains a reagent for detecting the expression level of the protein combination.

[0009] In a specific embodiment, the expression level of the protein combination is detected by liquid chromatography and mass spectrometry.

[0010] In another aspect, the present invention provides a kit for predicting the prognosis of medullary thyroid carcinoma, the kit comprising a reagent for detecting the expression level of a protein combination, wherein the protein combination is composed of:

[0011] SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK.

[0012] In another aspect, the present invention provides a system for prognostic stratification of medullary thyroid carcinoma, the system comprising:

[0013] A subsystem for determining the expression level of a protein combination, wherein the protein combination is composed of:

[0014] SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK;

[0015] a data storage module, which is used to store information on the expression level of the protein combination, lymph node metastasis, and maximum tumor diameter;

[0016] A data processing module processes the information of the data storage module, wherein the fitness of the feature subset is evaluated using a random forest classifier as the main performance indicator;

[0017] The output module outputs the data processing results of the data processing module as the recurrence probability predicted by the model, ranging from 0 to 1, where greater than 0.5 indicates a high-risk group and less than or equal to 0.5 indicates a low-risk group.

[0018] In another aspect, the present invention provides a use of a protein combination in preparing a kit for predicting the prognosis of medullary thyroid cancer, wherein the protein combination consists of:

[0019] HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1,

[0020] Wherein, the kit contains a reagent for detecting the expression level of the protein combination.

[0021] In a specific embodiment, the expression level of the protein combination is detected by liquid chromatography and mass spectrometry.

[0022] In another aspect, the present invention provides a kit for predicting the prognosis of medullary thyroid carcinoma, the kit comprising a reagent for detecting the expression level of a protein combination, wherein the protein combination is composed of:

[0023] HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1.

[0024] In another aspect, the present invention provides a system for prognostic stratification of medullary thyroid carcinoma, the system comprising:

[0025] A subsystem for detecting the expression level of the protein combination, wherein the protein combination is composed of:

[0026] HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1;

[0027] A data storage module, which is used to store information on the expression level of the protein combination;

[0028] A data processing module processes the information of the data storage module, wherein the fitness of the feature subset is evaluated using a random forest classifier as the main performance indicator;

[0029] The output module outputs the data processing results of the data processing module as the recurrence probability predicted by the model, ranging from 0 to 1, where greater than 0.5 indicates a high-risk group and less than or equal to 0.5 indicates a low-risk group.

[0030] Beneficial effects

[0031] This application proposes a multimodal classifier, comprising a combination of 18 proteins and two clinical data points, or a combination of 29 proteins, that can predict the risk of postoperative recurrence of medullary thyroid carcinoma and stratify patients according to their risk of recurrence. Compared to existing stratification methods based on postoperative pathology, this method is more objective and has stronger predictive power, and has demonstrated effectiveness in patient cohorts from multiple hospitals in China.

[0032] 2. In the cohort design of this application, the six hospitals in the discovery set and the four hospitals in the test set are independent of each other. The classifier screened from the discovery set is still effective in the independent test set, demonstrating the versatility of the classifier.

[0033] 3. After initial surgery, patients with medullary thyroid carcinoma can be divided into different risk levels based on their future recurrence risk. This can effectively guide postoperative follow-up strategies and clinical medication, facilitating physicians to develop personalized postoperative follow-up and treatment strategies. It should be noted that the recurrence mentioned in this invention only considers structural recurrence, i.e., the reappearance of histological or radiological evidence of MTC after radical surgery.

[0034] 4. The two clinical indicators used in the classifier are simple and easy to obtain, representing routine postoperative pathological information. Furthermore, protein detection can be performed using a small amount of intraoperatively removed tissue sample for pathological analysis (1 mg of tissue), without causing any trauma to the patient beyond the surgical procedure, facilitating practical application and widespread adoption. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A flowchart showing the model development of this application.

[0036] Figure 2 Shown are the ROC curves of the random forest model in the independent test set and the FUSCC dataset.

[0037] Figure 3 The Kaplan-Meier survival curves of recurrence-free survival (RFS) in the independent test set and the FUSCC dataset are shown.

[0038] Figure 4 Shown is a ranking of the importance of 20 features in the integrated model based on Shapley additive explanations (SHAP) values.

[0039] Figure 5 The importance of model features is ranked by SHAP value. (A) Clinical model features; (B) Gene model features; and (C) Proteomic model features. DETAILED DESCRIPTION

[0040] The modeling process and verification process of the present application are described in detail below through specific examples so that those skilled in the art can understand the present application more clearly. However, it should be understood that these specific implementations are not intended to limit the scope of the present application.

[0041] the term

[0042] The “classifier” herein refers to the risk prediction model specifically for the prognosis of medullary thyroid cancer proposed in this application.

[0043] The numbers, search names and English names of the two protein combinations involved in this application (combination 1 includes 18 species; combination 2 includes 29 species) in the Uniprot database are shown in Tables 1 and 2 below, respectively.

[0044] Table 1 18 protein combinations

[0045]

[0046]

[0047] Table 2 Combinations of 29 proteins

[0048]

[0049]

[0050]

[0051] The information of the biomarkers in Tables 1-2 above can be found in the Uniprot database (https: / / www.uniprot.org).

[0052] Example 1:

[0053] Sample collection:

[0054] The discovery cohort consisted of 376 MTC samples from six independent hospitals. The test cohort consisted of 105 MTC samples from four independent hospitals that were not duplicates of the discovery cohort. All samples were evaluated and confirmed as medullary thyroid carcinoma by at least two pathologists.

[0055] Gene sequencing:

[0056] The sequencing library was prepared using a next-generation gene panel developed by RigenBio to detect mutations in 28 genes associated with thyroid cancer (see Table 3 below for a list of genes). DNA was extracted from 476 formalin-fixed and paraffin-embedded (FFPE) thyroid samples using a DNA extraction kit (Rigen Biotechnology Co., Ltd., China), excluding the five samples in the discovery dataset. The DNA extraction and sequencing protocols have been described in previously published articles. [2] The process is described in detail here. Briefly, the extracted DNA undergoes multiplex amplification of the target region, followed by PCR amplification to incorporate unique dual-index adapters and Illumina sequencing adapters. After purification using magnetic beads, the indexed library is quantified using a Thermo Fisher Qubit fluorometer and sequenced on an Illumina NovaSeq 6000 system, generating 150bp paired-end reads.

[0057] The quality of raw sequencing data was assessed using FastQC (version 0.11.9). Raw reads were preprocessed using Cutadapt (version 1.18) to remove adapters and low-quality bases. The processed reads were then aligned to the hg19 human reference genome using Burrows-WheelerAligner software (version 0.7.17). Single nucleotide variants (SNVs) and insertions / deletions (InDels) were identified using VarScan2 (version 2.4.4), and variant annotation was performed using the Ensembl variant effect predictor (VEP) to assess potential effects.

[0058] Table 3 Detailed information of the 28-gene panel

[0059]

[0060]

[0061] Proteomic sample preparation

[0062] FFPE tissues were prepared as previously described. [3,4] . Briefly, FFPE sections were dewaxed, rehydrated, and decrosslinked using heptane, three different concentrations of ethanol (100%, 90%, and 75%), 100% water, and 100mM Tris-HCl solution (pH = 10.0). The samples were then lysed in a buffer containing 6M urea, 2M thiourea, 10mM tris(2-carboxyethyl)phosphine, and 40mM iodoacetamide with the assistance of pressure cycling technology (PCT). Trypsin and lysC protease were mixed and used for PCT digestion. Finally, the digested peptides were quenched with trifluoroacetic acid and desalted using a C18 column (Thermo Fisher Scientific, USA).

[0063] DIA-MS mass spectrometry data acquisition and analysis

[0064] The peptide sample was injected into the The system (Bruker Daltonics, Germany) was packed with a homemade C18 separation column (specifications: 15 cm × 75 μm × 1.9 μm, The sample was then separated by liquid chromatography (LC) gradient separation over 60 minutes, with a concentration gradient from 5% buffer B to 27% buffer B over 50 minutes, and then increased to 40% buffer B over 10 minutes. Buffer A was an aqueous solution containing 0.1% formic acid, and buffer B was an acetonitrile solution containing 0.1% formic acid.

[0065] Peptide samples separated by liquid chromatography were analyzed using a trapped ion mobility spectrometry quadrupole time-of-flight mass spectrometry system (timsTOF Pro, Bruker Daltonics, Germany). This instrument, equipped with a CaptiveSpray nanoflow electrospray ionization source, achieved secondary separation of peptides via ion mobility separation. Parallel accumulation–serial fragmentation (PASEF) was performed in data-independent acquisition (DIA) mode. The dual TIMS analyzer had an accumulation and ramp time of 100 milliseconds, resulting in a total cycle time of 1.17 seconds, consisting of 14 PASEF scans, each with four ion mobility-m / z two-dimensional separation windows. The ion mobility scan range was 0.6 to 1.6 Vs / cm². MS1 and MS2 acquisitions were performed over the m / z range of 100 to 1700 Th. Singly charged precursor ions were excluded.

[0066] DIA original files by DIA-NN [5,6] (Version 1.8.1) Reference thyroid-specific spectrum library [7] The specific spectral library for analysis contained 12,000 proteins and 215,000 precursor ions. Parameters were set for variable modifications including methionine oxidation and N-terminal acetylation, and fixed modifications including cysteine ​​carbamidomethylation. The peptide length range, precursor ion m / z range, and fragment ion m / z range were set to 6-30, 300-1800, and 200-1800, respectively. The false discovery rate for both precursor ions and proteins was set to 1%. The "Irrelevant Run" and "Use Isotopes" options were selected. Protein inference was set to "Off." All other parameters were kept to their default values.

[0067] Proteomic data quality control and preprocessing

[0068] To minimize potential bias during sample preparation and mass spectrometry acquisition, relapse and nonrelapse samples were randomly assigned. Each batch included 15 tissue samples and a thyroid peptide pool sample (as a quality control). Different samples from the same patient served as biological replicates. One sample from each batch was randomly selected as a technical replicate and acquired twice on the mass spectrometer to assess the stability of mass spectrometry quantification.

[0069] Missing values ​​in the protein matrix are calculated using ridge regression. [8] and NAguideR[9] Software estimation. Using the R package sva

[10] The resulting protein matrix was corrected for batch effects using the empirical Bayesian framework Combat (version 3.48.0). Batch effects were corrected for different clinical centers and sample batches. Each pair of technical replicates was combined into a single sample by calculating the mean protein abundance.

[0070] Example 2: Dataset Partitioning and Cross-Validation in Machine Learning

[0071] The dataset includes a discovery set and two test sets. The two test sets are an independent test set (n=105) from four independent medical centers and a test set from published literature.

[11] The FUSCC dataset (n = 93) was used. Of the 93 patients in the FUSCC dataset, 29 were already included in the discovery set. To avoid data duplication, overlapping patients were removed, resulting in a final FUSCC test set of 64 patients. To optimize model performance, the discovery dataset was evenly divided into five subsets for cross-validation. In each iteration, four subsets were selected for model training, and the remaining subset was used for model validation.

[0072] Feature selection and model building

[0073] First, proteins with more than 90% missing values ​​were removed, retaining 9,380 proteins. Subsequently, proteins significantly associated with prognosis were screened in the discovery dataset based on differential protein analysis and coefficient of variation (CV) values. Differentially expressed proteins were defined as having a |log2 fold change (FC)| > 0.25 between the SR (structural recurrence) and NR (no recurrence) groups or between the DSM (disease-specific mortality) and S (survival) groups, with a BH-corrected P value < 0.05. The top 200 proteins with the highest DEPs and CV values ​​were merged, ultimately resulting in a list of 610 proteins with the most significant prognostic dynamics. A total of 12 clinicopathological variables were collected for this study, including age, sex, genetics, tumor IMTCGS grade, presence of Hashimoto's thyroiditis, multifocality, bilaterality, maximum tumor diameter, presence of papillary thyroid carcinoma, extrathyroidal invasion, lymph node metastasis, and vascular invasion. In addition, 15 gene mutation sites were detected, of which 9 with a mutation rate > 1% were retained.

[0074] Subsequently, a genetic algorithm (GA) was used for feature selection. The specific method can be found in previous studies. [12,13]Briefly, the genetic algorithm was implemented using the eaSimple function in the Python package DEAP (version 1.4.1). The main parameters were set as follows: crossover probability (cxpb) was 0.5, mutation probability (mutpb) was 0.2, and the total number of iterations (ngen) was 400. Through iterative optimization, the feature subset that contributed most to the model performance was retained.

[0075] Based on the optimal feature subsets selected by GA, this application constructed three random forest (RF) models, using clinical, genomic, and proteomic features, respectively. Furthermore, the GA was applied again to 38 features of each of the three feature types to further reduce the number of features, ultimately resulting in a comprehensive model that included 18 proteins and two clinical indicators.

[0076] The RF model aggregates the predictions of all decision trees using majority voting to determine the final classification. At each iteration, a random forest classifier (RandomForestClassifier, random_state = 42) is used to evaluate the fitness of the feature subset as the primary performance metric. The classifier categorizes samples into high-risk and low-risk groups based on the contribution of the feature subset to the classification task. Finally, the model is optimized using 5-fold cross-validation (with a 4:1 training set to validation set ratio) on the discovery dataset.

[0077] Establishment and validation of a machine learning model for predicting postoperative recurrence

[0078] To predict the prognosis risk of MTC and develop personalized treatment and follow-up plans for patients, this application developed four machine learning models to predict the probability of structural recurrence after initial surgery. Three of the models were built using clinical indicators, gene mutation information, and proteomic data, respectively, and one model was a comprehensive model that integrated these three different data types.

[0079] Model construction includes three stages: feature selection, model training and cross-validation, and model prediction (see the above construction process for details). Figure 1 ).

[0080] Example 3: Verification on an independent test dataset and the FUSCC dataset

[0081] The aforementioned model was further validated on an independent test dataset (n=105) and a FUSCC dataset (n=64) to evaluate its generalization ability ( Figure 2). In the independent test set, the comprehensive model and the proteomic model obtained similar areas under the curve (AUCs) of 0.87 and 0.85, which were higher than the clinical model (AUC = 0.76), the gene model (AUC = 0.47) and the IMTCGS grading system (AUC = 0.51). In the FUSCC dataset, the proteomic model had the highest AUC (AUC = 0.78), followed by the comprehensive model (AUC = 0.77), the clinical model (AUC = 0.76) and the gene model (AUC = 0.53). It is worth noting that the samples in the FUSCC dataset are fresh frozen samples, which are different from the FFPE samples in the discovery dataset used to train the model. The similar performance of the models further demonstrates the robustness of the model of the present application. In addition, the number of features of the comprehensive model is 31% less than that of the protein model, so when the AUC is similar, the comprehensive model is considered to be better.

[0082] According to the protein expression and clinical indicators in the comprehensive model, the patients in the two test sets were divided into high-risk and low-risk groups, and the difference in their recurrence risk was as follows Figure 3 As shown (P = 1e-4 and P = 0.046).

[0083] Figure 4 This figure shows the importance of 20 features, including 18 proteins and 2 clinical indicators, in the comprehensive model screened as the optimal marker combination by the present invention. The horizontal axis represents the importance of the features, and the vertical axis represents each feature, with importance gradually decreasing from top to bottom. Orange represents clinical features, and green represents proteins (Uniprot database protein number + Uniprot database search name).

[0084] Figure 5 Schematic diagrams of the importance of features in clinical models, gene models, and protein models are shown respectively. The horizontal axis represents the importance of the feature, and the vertical axis represents each feature, with the importance gradually decreasing from top to bottom.

[0085] References:

[0086] [1 ]

[0087] [2]LI L,ZHANG L,JIANG W,et al.Mitochondrial Proteome DefinedMolecular Pathological Characteristics of Oncocytic Thyroid Tumors[J].EndocrPathol,2024,35(4):442-52.

[0088] [3]CAI X,XUE Z,WU C,et al.High-throughput proteomic samplepreparation using pressure cycling technology[J].Nat Protoc,2022,17(10):2307-25.

[0089] [4]ZHU Y,WEISS T,ZHANG Q,et al.High-throughput proteomic analysis ofFFPE tissue samples facilitates tumor stratification[J].Mol Oncol,2019,13(11):2305-28.

[0090] [5]DEMICHEV V,MESSNER C B,VERNARDIS S I,et al.DIA-NN:neural networksand interference correction enable deep proteome coverage in high throughput[J].Nat Methods,2020,17(1):41-4.

[0091] [6]DEMICHEV V,SZYRWIEL L,YU F,et al.dia-PASEF data analysis usingFragPipe and DIA-NN for deep proteomics of low sample amounts[J].Nat Commun,2022,13(1):3944.

[0092] [7]LI L,JIANG W,WEI W,et al.Comprehensive Mass Spectral Libraries ofHuman Thyroid Tissues and Cells[J].Sci Data,2024,11(1):1448.

[0093] [8]WEI R,WANG J,JIA E,et al.GSimp:A Gibbs sampler based left-censoredmissing value imputation approach for metabolomics studies[J].PLoS ComputBiol,2018,14(1):e1005973.

[0094] [9]WANG S,LI W,HU L,et al.NAguideR:performing and prioritizingmissing value imputations for consistent bottom-up proteomic analyses[J].Nucleic Acids Res,2020,48(14):e83.

[0095]

[10] LEEK J T,JOHNSON W E,PARKER H S,et al.The sva package forremoving batch effects and other unwanted variation in high-throughputexperiments[J].Bioinformatics,2012,28(6):882-3.

[0096]

[11] SHI X,SUN Y,SHEN C,et al.Integrated proteogenomiccharacterization of medullary thyroid carcinoma[J].Cell Discov,2022,8(1):120.

[0097]

[12] SUN Y,SELVARAJAN S,ZANG Z,et al.Artificial intelligence definesprotein-based classification of thyroid nodules[J].Cell Discov,2022,8(1):85.

[0098]

[13] ZANG Z,CHENG S,XIAH,et al.DMT-EV:An Explainable Deep Network forDimension Reduction[J].IEEE Trans Vis Comput Graph,2024,30(3):1710-27.

Claims

1. Use of a protein combination in preparing a kit for predicting the prognosis of medullary thyroid carcinoma, the protein combination comprising: SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK, Wherein, the kit contains a reagent for detecting the expression level of the protein combination.

2. The use according to claim 1, wherein The expression level of the protein combination is detected by liquid chromatography and mass spectrometry.

3. A kit for predicting the prognosis of medullary thyroid carcinoma, comprising a reagent for detecting the expression level of a protein combination, wherein the protein combination is composed of: SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK.

4. A system for prognostic stratification of medullary thyroid carcinoma, the system comprising: A subsystem for determining the expression level of a protein combination, wherein the protein combination is composed of: SRI, PTPRM, HSPB7, MELTF, PALS1, GDPD1, COL10A1, VGF, CBX4, ITIH1, TUBB2B, OLFM3, TMC4, PAFAH1B3, LAMA5, SELENOI, SCUBE3, PENK; A data storage module, which is used to store information on the protein combination expression level, lymph node metastasis, and maximum tumor diameter; A data processing module processes the information of the data storage module, wherein the fitness of the feature subset is evaluated using a random forest classifier as the main performance indicator; The output module outputs the data processing results of the data processing module as the recurrence probability predicted by the model, ranging from 0 to 1, where greater than 0.5 indicates a high-risk group and less than or equal to 0.5 indicates a low-risk group.

5. Use of a protein combination in preparing a kit for predicting the prognosis of medullary thyroid carcinoma, the protein combination comprising: HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1, Wherein, the kit contains a reagent for detecting the expression level of the protein combination.

6. The use according to claim 5, wherein The expression level of the protein combination is detected by liquid chromatography and mass spectrometry.

7. A kit for predicting the prognosis of medullary thyroid carcinoma, comprising a reagent for detecting the expression level of a protein combination, wherein the protein combination is composed of: HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1.

8. A system for prognostic stratification of medullary thyroid carcinoma, the system comprising: A subsystem for detecting the expression level of the protein combination, wherein the protein combination is composed of: HSPB7, PTPRM, MELTF, LAMA5, ITIH1, SERPINI1, COL10A1, PENK, OLFM3, PURB, CBX4, MGMT, PAFAH1B3, PALS1, EH D4, GOLGA6L7, GFOD2, TMC4, CALD1, SRI, GDPD1, SELENOI, CELSR3, SFRP2, GDA, TUBB2B, VGF, SCUBE3, HLA-DQA1; A data storage module, which is used to store information on the expression level of the protein combination; A data processing module processes the information of the data storage module, wherein the fitness of the feature subset is evaluated using a random forest classifier as the main performance indicator; The output module outputs the data processing results of the data processing module as the recurrence probability predicted by the model, ranging from 0 to 1, where greater than 0.5 indicates a high-risk group and less than or equal to 0.5 indicates a low-risk group.

Citation Information

Patent Citations

  • Application of protein combination in preparation of kit for prognosis stratification of children thyroid cancer as well as kit and system thereof

    CN115144599A

  • Methods for identifying, diagnosing, and predicting survival of lymphomas

    US20110152115A1