Protein marker combination for predicting lateral cervical lymph node metastasis of papillary thyroid carcinoma as well as screening method and application of protein marker combination
By screening multi-cohort proteomics data and constructing machine learning models, the insufficient sensitivity and specificity of existing technologies for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma have been addressed. A combination of highly stable and reproducible protein biomarkers has been provided for the accurate prediction and in vitro diagnosis of lateral cervical lymph node metastasis in papillary thyroid carcinoma.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENGJING HOSPITAL OF CHINA MEDICAL UNIVERSITY
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies lack a high-performance, high-stability combination of protein biomarkers specifically for predicting lateral cervical lymph node metastasis (LLNM) in papillary thyroid carcinoma. Existing prediction methods have insufficient sensitivity and specificity, and cross-platform validation is difficult.
By systematically screening multi-cohort proteomics data, 31 protein biomarker combinations were identified, and machine learning models were constructed for prediction. This included the integration of proteomics data, screening for common proteins, screening for consistency in expression direction, and performance evaluation, resulting in highly stable biomarker combinations that were then detected using protein microarrays or mass spectrometry.
It achieves accurate and objective assessment of the risk of lateral cervical lymph node metastasis in papillary thyroid carcinoma, with predictive performance significantly superior to existing methods. The biomarker combination has high stability and reproducibility, and is suitable for in vitro diagnosis of clinical puncture or surgical samples.
Smart Images

Figure CN121955408A_ABST
Abstract
Description
A combination of protein biomarkers for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, their screening methods, and applications. Technical Field
[0001] This invention relates to the field of biomedical technology, specifically to a combination of protein biomarkers for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, their screening methods, and applications. Background Technology
[0002] Papillary thyroid carcinoma (PTC) is the most common subtype of thyroid cancer, and its prognosis is generally good. However, lateral cervical lymph node metastasis (LLNM) is a key factor affecting patient treatment decisions and long-term prognosis. In current clinical practice, preoperative assessment of LLNM risk mainly relies on imaging examinations such as neck ultrasound, but its sensitivity is limited and it is highly dependent on the operator's experience, making it difficult to conduct objective and quantitative accurate assessment.
[0003] Existing technologies have attempted to utilize protein biomarkers for PTC diagnosis or risk prediction. For example, patent CN115856313A discloses six proteins, including ANXA1, PDLIM4, and FN1, for the diagnosis of papillary thyroid microcarcinoma (PTMC) and prediction of cervical lymph node metastasis. However, the AUC of the single protein FN1 used for predicting metastasis is only 0.690, indicating limited predictive efficacy, and it has not been independently validated across multiple centers and technology platforms.
[0004] Patent CN115792247A discloses a combination of six proteins, including DPP7 and DPLIM3, for preoperative risk stratification (high / intermediate / low risk) in PTC. This patent focuses on comprehensive risk stratification based on the risk of postoperative recurrence, rather than prediction of the specific clinical phenotype of lymph node metastasis (yes / no).
[0005] Patent CN111292801A discloses a list of proteins and a deep learning model, which is mainly used for the differentiation of benign and malignant thyroid nodules. Its application scenarios and objectives are different from those for lymph node metastasis prediction.
[0006] In existing technologies, some detection methods based on gene or protein biomarkers (such as certain multi-gene detection kits) are mainly used for differentiating between benign and malignant thyroid nodules or for overall prognostic assessment, and are not specifically designed for predicting LLNM, a specific clinical problem. These methods suffer from insufficient specificity and need to improve sensitivity in predicting LLNM. In addition, existing biomarker studies are mostly based on single cohorts or small samples, lacking systematic validation across different studies and different proteomics technology platforms, resulting in poor stability and reproducibility of biomarkers, which severely limits their clinical translation and application.
[0007] Therefore, there is an urgent clinical need to develop a combination of protein biomarkers specifically targeting LLNM in patients with PTC, based on multi-cohort validation, with high stability and high predictive performance, along with their predictive models and detection methods, in order to achieve accurate and objective preoperative assessment of LLNM risk. Summary of the Invention
[0008] To address the lack of high-performance, high-stability protein biomarker combinations specifically for predicting lateral cervical lymph node metastasis (LLNM) in existing technologies, as well as the insufficient sensitivity, specificity, and difficulties in cross-platform validation of existing prediction methods, this invention aims to provide a protein biomarker combination, its screening method, corresponding prediction model, and detection application based on systematic screening and validation of multi-cohort proteomics data, in order to achieve reliable prediction of LLNM risk in PTC patients.
[0009] To achieve the above-mentioned objectives, the present invention provides the following technical solutions.
[0010] This invention discloses a protein biomarker combination for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, characterized in that the protein biomarker combination consists of the following 31 proteins: VWA8, PDHX, DENR, ATP5MG, SFTPB, MAOA, PHB1, CHI3L1, ACAA2, PTPN9, AFM, ACADVL, HSPE1, PPP2R2A, REEP5, FABP5, NME3, AUH, TMEM214, NDUFAF2, PHB2, FUCA2, PAXX, APOO, TMEM109, CHID1, MAGT1, ARL8B, BCAP29, AFG3L2, and MTCH2.
[0011] This invention also discloses a method for screening the above-mentioned protein biomarker combinations, characterized by comprising the following steps: S1. Obtaining at least three independent papillary thyroid carcinoma proteomics datasets, each dataset containing samples with lateral cervical lymph node metastasis and samples without lateral cervical lymph node metastasis; S2. Screening for common proteins detected in all three datasets; S3. From the common proteins, screening for proteins whose expression change direction is consistent between the lateral cervical lymph node metastasis group and the non-metastasis group in all datasets, obtaining a set of proteins with consistent direction; S4. Evaluating the predictive performance of the proteins in the set of proteins with consistent direction, screening for proteins whose predictive performance meets a preset threshold, forming a candidate protein set; S5. From the candidate protein set, screening for proteins whose expression level in papillary thyroid carcinoma tumor tissue is significantly different from that in paired adjacent normal tissue, obtaining the final protein biomarker combination.
[0012] Furthermore, in step S1, the at least three independent datasets are derived from proteomics research cohorts employing different technology platforms or different research centers.
[0013] Further, in step S4, the predictive performance evaluation is to calculate the area under the receiver operating characteristic curve (AUC) of a single protein predicting lateral cervical lymph node metastasis, and the preset threshold is AUC ≥ 0.7.
[0014] This invention also discloses a method for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, characterized by comprising: detecting the expression level of the above-mentioned protein biomarker combination in a sample of papillary thyroid carcinoma to be tested; inputting the expression level data into a prediction model pre-trained based on the protein biomarker combination to obtain a prediction result of the risk of lateral cervical lymph node metastasis in the sample to be tested; wherein the prediction model is a classification model trained using a machine learning algorithm.
[0015] The present invention also discloses a detection kit for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, characterized in that it contains reagents for detecting the expression levels of the combination of protein markers described above.
[0016] Furthermore, the kit is a kit based on protein chip, immunological detection, or mass spectrometry analysis.
[0017] The present invention also discloses the use of the protein biomarker combination described in 1 above in the preparation of reagents or devices for assisting in the diagnosis or assessment of lateral cervical lymph node metastasis in papillary thyroid carcinoma.
[0018] The present invention also discloses an electronic device, including a memory and a processor, characterized in that the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0019] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method described above.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows.
[0021] Precise prediction: This invention is the first to systematically focus on the discovery of protein biomarkers for the specific clinical challenge of PTC lateral cervical lymph node metastasis (LLNM), clearly distinguishing it from existing products that are mainly used for benign or malignant differentiation or general prognostic assessment.
[0022] Robust biomarker discovery process: An innovative biomarker discovery process integrating multiple independent cohorts, cross-technology platforms, and combining directional consistency and performance screening was proposed and applied, ensuring the high stability and reproducibility of the discovered biomarkers and overcoming the limitation of unreliability of single research results.
[0023] Excellent predictive performance: The predictive model, built on the discovered protein biomarker library (e.g., 31 proteins), demonstrates extremely high prediction accuracy in multiple independent validation cohorts, significantly outperforming existing methods.
[0024] The biological significance is clear: the screened biomarkers are enriched in pathways closely related to tumor invasion and metastasis, such as mitochondrial fatty acid β-oxidation and endoplasmic reticulum stress, providing clues for mechanism research and exploration of potential therapeutic targets.
[0025] Strong clinical translation potential: The provided biomarker combinations and detection methods (such as targeted detection based on protein chips or mass spectrometry) are applicable to clinical puncture or surgical samples, making it easy to develop into in vitro diagnostic products. Attached Figure Description
[0026] Figure 1 is a flowchart of the method for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma according to the present invention.
[0027] Figure 2 is a Venn diagram showing the screening results of common proteins in the three datasets: Study 5, Study 6, and Study Shengjing (SHJ).
[0028] Figure 3 is a Venn diagram showing the screening of proteins (i.e., "direction-consistent proteins") that were either upregulated or downregulated in the transfer group in all three datasets: Study 5, Study 6, and Study Shengjing (SHJ).
[0029] Figure 4 shows the receiver operating characteristic curves of the ridge regression logistic regression model based on a library of 31 protein biomarkers in the training set (Study 5).
[0030] Figure 5 shows the receiver operating characteristic curves (ROCs) of cross-validation based on a library of 31 protein biomarkers in the training set (Study 5).
[0031] Figure 6 shows the receiver operating characteristic curves of the ridge regression logistic regression model based on a library of 31 protein biomarkers in the test set (Study 6).
[0032] Figure 7 shows the receiver operating characteristic curves of the ridge regression logistic regression model based on a library of 31 protein biomarkers on the independent validation set (StudyShengjing). Detailed Implementation
[0033] The present invention will be further described in detail below with reference to specific embodiments. However, this should not be construed as limiting the scope of the above-described subject matter of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.
[0034] Unless otherwise specified, all reagents and materials used in this invention are commercially available.
[0035] Example 1.
[0036] 1. Experimental materials.
[0037] 1.1 Source of clinical samples.
[0038] The papillary thyroid carcinoma tissue samples used in this study were obtained from the Department of Thyroid Surgery, Shengjing Hospital, China Medical University. Sample collection took place from August 2021 to April 2024. After surgical resection, the specimens were immediately selected by pathologists as representative tumor tissue and paired adjacent normal tissue (≥2 cm from the tumor margin, confirmed by postoperative pathology to be free of cancer cell infiltration). These were then divided into approximately 50 mg tissue blocks, rapidly frozen in liquid nitrogen, and subsequently transferred to a -80°C ultra-low temperature freezer for storage. All samples underwent HE staining for pathological confirmation, ensuring tumor purity >70%.
[0039] Based on the postoperative pathological lymph node dissection results, patients were divided into a lateral cervical lymph node metastasis group (LLNM+, postoperative pathology confirmed the presence of metastatic lesions in any region of the lateral cervical region II-V) and a non-lateral cervical lymph node metastasis group (LLNM-, no lymph node metastasis in the lateral cervical region). A total of 14 LLNM+ samples and 14 LLNM- samples were included, designated as Study Shengjing.
[0040] 1.2 Main reagents.
[0041] Reagents used for protein extraction and quantification include: urea (purity ≥99%, Amresco), ammonium bicarbonate (Sigma-Aldrich, catalog number A6141-500G), tetraethylammonium bromide (TEAB, Sigma-Aldrich, catalog number T7408-100mL), Bradford protein quantification staining solution (Huaxingbio, catalog number HXJ5137), and bovine serum albumin standard (Thermo Scientific, catalog number 23209).
[0042] Protein reduction and alkylation reagents include: dithiothreitol (DTT, Amresco, catalog number M109-5G) and iodoacetamide (IAM, Amresco, catalog number M216-30G).
[0043] The protease digestion and labeling reagents include: sequencing-grade trypsin (Promega, catalog number V5280) and TMT10plex isotope labeling reagent kit (Thermo Fisher, catalog number 90111).
[0044] Chromatographic separation reagents include: acetonitrile (chromatographic grade, JTBaker, catalog number 34851), formic acid (mass spectrometry grade, Sigma-Aldrich, catalog number T79708), and ammonia (Wako Pure Chemical Industries, catalog number 013-23355).
[0045] Sample processing consumables include: C18 desalting column (Ziptip, Millipore, catalog number ZTC18M096), XBridge Peptide BEH C18 column (Waters, 5μm particle size, catalog number 186003581), mass spectrometer vials and matching caps (Thermo, catalog numbers 11190533 and 11150635 respectively).
[0046] 1.3 Main instruments.
[0047] The protein extraction and pretreatment instruments include: a high-throughput tissue homogenizer (Shanghai Hefan Instrument Co., Ltd., model hf-48), an ultrasonic homogenizer (Shanghai Huxi Industrial Co., Ltd., model JY96-IIN), a vortex mixer (SCILOGEX, model MX-S), a benchtop high-speed centrifuge (Eppendorf), and an electric thermostatic water bath (Beijing Guangming Medical Instrument Co., Ltd., model XMTD-7000).
[0048] The protein quantification and separation instruments include: an ELISA reader (DR200B model) and a polyacrylamide gel electrophoresis system (Bio-Rad).
[0049] The sample concentration and chromatographic separation instruments include: a vacuum centrifuge concentrator (Beijing Genesys Logic Technology Co., Ltd., model CV100-DNA) and a high performance liquid chromatography system (Beijing RIGOL Technology Co., Ltd., model RIGOL L-3000).
[0050] The mass spectrometry instruments include: a nanoliter electrospray ion source quadrupole-electrostatic field orbital trap high-resolution mass spectrometer (Thermo Scientific, Q ExactiveHF-X model), equipped with a Nanospray Flex ion source.
[0051] The data processing workstation is equipped with the Proteome Discoverer 2.4 software platform.
[0052] 1.4 Sources and Acquisition of Public Data.
[0053] Study 5 Dataset: The Study 5 dataset originates from a 2022 proteomics study on lymph node metastasis associated with papillary thyroid carcinoma published in *Frontiers in Oncology* (Cao et al., 2022, *Frontiers in Oncology*, 12:887977). This study employed liquid chromatography-tandem mass spectrometry (LC-MS / MS) in a direct data-independent acquisition mode to perform systematic proteomics analysis on primary lesions of patients with surgically treated papillary thyroid carcinoma. Thirty-three primary lesion tissue samples from papillary thyroid carcinoma were included, of which 17 were without lymph node metastasis (NLNM group) and 16 were with lymph node metastasis (LNM group). The dataset was obtained from the ProteomeXchange Consortium database (http: / / proteomecentral.proteomexchange.org) through the iProX partner repository, with the dataset identifier PXD031838. The downloaded data includes raw mass spectrometry data files and protein quantification matrices.
[0054] Study 6 Dataset: The Study 6 dataset originates from a proteomics study on lymph node metastasis associated with papillary thyroid carcinoma published in the *Journal of Proteome Research* in 2025 (Zhang et al., 2025, J. Proteome Res.24:256-267). This study also employed liquid chromatography-tandem mass spectrometry in a data-independent acquisition mode, focusing on the molecular mechanisms of lateral cervical lymph node metastasis (LLNM). The study included 10 patients with papillary thyroid carcinoma, divided into two groups of 5: no lymph node metastasis (NLNM) (n=5) and lateral cervical lymph node metastasis (LLNM) (n=5). The dataset was obtained from the ProteomeXchange Consortium database through the iProX partner repository, with the dataset identifier PXD057473. Since the supplementary materials of this study already provided a protein quantitative expression matrix, the raw mass spectrometry data and the normalized protein expression matrix were also downloaded for analysis in this study.
[0055] 2. Experimental methods.
[0056] 2.1 Extraction of tissue proteins.
[0057] Frozen tissue samples were removed from the cryogenic freezer and ground into powder using a high-throughput tissue homogenizer in liquid nitrogen. An appropriate amount of tissue powder was weighed into a centrifuge tube, and pre-chilled protein lysis buffer (prepared by a 10:1 volume ratio of 8M urea aqueous solution to a protease inhibitor) was added. The sample was ultrasonically disrupted in an ice bath on intermittent mode until complete lysis and a clear solution. The sample was then centrifuged at 14,100 times the gravitational acceleration for 20 minutes in a cryogenic high-speed centrifuge, carefully collecting the supernatant while avoiding the aspiration of precipitate. Protein concentration was determined using the Bradford method, establishing a concentration gradient standard curve using bovine serum albumin as a standard. Absorbance values were read at a specific wavelength using a microplate reader to calculate the protein concentration for each sample. The quantified protein samples were aliquoted and stored at -80°C for later use.
[0058] 2.2 Protein reduction and enzymatic digestion.
[0059] 50 μg of protein was taken from each sample for reduction treatment. Dithiothreitol stock solution was added to the protein solution to achieve a final concentration of 200 mmol / L. The sample was incubated in a 37°C water bath for 1 hour to ensure complete cleavage of the protein disulfide bonds. Subsequently, the sample was diluted 8-fold with ammonium bicarbonate buffer to reduce the urea concentration. Sequencing-grade trypsin was added at a trypsin-to-substrate protein ratio of 1:25, and the mixture was incubated overnight at 37°C.
[0060] For samples requiring desalting, the following procedure was used: 100 μg of protein was taken, and a reducing agent was added to bring the final concentration to 10 mmol / L. The mixture was incubated at 37°C for 1 hour, then allowed to return to room temperature. Iodoacetamide, an alkylating agent, was added to bring the final concentration to 40 mmol / L. The mixture was reacted at room temperature for 45 minutes in the dark to block free thiol groups. The sample was further diluted with ammonium bicarbonate buffer, and the pH of the solution was measured to approximately 8. Trypsin was added at a protein-to-trypsin mass ratio of 50:1, and the mixture was incubated overnight at 37°C. The next day, formic acid aqueous solution was added to bring the final volume fraction to 0.1%, thus terminating the enzymatic digestion reaction.
[0061] 2.3 Peptide desalting and purification.
[0062] The enzymatically digested peptide mixture was purified using a reverse-phase C18 desalting column. First, the desalting column packing material was activated with pure acetonitrile, followed by column equilibration with 0.1% formic acid aqueous solution. The sample was loaded onto the column, ensuring the peptides were fully bound to the stationary phase. The column was washed with 0.1% formic acid aqueous solution to remove salts and polar impurities. Finally, the peptides were eluted with an aqueous solution containing 70% acetonitrile, and the eluent was collected, concentrated by vacuum centrifugation, and freeze-dried to obtain purified peptide powder.
[0063] 2.4 Isotope labeling.
[0064] Remove the TMT10plex labeling reagent from -20℃ and allow it to equilibrate at room temperature before opening the cap. Add 41 μL of anhydrous acetonitrile to each tube of labeling reagent, vortex for 5 minutes to fully dissolve the reagent, and briefly centrifuge to allow the liquid to collect at the bottom of the tube. Add the dissolved labeling reagent to 100 μg of the corresponding enzyme-digested peptide sample and react at room temperature for 1 hour to allow the labeling reagent to covalently bind to the amino groups of the peptide. After the reaction is complete, add ammonia solution to terminate the labeling reaction. Transfer all labeled samples to the same container, vortex thoroughly to mix, briefly centrifuge, and then freeze-dry under vacuum to obtain the mixed labeled sample.
[0065] 2.5 Reversed-phase chromatography fractionation.
[0066] The lyophilized mixed-labeled sample was dissolved in mobile phase A, which consisted of 100 μL of aqueous solution. The sample was centrifuged at 14000x gravity for 20 min, and the supernatant was used for high-performance liquid chromatography (HPLC) fractionation. Chromatographic separation was performed using an XBridge Peptide BEH C18 column at a flow rate of 0.7 mg / min. The elution gradient was set as follows: initial mobile phase B volume fraction of 5%, maintained for 5 min, linearly increased to 8%, then reached 18% at 40 min, increased to 32% at 62 min, rapidly increased to 95% at 64 min and maintained until 68 min, and finally returned to the initial proportion of 5% at 72 min. Mobile phase A was an aqueous solution containing 5% acetonitrile, and mobile phase B was an aqueous solution containing 95% acetonitrile. Based on the chromatographic peak shape, the eluent was collected at time intervals into multiple fractions, and each fraction was lyophilized separately for later use.
[0067] 2.6 Liquid chromatography-tandem mass spectrometry analysis.
[0068] Chromatographic conditions: The fractionated lyophilized powders were dissolved in mobile phase A, which consisted of 100% water and 0.1% formic acid. The dissolved samples were centrifuged at 14000x gravity for 20 minutes at 4°C, and the supernatant was used for mass spectrometry analysis. 1 μg of peptide was loaded and separated using a nano-flow chromatography system. Mobile phase A was an aqueous solution containing 0.1% formic acid, and mobile phase B was an aqueous solution containing 80% acetonitrile and 0.1% formic acid. The chromatographic elution program was as follows: initially, the proportion of mobile phase B was 7%, increasing to 15% at 11 minutes, reaching 25% at 48 minutes, increasing to 40% at 68 minutes, and rapidly increasing to 100% at 69 minutes and maintaining this level until 80 minutes.
[0069] Mass spectrometry acquisition parameters: A Q Exactive HF-X mass spectrometer was used for detection, equipped with a nanoliter electrospray ionization source. Ion source parameter settings: ion spray voltage 2.4 kV, ion transfer tube temperature 275℃. The mass spectrometry employed a data-dependent acquisition mode, with the first-stage full scan mass range set to a mass-to-charge ratio of 407 to 1500, a resolution set to 60000 (with a mass-to-charge ratio of 200 as a reference), and an automatic gain control target value of 3 × 10⁻⁶. 6 The maximum C-trap injection time is 20 ms. After the first-stage scan, the 40 precursor ions with the highest signal intensity are automatically selected for second-stage fragmentation using a high-energy collision-induced dissociation mode. The second-stage resolution is 45,000 (with a mass-to-charge ratio of 200 as a reference), and the automatic gain control target value is 5 × 10⁻⁶. 4 The maximum injection time was 86 ms, the normalized collision energy was 32%, and the isolation window was 1.2 mass-to-charge ratio units. The raw mass spectrometry data were saved in .raw format.
[0070] 2.7 Database retrieval and protein quantification.
[0071] The Proteome Discoverer 2.4 software platform was used for database searching. The human UniProt database was selected, and trypsin was used for enzyme digestion, allowing a maximum of two missed cleavage sites. Fixed modifications were defined as urea methylation of cysteine, with variable modifications including methionine oxidation, TMT10plex labeling (lysine side chain and N-terminus of peptide), and N-terminal acetylation of protein. The precursor ion mass tolerance was set to ±10 ppm, and the fragment ion mass tolerance was set to ±0.02D. False discovery rates at both peptide and protein levels were controlled below 1%. TMT reporter ion intensity was used for relative protein quantification, and the relative expression level of protein in different samples was calculated as the ratio of signal intensities in each channel.
[0072] 3. Integration of multi-cohort proteomics data and construction of candidate biomarker libraries.
[0073] 3.1 Data Sources: Three independent PTC proteomics research datasets (including Study 5, Study 6, and Study Shengjing) were integrated. Considering the sample size, technology platform characteristics, and complementary research designs of each cohort, the following partitioning strategy was adopted: Training set: Study 5 dataset, used for model training, feature selection, and hyperparameter tuning; Test set: Study 6 dataset, used for external validation on the same technology platform; Independent validation set: Study Shengjing dataset, used for independent validation across technology platforms and centers.
[0074] The rationale for the above division strategy is as follows: Study 5 has the largest sample size (33 cases), which can provide sufficient training data; Study 6 uses the same technology platform (DIA-MS) as Study 5 but is an independent patient group, which can evaluate the model's generalization ability within the same platform; StudyShengjing uses a different technology platform (TMT labeling and quantification) and is collected from different centers, which can evaluate the model's cross-platform and cross-center transferability.
[0075] 3.2 Protein Intersection: Study Shengjing and Study 5 share 3964 proteins; Study Shengjing and Study 6 share 3729 proteins; Study 5 and Study 6 share 3975 proteins; Study 5, Study 6, and Study Shengjing share a total of 3430 proteins. Subsequent cross-study consistency analysis and biomarker screening were all based on these 3430 common proteins.
[0076] 3.3 Directional Consistency Screening: Differential expression analysis of LLNM+ vs LLNM- was performed on each dataset. The significance of the difference was calculated using the Wilcoxon rank-sum test. Proteins with completely consistent change directions in the three datasets were screened, resulting in 1013 "directionally consistent proteins" (429 upregulated and 584 downregulated).
[0077] 3.4 Preliminary Performance Screening: In each dataset, the AUC value of the predicted LLNM for each protein in the "Orientation Consistent Proteins" group was calculated individually. Proteins with an AUC ≥ 0.7 in at least one dataset were screened, resulting in 42 candidate proteins.
[0078] 3.5 Specific Enhancement Screening: To exclude proteins that show general differences between cancer and adjacent normal tissues, the focus was on LLNM-specific signals, and the differences of these 42 proteins in tumor tissue and paired adjacent normal tissues were calculated. Proteins that were significantly different in expression in tumor tissue from those in adjacent normal tissue and showed significant changes in LLNM were ultimately retained, forming a final candidate biomarker library of 31 proteins (VWA8, PDHX, DENR, ATP5MG, SFTPB, MAOA, PHB1, CHI3L1, ACAA2, PTPN9, AFM, ACADVL, HSPE1, PPP2R2A, REEP5, FABP5, NME3, AUH, TMEM214, NDUFAF2, PHB2, FUCA2, PAXX, APOO, TMEM109, CHID1, MAGT1, ARL8B, BCAP29, AFG3L2, MTCH2).
[0079] 4. Predictive model construction and validation (based on a library of 31 protein biomarkers).
[0080] 4.1 Feature and Data Preparation: The relative quantitative values of 31 proteins were used as modeling features. The Study 5 dataset was used as the training set. For the test set (Study 6) and the independent validation set (Study Shengjing), the data were pre-corrected for distribution based on the mean and variance of the training set to eliminate differences between technology platforms.
[0081] 4.2 Model Construction.
[0082] Algorithm selection: Ridge Logistic Regression was used on the Study 5 training set. This algorithm constrains the model complexity through L2 regularization and is suitable for scenarios where the high-dimensional features (31 proteins) are relatively small compared to the sample size (33 cases), which can effectively reduce the risk of overfitting.
[0083] Hyperparameter tuning: Repeated 3-Fold Cross-Validation (repeated 50 times) was used for hyperparameter tuning. In each repetition, the training set was randomly divided into 3 folds: 2 folds were used for training and 1 fold was used for validation. This process was repeated 3 times to ensure that each fold was used as a validation set once. The optimal regularization parameter λ was determined through grid search, with the search range being a preset sequence. The optimal λ was ultimately determined to be 4.317.
[0084] Model training: The final model was trained on the entire training set (33 cases) using the optimal regularization parameters to obtain the weight coefficients of 31 proteins and the model intercept.
[0085] Decision threshold determination: Based on the predicted probabilities of the training set, the optimal decision threshold is determined by maximizing the Youden index (sensitivity + specificity - 1). The optimal threshold for the training set is 0.471, and the optimal threshold for cross-validation is 0.486.
[0086] 4.3 Model Validation Results: The model was evaluated on the training set using the optimal model parameters (λ=4.317) and decision threshold (0.471). The model's AUC on the training set was 0.893, accuracy was 0.879, sensitivity was 0.938, specificity was 0.824, positive predictive value was 0.833, and negative predictive value was 0.933.
[0087] 4.4 Cross-validation robustness: The average performance obtained through repeated cross-validation is as follows: AUC 0.824, accuracy 0.848, sensitivity 0.938, specificity 0.765, positive predictive value 0.789, negative predictive value 0.929, and average optimal threshold 0.486. The cross-validation performance is slightly lower than that of the training set, but the AUC remains above 0.8, indicating that the model has a certain generalization ability under small sample conditions and does not exhibit severe overfitting.
[0088] 4.5 Test set (Study 6): The trained model and threshold (0.471) were applied to the Study 6 data. The prediction AUC reached 0.960, accuracy 0.9, sensitivity 0.8, and specificity 1.0.
[0089] 4.6 Independent Validation Set (Study Shengjing): The trained model and threshold (0.471) were applied to the Study Shengjing data, achieving accurate classification of all samples with an AUC of 1.000, and accuracy, sensitivity, and specificity were all 1.0.
[0090] 5. Data Usage Compliance Statement: The Study 5 and Study 6 datasets used in this study are both publicly published scientific research data and comply with public data usage guidelines.
[0091] The above description is merely a preferred embodiment of the present invention and is not intended to limit the patent scope of the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A combination of protein biomarkers for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, characterized in that, The protein biomarker group consists of the following 31 proteins: VWA8, PDHX, DENR, ATP5MG, SFTPB, MAOA, PHB1, CHI3L1, ACAA2, PTPN9, AFM, ACADVL, HSPE1, PPP2R2A, REEP5, FABP5, NME3, AUH, TMEM214, NDUFAF2, PHB2, FUCA2, PAXX, APOO, TMEM109, CHID1, MAGT1, ARL8B, BCAP29, AFG3L2, and MTCH2.
2. A method for screening combinations of protein biomarkers as described in claim 1, characterized in that, Includes the following steps: S1. Obtain at least three independent papillary thyroid carcinoma proteomics datasets, each containing samples with lateral cervical lymph node metastasis and samples without lateral cervical lymph node metastasis; S2. Screen for common proteins that are detected in all three datasets; S3. From the common proteins, proteins whose expression change direction is consistent between the lateral cervical lymph node metastasis group and the non-metastasis group in all datasets are selected to obtain a set of proteins with consistent direction. S4. Evaluate the predictive performance of the proteins in the set of proteins with consistent orientation, and select the proteins whose predictive performance meets the preset threshold to form a candidate protein set; S5. From the candidate protein set, proteins whose expression levels in papillary thyroid carcinoma tumor tissue are significantly different from those in paired adjacent normal tissue are selected to obtain the final protein biomarker combination.
3. The method according to claim 2, characterized in that, In step S1, the at least three independent datasets are derived from proteomics research cohorts using different technology platforms or different research centers.
4. The method according to claim 2, characterized in that, In step S4, the predictive performance evaluation is to calculate the area under the receiver operating characteristic curve (AUC) of a single protein predicting lateral cervical lymph node metastasis, and the preset threshold is AUC ≥ 0.
7.
5. A method for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, characterized in that, include: The expression level of the protein biomarker combination described in claim 1 is detected in a papillary thyroid carcinoma sample to be tested; the expression level data is input into a prediction model pre-trained based on the protein biomarker combination to obtain a risk prediction result of lateral cervical lymph node metastasis in the sample to be tested; the prediction model is a classification model trained using a machine learning algorithm.
6. A detection kit for predicting lateral cervical lymph node metastasis in papillary thyroid carcinoma, characterized in that, It includes a reagent for detecting the expression level of at least one protein in the protein biomarker combination of claim 1.
7. The detection kit according to claim 6, characterized in that, The kit is based on protein chip, immunological detection, or mass spectrometry analysis.
8. Use of the protein biomarker combination of claim 1 in the preparation of reagents or devices for assisting in the diagnosis or assessment of lateral cervical lymph node metastasis in papillary thyroid carcinoma.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method described in claim 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in claim 5.