Genetic biomarkers for predicting the risk of progression in papillary thyroid carcinoma and their applications
By detecting the expression levels of SPC24, TCEB3B, PLXNA3, AZGP1, GRP156, and NRGN genes, a kit and assessment system for predicting the prognostic progression risk of papillary thyroid carcinoma were constructed. This solves the problem of the lack of accurate prediction methods in existing technologies and enables accurate prediction and prevention of papillary thyroid carcinoma prognosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
- Filing Date
- 2025-06-09
- Publication Date
- 2026-08-04
AI Technical Summary
The lack of effective biomarkers in current technologies for predicting the recurrence risk of papillary thyroid carcinoma leads to a high recurrence rate after treatment. Existing methods mainly rely on drug intervention and imaging monitoring, lacking precise predictive means.
Gene markers such as SPC24, TCEB3B, PLXNA3, AZGP1, GRP156, and NRGN are provided. By detecting their expression levels, a kit and assessment system for predicting the prognostic progression risk of papillary thyroid carcinoma are constructed, including data acquisition, processing, and output modules. A prognostic model is constructed using multivariate Cox regression analysis.
It can accurately predict the prognostic progression risk of papillary thyroid carcinoma, provide patients with a reference for prognostic prevention and treatment, improve the accuracy and effectiveness of prediction, and reduce the risk of recurrence.
Smart Images

Figure CN120758625B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene technology, specifically to a gene biomarker for predicting the risk of progression of papillary thyroid carcinoma and its application. Background Technology
[0002] The incidence of thyroid cancer is rising globally, primarily driven by the increase in papillary thyroid carcinoma (PTC). Papillary thyroid carcinoma (PTC), the most common subtype of thyroid cancer, accounts for 96.0% of all new thyroid cancer cases and 66.8% of cancer deaths, with an extremely high recurrence rate. Reducing thyroid cancer recurrence is a key goal after treatment, especially for differentiated thyroid cancers (papillary and follicular carcinomas). Currently, methods for preventing thyroid cancer recurrence mainly rely on drug intervention and imaging monitoring.
[0003] PTC recurrence is associated with many factors. Currently, independent risk factors for recurrence have been identified, including tumor size, multifocality, extrathyroidal spread, and lymph node metastasis. In addition, extranodal extension, the ratio of metastatic lymph nodes, and non-initial treatment are potential risk factors for recurrence. Besides pathological and clinical features serving as predictors of poor prognosis in PTC, studies have also found that biomolecular markers are associated with the risk of PTC recurrence.
[0004] An analysis based on the TCGA database predicted mRNAs such as FN1 and ITGα3, miRNAs such as miR-486 and miR-1179, and 59 genetic variants as potential biomarkers associated with PTC recurrence. Recently, the WHO made significant adjustments to the classification of TC, emphasizing the differences in cellular origin, pathological features, molecular classification, and biological behavior among benign, low-risk, and malignant tumors, while also highlighting the role of biomarkers in aiding diagnosis and prognosis. Preventing recurrence is crucial for improving survival; identifying and intervening in risk factors for PTC recurrence can help prevent it.
[0005] In eukaryotic cells, biological activities are carried out in different cellular compartments or organelles. In addition to the typical organelles with membrane structures, the cell itself also contains multifunctional non-membrane organelles, and the biophysical process of liquid-liquid phase separation (LLPS) elucidates the assembly and formation of non-membrane organelles. LLPS is causally related to many dysregulated cellular processes in cancer and participates in tumor processes by promoting cancer cell proliferation and tumor metastasis, helping cancer cells evade growth inhibition and immune destruction, and facilitating apoptosis under stress.
[0006] There are reports that LLPS participates in the regulation of various malignant tumors such as lung cancer, liver cancer, breast cancer, and prostate cancer, but there are no publicly available research reports on whether LLPS participates in the regulation of PTC. Summary of the Invention
[0007] To address the gaps in existing technologies, this invention provides a genetic biomarker for predicting the risk of progression of papillary thyroid carcinoma, wherein the biomarker includes at least one of SPC24, TCEB3B, PLXNA3, AZGP1, GRP156, and NRGN.
[0008] Furthermore, high expression of at least one of the gene markers SPC24, TCEB3B, PLXNA3, AZGP1, and GRP156 increases the risk of progression of papillary thyroid carcinoma; low expression of NRGN increases the risk of progression of papillary thyroid carcinoma.
[0009] The second invention provides the application of the genetic marker in any of the following aspects:
[0010] A1) Application in the preparation of kits for predicting the risk of progression of papillary thyroid carcinoma;
[0011] A2) Application in the development of a system for predicting the risk of progression of papillary thyroid carcinoma.
[0012] In a third aspect, the present invention provides a kit for predicting the risk of progression of papillary thyroid carcinoma, the kit comprising reagents for detecting the genetic markers.
[0013] Furthermore, the kit includes reagents for detecting mRNA expression and / or protein expression of genes.
[0014] Furthermore, the kit also includes at least one of a transfection reagent and a carrier RNA.
[0015] Furthermore, the reagents for detecting mRNA expression actually include at least one of RNA extraction reagents, mRNA purification reagents, cDNA synthesis reagents, qPCR / RT-PCR detection reagents, and high-throughput sequencing library construction reagents; the reagents for detecting protein expression include at least one of protein extraction reagents, immunoassay reagents, protein purification reagents, and protein labeling reagents.
[0016] Furthermore, the reagents in the kit include any one of the following:
[0017] B1) A gene or its transcript or protein, wherein the sequence of the transcript or protein is the sequence of one or more or all transcripts or proteins of the corresponding gene;
[0018] B2) Primer molecules and / or probe molecules that hybridize with the gene or its transcript;
[0019] B3) Immune molecules of the protein expression product of the gene.
[0020] Furthermore, the kit is suitable for samples containing nucleic acids, and preferably, the samples include at least one of blood samples, tissue samples, body fluid samples, exfoliated cell samples, fecal samples, and urine samples.
[0021] In a fourth aspect, the present invention provides an assessment system for predicting the risk of progression of papillary thyroid carcinoma, the system comprising at least the following modules:
[0022] The data acquisition module is at least used to acquire the relative expression levels of the gene markers as described in claim 1;
[0023] The data processing module is at least used to calculate a score according to a formula, wherein the formula is: Risk Score = 1.3929258 × SPC24 expression value + 14.9827438 × TCEB3B expression value + 0.7688884 × PLXNA3 expression value - 0.4758786 × NRGN expression value + 0.6593675 × AZGP1 expression value + 1.20117 × GPR156 expression value;
[0024] The judgment module is at least used to compare the score obtained by the data processing module with a threshold. If the score is higher than the threshold, the group is considered high-risk; if the score is lower than the threshold, the group is considered low-risk. The threshold is 0.852704422136068.
[0025] A data output module, which is at least used to output results.
[0026] Furthermore, the system also includes a sample extraction module, which is used at least to extract nucleic acids from the sample.
[0027] In a fifth aspect, the present invention provides the use of the system in the preparation of instruments, devices, or reagent kits for predicting the risk of progression of papillary thyroid carcinoma.
[0028] The beneficial effects of the present invention include, but are not limited to:
[0029] This invention provides six genes for predicting the risk of progression of papillary thyroid carcinoma and a system for predicting the risk of progression of papillary thyroid carcinoma. The system can accurately predict the risk of progression of papillary thyroid carcinoma, providing a reference for the prognosis and prevention of patients with papillary thyroid carcinoma. Attached Figure Description
[0030] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0031] Figure 1 This is a schematic diagram of the differentially expressed gene results between papillary thyroid carcinoma and normal tissue in an embodiment of the present invention, where blue represents low-expression genes and red represents high-expression genes;
[0032] Figure 2 This is a schematic diagram of the differentially expressed gene results related to liquid-liquid phase separation between PTC and adjacent normal tissue in the TCGA cohort of this invention.
[0033] Figure 3 This is a schematic diagram of the functional enrichment analysis results of differentially expressed genes in an embodiment of the present invention;
[0034] Figure 4 This is a schematic diagram of the results of further screening of characteristic genes by LASSO regression analysis in an embodiment of the present invention, wherein Figure A is a coefficient path diagram of LASSO regression; Figure B is a 10-fold cross-validation curve of LASSO regression analysis; and Figure C is the ROC curve of the model.
[0035] Figure 5 This is a forest plot drawn from the prognostic risk assessment model for liquid-liquid phase separation in this embodiment of the invention;
[0036] Figure 6 This is a schematic diagram of the survival analysis results of the prognostic risk assessment model in an embodiment of the present invention;
[0037] Figure 7 This is a schematic diagram of the results of calculating the optimal cutoff value of the LRS score in the training set in an embodiment of the present invention. Figure A is a diagram of the risk score distribution of the high- and low-risk groups in the training set; Figure B is a scatter plot of the PTC progression event outcomes and occurrence time of each patient in the high- and low-risk groups in the training set; Figure C is a heatmap of the differential expression of the 6 core genes in the prognostic model in the high- and low-risk groups in the training set.
[0038] Figure 8 This is a schematic diagram of the survival analysis results of progression-free survival as the clinical outcome in an embodiment of the present invention, where Figure A is the Kaplan-Meier curve and Figure B is the time-ROC curve;
[0039] Figure 9 Figure A is a diagram of the survival analysis results of the prognostic risk assessment model in the test set in this embodiment of the invention. Figure B is a diagram of the risk score results of the high- and low-risk groups in the training set; Figure C is a scatter plot of the PTC progression event outcomes and occurrence time of each patient in the high- and low-risk groups in the training set; Figure D is a heatmap of the differential expression of the 6 core genes in the prognostic model in the high- and low-risk groups in the training set.
[0040] Figure 10This is a schematic diagram of the survival analysis results of testing progression-free survival as a clinical outcome in an embodiment of the present invention, where Figure A is the Kaplan-Meier curve and Figure B is the time-ROC curve. Detailed Implementation
[0041] The present invention is described in detail below with reference to the embodiments, but the present invention is not limited to these embodiments. Unless otherwise specified, the raw materials and catalysts in the embodiments of the present invention are all purchased through commercial channels.
[0042] Example 1: Constructing an Evaluation System
[0043] 1. Data Collection
[0044] (1) Data related to papillary thyroid carcinoma from the TCGA database: RNA-seq data, survival information, and clinical information of the THCA cohort were downloaded from the Cancer and Genome Atlas Database (TCGA) using code in R software (download code is...).
[0045] proj="TCGA-THCA"
[0046] if(F){download.file(url=paste0("https: / / gdc.xenahubs.net / download / ",proj,".htseq_counts.tsv.gz"),destfile=paste0(proj,".htseq_counts.tsv.gz"))
[0047] download.file(url=paste0("https: / / gdc.xenahubs.net / download / ",proj,".GDC_phenotype.tsv.gz"),destfile=paste0(proj,".GDC_phenotype.tsv.gz"))
[0048] download.file(url=paste0("https: / / gdc.xenahubs.net / download / ",proj,".survival.tsv"),destfile=paste0(proj,".survival.tsv"))}.
[0049] Based on the inclusion criteria of complete clinical information and survival time of ≥30 days, we obtained the clinical data of 56 normal thyroid tissues and 498 papillary thyroid carcinoma tissues in the THCA cohort.
[0050] (2) Liquid-liquid phase separation-related genes: Liquid-liquid phase separation-related genes (LRGs) were obtained from the PhaSepDB database (http: / / db.phasep.pro / ). First, files under the phase separation and membraneless organelle entries were downloaded, and R software was used to extract genes and remove duplicate genes. Finally, 5541 genes were obtained as liquid-liquid phase separation-related genes for subsequent analysis. The THCA cohort transcriptome data in the TCGA database included 498 papillary thyroid carcinoma samples and 56 adjacent normal samples, containing approximately 30956 genes. Differentially expressed genes were annotated with probes, and a fold change (FC) greater than 1 and P < 0.05 were used as screening criteria. A total of 5177 genes were found to be differentially expressed between papillary thyroid carcinoma and normal tissues; among them, 2800 genes were upregulated in PTC and 2377 genes were downregulated in PTC. The results are as follows: Figure 1 As shown.
[0051] 2. Data Processing
[0052] (1) Data Preprocessing: R software was used to perform count scaling at the probe level and filter out genes with zero or low expression levels in the samples, retaining only genes expressed in more than half of the samples. The "DESeq2" R package was used to normalize each dataset and remove batch effects. GTF files were downloaded, and RNA sequencing data from the TCGA cohort were annotated to obtain the expression matrices of all genes. The expression matrices were matched with clinical information, and only one sample was retained for each patient. The matched data were renamed and classified according to clinical characteristics, retaining only clinical information such as "gender", "race_list", "age_at_initial_pathologic_diagnosis", "stage_event_pathologic_stage", "stage_event_tnm_categories", "PFS", and "PFS.time". The clinical characteristics of each sample were further processed according to AJCC (American Joint Committee on Cancer) staging, whether there was extrathyroidal invasion, lymph node metastasis, or remote metastasis. No dataset containing the clinical endpoint of "progression-free survival (PFS)" was found in the GEO database.
[0053] (2) Differentially Expressed Gene Screening: The “DESeq2” R package was used with threshold conditions of |log2fold change|>1 and False Discovery Rate (FDR) <0.05 to identify differentially expressed genes (DEGs) between papillary thyroid carcinoma (PTC) tissues and adjacent normal tissues in the TCGA cohort. The intersection of the obtained differentially expressed genes with those related to liquid-liquid phase separation was used to obtain differentially expressed genes related to liquid-liquid phase separation between PTC and adjacent normal tissues in the TCGA cohort. Figure 2 It contains 318 genes, of which 195 genes are upregulated in papillary thyroid carcinoma tissue and 124 genes are downregulated in papillary thyroid carcinoma tissue.
[0054] (3) Functional enrichment analysis of differentially expressed genes: The biological pathways and functions of differentially expressed genes related to liquid-liquid phase separation were explored using the "clusterProfiler" package in R. First, the biological pathways associated with differentially expressed genes were assessed using the Kyoto Encyclopedia of Genes and Genomes (KEGG). Then, based on Gene Ontology (GO), further functional analyses were performed on three types of biological processes regulated by DEGs: cellular component (CC), molecular function (MF), and biological process (BP). GO annotation and KEGG enrichment analysis were performed on 318 differentially expressed genes related to liquid-liquid phase separation in papillary thyroid carcinoma. Figure A shows the top 10 results of GO annotation. GO annotation results show that liquid-liquid phase separation-related genes in papillary thyroid carcinoma are mainly distributed in cellular components such as the cytoskeleton, cell-basement membrane junctions, contractile fibroblasts, myofibrils, and sarcomeres. They primarily regulate molecular functions such as actin binding, protein kinase activity, transmembrane receptor kinase activity, growth factor binding, and cytoskeleton formation, participating in biological processes such as cell migration, post-translational protein modification, tissue migration, cell development, epithelial cell development, and promoting cell protrusion formation. Figure BKEGG enrichment analysis results show that differentially expressed phase separation-related genes in papillary thyroid carcinoma are mainly involved in pathways such as the muscle cytoskeleton, MAPK signaling pathway, PI3K-Akt signaling pathway, transcriptional dysregulation in cancer, nucleotide metabolism, purine metabolism, acute myeloid leukemia, bladder cancer, viral particles (Lassa virus and SFTS virus), and axonal guidance. The results are as follows... Figure 3 As shown.
[0055] 3. Construct a prognostic model for liquid-liquid phase separation
[0056] (1) Obtaining candidate genes for prognostic models: Univariate Cox regression analysis was performed on differentially expressed genes related to liquid-liquid phase separation using the "survival" R package in the training set. Progression-free survival (PFS) was used as the clinical endpoint; follow-up time was converted to months, and genes with a hazard ratio (HR) greater than 1 and p < 0.05 were selected as candidate genes for modeling. Normal tissue samples and thyroid papillary carcinoma samples that did not meet the inclusion criteria were removed from the TCGA thyroid papillary carcinoma dataset, leaving 498 thyroid papillary carcinoma samples that met the criteria. These 498 PTC patients were randomly divided into training and test sets in a 7:3 ratio. Univariate Cox regression analysis was used in the training set to initially screen genes for PTC progression prognosis. Using p < 0.05 as the screening criterion, 18 liquid-liquid phase separation-related genes were selected as candidate genes for thyroid papillary carcinoma progression. Specific information is shown in Table 1.
[0057] Table 1
[0058] Gene name p-value HR Gene name p-value HR TCEB3B 0.011 20856.203 MAST1 0.01 2.013 TRIM71 0.027 34.761 IQGAP3 0.003 1.935 GPR156 0.002 6.912 MAPK13 0.033 1.899 TNNT2 0.003 3.566 ABCA12 0.013 1.887 EN1 0.02 3.537 ANLN 0.009 1.809 SPC24 0.001 3.408 PAQR6 0.022 1.657 AZGP1 0.003 2.346 FHL1 0.016 0.788 PLXNA3 0.018 2.246 NRGN 0.048 0.66 WT1 0.049 22.039 PGM5 0.004 0.536
[0059] LASSO regression analysis further screens feature genes: The Least Absolute Shrinkage and Selection Operator (LASSO) algorithm, also known as the lasso algorithm, is a machine learning method based on linear regression. The LASSO algorithm performs L1 regularization shrinkage on the regression coefficients of variables, eliminating variables with zero regression coefficients and leaving only those with non-zero coefficients. LASSO regression is particularly suitable for dimensionality reduction and feature selection of high-dimensional, small-sample data such as gene expression profiles. LASSO regression analysis was performed using the "glmnet" R package on the training set. Using the minimum λ as a threshold, the optimal model was determined, and further screening for LLPS risk genes associated with the progression of papillary thyroid carcinoma was conducted. The results are as follows... Figure 4As shown in Figure A, the coefficient path diagram of LASSO regression is presented. Each colored curve represents a gene. The upper x-axis represents the number of genes with non-zero coefficients at different λ values in the model, the lower x-axis represents the standardized coefficient vector (log(λ)), and the y-axis represents the regression coefficient value for each gene. Each gene's value ends with its corresponding coefficient. Figure B shows the 10-fold cross-validation curve for LASSO regression analysis. The upper x-axis represents the number of genes at different λ values, the lower x-axis is the logarithm of the penalty coefficient for λ, and the y-axis is the cross-validation error. A smaller y-axis value indicates a better fit. The dashed line on the left represents the λ value at the minimum error (lambda.min), where the model fit is best. The dashed line on the right represents the λ value at the minimum error (lambda.1se), where the model fit is also relatively good, with the fewest included variables. Figure C shows the ROC curve of the model. The blue and red curves are the ROC curves modeled with lambda.min and lambda.1se, respectively. The AUC (area under the ROC curve) modeled with lambda.min is 0.76. The minimum λ was selected as the threshold to construct the optimal regression model, corresponding to 14 genes. The 14 genes used to construct the optimal regression model were extracted and identified as "WT1", "TRIM71", "TNNT2", "TCEB3B", "SPC24", "PLXNA3", "PGM5", "NRGN", "MAST1", "MAPK13", "GPR156", "FHL1", "EN1", and "AZGP1". The results are as follows... Figure 4 As shown.
[0060] (4) Multivariate Cox regression analysis to construct a prognostic model: The "stringr", "survival", "survminer", and "My.stepwise" packages were used in the training set to perform multivariate Cox analysis on the genes obtained after LASSO regression, constructing a liquid-liquid phase separation prognostic risk assessment model. Through multivariate Cox regression analysis, a prognostic model consisting of six genes—"SPC24", "TCEB3B", "PLXNA3", "NRGN", "AZGP1", and "GPR156"—was constructed, and a forest plot was drawn as shown below. Figure 5 As shown.
[0061] 4. Training set model prediction and evaluation
[0062] (1) Calculate the LLPS-related risk score: Calculate the regression coefficients of all non-zero variables in the training set and the corresponding multivariate Cox regression coefficients. Combine the regression coefficients with the mRNA expression levels of the selected genes to calculate the LLPS-related risk score (LRS) for each patient. LRS = ∑Coef(gene) i)*Expr(gene i (Coef represents the coefficient, gene represents the gene variable with a non-zero coefficient, and Expr represents the gene expression level.)
[0063] The risk score of the prognostic model = 1.3929258 × SPC24 expression value + 14.9827438 × TCEB3B expression value + 0.7688884 × PLXNA3 expression value - 0.4758786 × NRGN expression value + 0.6593675 × AZGP1 expression value + 1.20117 × GPR156 expression value.
[0064] (2) Evaluation of the prognostic risk assessment model in the training set: Survival analysis of progression risk was performed on the training set using the R packages "survival" and "survminer": Survival analysis was conducted on the relationship between the high and low expression of 6 core genes and PTC progression. The results showed that increased expression of SPC24, TCEB3B, PLXNA3, and AZGP1 increased the risk of PTC progression, while increased expression of NRGN decreased the risk of PTC progression. Figure 6 (where red represents high expression and gray represents low expression).
[0065] The optimal cutoff value (i.e., threshold) for the LRS score in the training set was calculated, and the threshold was 0.852704422136068. Samples with a risk score > optimal cutoff value were classified as high-risk, and samples with a risk score < optimal cutoff value were classified as low-risk. In the prognostic model, compared with the low-risk group in the training set, the progression-free survival time of the high-risk PTC progression population was significantly shortened, and most progression occurred within 50 months. The six core genes in the prognostic model showed differential expression between the two groups, with SPC24, TCEB3B, PLXNA3, AZGP1, and GPR156 highly expressed in the high-risk group, while NRGN was highly expressed in the high-risk low-risk group. Figure 7The figures, from top to bottom, are as follows: (1) Distribution of high- and low-risk groups in the training set (the horizontal axis represents the sample and number, and the vertical axis represents the risk score; the red part of the curve represents high risk, and the blue part represents low risk); (2) Scatter plot of PTC progression event outcomes and occurrence time for each patient in the high- and low-risk groups in the training set (the horizontal axis represents the sample and number, and the vertical axis represents the progression-free survival time in months; event represents progression-free survival outcome, red dots death represents PTC progression, and blue dots alive represents no PTC progression); (3) Heatmap of differential expression of 6 core genes in the prognostic model in the high- and low-risk groups in the training set (the horizontal axis represents the sample, where red highrisk represents the high-risk group and blue lowrisk represents the low-risk group; the vertical axis represents the gene, each column represents the expression of different genes in the same sample, and each row represents the expression of the same gene in different samples, where red represents increased expression and blue represents decreased expression).
[0066] Survival analysis was performed using progression-free survival as the clinical outcome, and Kaplan-Meier curves were plotted. Figure 8 A, the horizontal axis represents progression-free survival time, the upper vertical axis represents the progression-free survival outcome, and the lower vertical axis represents risk stratification; blue ri = high risk represents high risk, yellow ri = low risk represents low risk, and the number corresponding to each risk stratification represents the number of people who did not experience PTC progression at a certain time point. The results showed that there was a significant difference in PTC progression risk between the high-risk group and the low-risk group, and the progression-free survival time was significantly shortened (p < 0.0001); the high-risk group was more likely to experience PTC progression, and progression mainly occurred in the first 5 years (60 months) after initial treatment. The R package "timeROC" was used to perform receiver operating characteristic (ROC) analysis on the training set to assess progression risk, and time-ROC curves were plotted. Figure 8 B, where the horizontal axis represents specificity and the vertical axis represents sensitivity; the blue curve represents the ROC curve and AUC value for 1 year, the red curve represents the ROC curve and AUC value for 3 years, and the yellow curve represents the ROC curve and AUC value for 5 years. As shown, the AUC of the prognostic model at 1 year, 3 years, and 5 years are 0.71, 0.74, and 0.83, respectively, indicating that the prognostic model has good predictive performance on the training set.
[0067] 5. Test Set Model Prediction and Evaluation
[0068] (1) Calculate the LLPS-related risk score: Calculate the regression coefficients of all non-zero variables in the validation set and the multivariate Cox regression coefficients. Combine the regression coefficients with the mRNA expression level of the selected genes to calculate the LLPS-related risk score (LRS) for each patient.
[0069] (2) Evaluation of the prognostic risk assessment model for the test set: Survival analysis of the progression risk was performed on the test set using the R packages "survival" and "survminer". Results ( Figure 9 The prognostic model showed that the progression-free survival (PFS) time was significantly shorter in the high-risk group of PTC patients in the test set compared with that in the low-risk group, and most progressions occurred within 50 months. The six core genes in the prognostic model were differentially expressed between the two groups, with SPC24, TCEB3B, PLXNA3, AZGP1 and GPR156 being highly expressed in the high-risk group, while NRGN was lowly expressed in the high-risk group.
[0070] Survival analysis was performed using progression-free survival as the clinical outcome, and Kaplan-Meier curves were plotted. Figure 10 A, the horizontal axis represents progression-free survival (PFS), the upper vertical axis represents the PFS outcome, and the lower vertical axis represents risk stratification; blue ri = high risk represents high risk, yellow ri = low risk represents low risk, and the number corresponding to each risk stratification represents the number of people who did not experience PTC progression at a certain time point. Results showed a significant difference in PTC progression risk between the high-risk and low-risk groups, and a significantly shorter PFS time (p = 0.011); the high-risk group was more likely to experience PTC progression, with progression mainly occurring 50 months (4.17 years) after initial treatment. Receiver operating characteristic (ROC) analysis of progression risk was performed on the validation set using the R package "timeROC". Figure 10 B, where the horizontal axis represents specificity and the vertical axis represents sensitivity; the blue curve represents the 1-year ROC curve and AUC value, the red curve represents the 3-year ROC curve and AUC value, and the yellow curve represents the 5-year ROC curve and AUC value. The results show that the AUC of the prognostic model at 1 year, 3 years, and 5 years are 0.71, 0.66, and 0.65, respectively. In summary, the predictive performance of the prognostic model on the test set is consistent with that on the training set. This demonstrates that the gene markers and predictive model of this invention can accurately predict the risk of prognostic progression in papillary thyroid carcinoma.
[0071] 6. Statistical Analysis
[0072] All statistical analyses and chart visualizations were performed using R software (version 4.0.1). The correlation between clinical characteristics between the training and test sets was analyzed using the χ² method. 2Chi-square test or Fisher's exact test. Wilcoxon rank-sum test was used to estimate and test the distribution of selected genes among groups. Differential analysis was performed using the Deseq2 package. GO and KEGG analyses were performed using the clusterProfiler package, and a LASSO model was built based on glmnet. ROC curves were plotted using the pROC package in R. The survival package in R was used to plot Kaplan-Meier curves and Cox proportional hazards regression analysis. Forest plots were drawn using the forestplot package in R. p < 0.05 was considered statistically significant.
[0073] The above description is merely an embodiment of the present invention, and the scope of protection of the present invention is not limited to these specific embodiments, but is determined by the claims of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the technical concept and principle of the present invention should be included within the scope of protection of the present invention.
Claims
1. A genetic biomarker for predicting the risk of progression of papillary thyroid carcinoma, characterized in that, The biomarkers are SPC24, TCEB3B, PLXNA3, AZGP1, GPR156, and NRGN. High expression of SPC24, TCEB3B, PLXNA3, AZGP1, and GPR156 increases the risk of progression of papillary thyroid carcinoma, while low expression of NRGN increases the risk of progression of papillary thyroid carcinoma.
2. The application of the reagent for detecting the expression of the gene marker as described in claim 1 in any of the following aspects: A1) Application in the preparation of kits for predicting the risk of progression of papillary thyroid carcinoma; A2) Application in the development of a system for predicting the risk of progression of papillary thyroid carcinoma.
3. A kit for predicting the risk of progression of papillary thyroid carcinoma, characterized in that, The kit includes reagents for detecting the expression of the gene marker of claim 1.
4. The reagent kit according to claim 3, characterized in that, The kit includes reagents for detecting the mRNA expression and / or protein expression of the genetic marker.
5. The reagent kit according to claim 3, characterized in that, The kit includes any one of the following: B1) A gene or its transcript or protein, wherein the sequence of the transcript or protein is the sequence of one or more or all transcripts or proteins of the corresponding gene; B2) Primer molecules and / or probe molecules that hybridize with the gene or its transcript; B3) Immune molecules of the protein expression products of genes.
6. The reagent kit according to claim 3, characterized in that, The kit is suitable for tissue samples.
7. An assessment system for predicting the risk of progression of papillary thyroid carcinoma, characterized in that, The system includes at least the following modules: The data acquisition module is used to acquire at least the relative expression levels of target genes, wherein the target genes are SPC24, TCEB3B, PLXNA3, AZGP1, GPR156 and NRGN; The data processing module is at least used to calculate a score according to a formula, wherein the formula is: Risk Score = 1.3929258 × SPC24 expression value + 14.9827438 × TCEB3B expression value + 0.7688884 × PLXNA3 expression value - 0.4758786 × NRGN expression value + 0.6593675 × AZGP1 expression value + 1.20117 × GPR156 expression value; The judgment module is at least used to compare the score obtained by the data processing module with a threshold. If the score is higher than the threshold, the group is considered high-risk; if the score is lower than the threshold, the group is considered low-risk. The threshold is 0.852704422136068. A data output module, which is at least used to output results.
8. The system according to claim 7, characterized in that, The system also includes a sample extraction module, which is used at least to extract nucleic acids from samples.
9. The system according to claim 8, characterized in that, The sample is a tissue sample.
10. The use of the system according to any one of claims 7-9 in the preparation of instruments, devices or kits for predicting the risk of progression of papillary thyroid carcinoma.