Gene marker for predicting prognosis risk of papillary thyroid cancer and application of gene marker
By detecting the expression of SPC24, TCEB3B, PLXNA3, AZGP1, GRP156 and NRGN genes, a kit and evaluation system for predicting the prognosis and progression risk of papillary thyroid cancer were constructed, which solved the problem of lack of biomarker-assisted diagnosis in existing technologies and achieved accurate prediction of the prognosis and progression risk of papillary thyroid cancer and personalized treatment.
Patent Information
- Application Number
- CN202510762697.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing technologies have not fully studied the regulatory role of liquid-liquid phase separation (LLPS) in papillary thyroid cancer, making it difficult to effectively predict its recurrence risk. Existing prevention measures mainly rely on drug intervention and imaging monitoring, and lack biomarker-assisted diagnosis and prognosis assessment.
SPC24, TCEB3B, PLXNA3, AZGP1, GRP156 and NRGN genes were used as markers. By detecting their expression levels, a kit and evaluation system for predicting the prognosis and progression risk of papillary thyroid cancer were constructed, including data acquisition, processing and judgment modules, and a formula was used to calculate the score to distinguish between high-risk and low-risk groups.
It provides an accurate method for predicting the risk of prognosis progression of papillary thyroid cancer, helping to develop personalized treatment plans, improve patient survival and reduce the risk of recurrence.
Smart Images

Figure CN120758625A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gene technology, and in particular to a gene marker for predicting the prognosis and progression risk of papillary thyroid cancer and its application. Background Art
[0002] The incidence of thyroid cancer continues to rise worldwide, primarily driven by an increase in PTC. Papillary thyroid carcinoma (PTC), the most common subtype of thyroid cancer, accounts for 96.0% of all new thyroid cancer cases and 66.8% of deaths, and has an extremely high recurrence rate. Reducing thyroid cancer recurrence is a key goal after treatment, especially for differentiated thyroid cancer (papillary and follicular). Current methods for preventing thyroid cancer recurrence primarily rely on drug intervention and imaging surveillance.
[0003] PTC recurrence is associated with numerous factors. Independent risk factors for recurrence have been shown to include tumor size, multifocality, extrathyroidal spread, and lymph node metastasis. Furthermore, extranodal extension, metastatic lymph node ratio, and non-initial treatment are potential risk factors for recurrence. In addition to pathological and clinical features that can predict a poor prognosis in PTC, studies have found that biomolecular markers are also associated with the risk of PTC recurrence.
[0004] An analysis based on the TCGA database identified mRNAs such as FN1 and ITGα3, miRNAs such as miR-486 and miR-1179, and 59 genetic variants as potential biomarkers associated with PTC recurrence. The WHO recently made significant adjustments to the classification of TC, emphasizing the differences in cell origin, pathological characteristics, molecular classification, and biological behavior between benign, low-risk, and malignant tumors. It also emphasized the role of biomarkers in aiding diagnosis and determining prognosis. Preventing recurrence is crucial for improving survival. Identifying risk factors for PTC recurrence and implementing interventions can help prevent recurrence.
[0005] In eukaryotic cells, biological activities are carried out in distinct cellular compartments, or organelles. In addition to typical membrane-bound organelles, cells also contain multifunctional membrane-less organelles. Liquid-liquid phase separation (LLPS), a biophysical process, elucidates the assembly and formation of membrane-less organelles. LLPS is causally linked to many dysregulated cellular processes in cancer, participating in tumorigenesis by promoting cancer cell proliferation and metastasis, helping cancer cells evade growth suppression and immune destruction, and inducing apoptosis under stress.
[0006] It has been reported that LLPS is involved in regulating various malignant tumors such as lung cancer, liver cancer, breast cancer, and prostate cancer. However, there have been no public research reports on whether LLPS is involved in regulating PTC. Summary of the Invention
[0007] In response to the gaps in the prior art, the present invention provides a gene marker for predicting the risk of prognosis and progression of papillary thyroid cancer, wherein the marker includes at least one of SPC24, TCEB3B, PLXNA3, AZGP1, GRP156 and NRGN.
[0008] Furthermore, high expression of at least one of the gene markers SPC24, TCEB3B, PLXNA3, AZGP1, and GRP156 increases the risk of progression of papillary thyroid carcinoma; low expression of NRGN increases the risk of progression of papillary thyroid carcinoma.
[0009] The second invention of the present invention provides the use of the gene marker in any of the following aspects:
[0010] A1) Use in the preparation of a kit for predicting the risk of prognosis and progression of papillary thyroid cancer;
[0011] A2) Application in preparing a system for predicting the risk of prognosis and progression of papillary thyroid cancer.
[0012] In a third aspect, the present invention provides a kit for predicting the risk of prognosis and progression of papillary thyroid cancer, wherein the kit comprises a reagent for detecting the gene marker.
[0013] Furthermore, the kit includes reagents for detecting mRNA expression and / or protein expression of a gene.
[0014] Furthermore, the kit further comprises at least one of a transfection reagent and a carrier RNA.
[0015] Furthermore, the reagents for detecting mRNA expression include at least one of RNA extraction reagents, mRNA purification reagents, cDNA synthesis reagents, qPCR / RT-PCR detection reagents, and high-throughput sequencing library construction reagents; the reagents for detecting protein expression include at least one of protein extraction reagents, immunoassay reagents, protein purification reagents, and protein labeling reagents.
[0016] Furthermore, the reagents in the kit include any one of the following:
[0017] B1) a gene or its transcript or protein, wherein the sequence of the transcript or protein is the sequence of one or more or all transcripts or proteins of the corresponding gene;
[0018] B2) primer molecules and / or probe molecules that hybridize to the gene or its transcript;
[0019] B3) Immune molecules of the protein expression products of the genes.
[0020] Furthermore, the kit is applicable to samples containing nucleic acid. Preferably, the sample includes at least one of a blood sample, a tissue sample, a body fluid sample, an exfoliated cell sample, a stool sample, and a urine sample.
[0021] A fourth aspect of the present invention provides an evaluation system for predicting the prognosis and progression risk of papillary thyroid cancer, the system comprising at least the following modules:
[0022] A data acquisition module, the data acquisition module being used at least to obtain the relative expression level of the gene marker according to claim 1;
[0023] A data processing module, the data processing module is at least used to calculate a score according to a formula, wherein the formula is: risk score = 1.3929258 × SPC24 expression value + 14.9827438 × TCEB3B expression value + 0.7688884 × PLXNA3 expression value - 0.4758786 × NRGN expression value + 0.6593675 × AZGP1 expression value + 1.20117 × GPR156 expression value;
[0024] a judgment module, the judgment module being at least configured to compare the score obtained by the data processing module with a threshold, where a score higher than the threshold is a high-risk group; a score lower than the threshold is a low-risk group, wherein the threshold is 0.852704422136068;
[0025] A data output module is used at least to output a result.
[0026] Furthermore, the system further comprises a sample extraction module, which is at least used to extract nucleic acid from the sample.
[0027] The fifth aspect of the present invention provides the use of the system in preparing an instrument, device or kit for predicting the prognosis and progression risk of papillary thyroid cancer.
[0028] The beneficial effects of the present invention include but are not limited to:
[0029] The present invention provides six genes for predicting the prognosis and progression risk of papillary thyroid cancer and a system for predicting the prognosis and progression risk of papillary thyroid cancer. The system can accurately predict the prognosis and progression risk of papillary thyroid cancer and provide a reference basis for the prognosis and prevention of papillary thyroid cancer patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0031] Figure 1 1 is a schematic diagram of the differentially expressed genes between papillary thyroid carcinoma and normal tissues in an embodiment of the present invention, wherein blue represents low-expression genes and red represents high-expression genes;
[0032] Figure 2 Schematic diagram of differentially expressed genes related to liquid-liquid phase separation between PTC and adjacent normal tissues in the TCGA cohort in an embodiment of the present invention;
[0033] Figure 3 Schematic diagram of the functional enrichment analysis results of differentially expressed genes in an embodiment of the present invention;
[0034] Figure 4 1 is a schematic diagram of the results of further screening characteristic genes by LASSO regression analysis in an embodiment of the present invention, wherein Figure A is a coefficient path diagram of LASSO regression; Figure B is a 10-fold cross-validation curve of LASSO regression analysis; and Figure C is a ROC curve of the model;
[0035] Figure 5 is a forest plot drawn by the liquid-liquid phase separation prognostic risk assessment model in an embodiment of the present invention;
[0036] Figure 6 Schematic diagram of the survival analysis results of the prognostic risk assessment model in an embodiment of the present invention;
[0037] Figure 7 Schematic diagram of the results of calculating the optimal cutoff value of the LRS score of the training set in an embodiment of the present invention, wherein Figure A is a graph showing the distribution of risk scores for the high-risk and low-risk groups of the training set; Figure B is a scatter plot of the PTC progression event outcomes and occurrence time of each patient in the high-risk and low-risk groups of the training set; Figure C is a heat map of the differential expression of the six core genes in the prognostic model in the high-risk and low-risk groups of the training set;
[0038] Figure 8 Schematic diagram of the results of survival analysis using progression-free survival as the clinical outcome in the embodiments of the present invention, wherein Figure A is a Kaplan-Meier curve and Figure B is a time-ROC curve;
[0039] Figure 9 Figure 1 is a schematic diagram of the survival analysis results of the prognostic risk assessment model for the test set in an embodiment of the present invention, wherein Figure A is a graph showing the risk score distribution of the high-risk and low-risk groups in the training set; Figure B is a scatter plot of the PTC progression event outcomes and occurrence times of each patient in the high-risk and low-risk groups in the training set; and Figure C is a heat map showing the differential expression of the six core genes in the prognostic model in the high-risk and low-risk groups in the training set.
[0040] Figure 10Schematic diagram of the results of survival analysis performed as a clinical outcome for testing progression-free survival in an embodiment of the present invention, wherein Figure A is a Kaplan-Meier curve and Figure B is a time-ROC curve. DETAILED DESCRIPTION
[0041] The present invention is described in detail below with reference to the examples, but the present invention is not limited to these examples. Unless otherwise specified, the raw materials and catalysts in the examples of the present invention are purchased through commercial channels.
[0042] Example 1: Building an evaluation system
[0043] 1. Data Collection
[0044] (1) Data related to papillary thyroid carcinoma in the TCGA database: RNA-seq data, survival information, and clinical information of the THCA cohort were downloaded from the Cancer Genome Atlas database (TCGA database) using the code in R software (download code is
[0045] proj="TCGA-THCA"
[0046] if(F){download.file(url=paste0("https: / / gdc.xenahubs.net / download / ",proj,".htseq_counts.tsv.gz"),destfile=paste0(proj,".htseq_counts.tsv.gz"))
[0047] download.file(url=paste0("https: / / gdc.xenahubs.net / download / ",proj,".GDC_phenotype.tsv.gz"),destfile=paste0(proj,".GDC_phenotype.tsv.gz"))
[0048] download.file(url=paste0("https: / / gdc.xenahubs.net / download / ",proj,".survival.tsv"),destfile=paste0(proj,".survival.tsv"))}.
[0049] With complete clinical information and survival time greater than or equal to 30 days as inclusion criteria, clinical data of 56 normal thyroid tissues and 498 papillary thyroid carcinoma tissues in the THCA cohort were obtained.
[0050] (2) Liquid-liquid phase separation-related genes: Liquid-liquid phase separation-related genes (LLPS-related genes, LRGs) were obtained from the PhaSepDB database (http: / / db.phasep.pro / ). First, the files under the phase separation and membraneless organelle entries were downloaded, and the genes therein were extracted using R software and duplicate genes were removed. Finally, 5541 genes were obtained as liquid-liquid phase separation-related genes for subsequent analysis: The THCA cohort transcriptome data in the TCGA database includes 498 papillary thyroid carcinoma samples and 56 adjacent normal samples and approximately 30956 genes. The probes of the differentially expressed genes were annotated, and the fold change (FC) was greater than 1 and P < 0.05 was used as the screening criteria. A total of 5177 genes were differentially expressed between papillary thyroid carcinoma and normal tissues; among them, 2800 genes were upregulated in PTC and 2377 genes were downregulated in PTC. The results are as follows Figure 1 shown.
[0051] 2. Data Processing
[0052] (1) Data preprocessing: R software was used to perform count scaling at the probe level and filter out genes with zero or low expression in the samples, retaining only genes expressed in more than half of the samples. The “DESeq2” R package was used to normalize each dataset and remove batch effects. gtf files were downloaded, and the RNA sequencing data of the TCGA cohort were annotated to obtain the expression matrix of all genes. The expression matrix was matched with clinical information, and only one sample was retained for each patient. The matched data were renamed and classified according to clinical characteristics, retaining only clinical information such as “gender”, “race_list”, “age_at_initial_pathologic_diagnosis”, “stage_event_pathologic_stage”, “stage_event_tnm_categories”, “PFS”, and “PFS.time”. The clinical characteristics of each sample were processed according to the AJCC (American Joint Committee on Cancer) stage and whether there was extrathyroidal infiltration, lymph node metastasis, and distant metastasis. The GEO database did not retrieve a data set containing the clinical endpoint of "Progression-free Survival (PFS)".
[0053] (2) Screening of differentially expressed genes: The “DESeq2” R package was used with |log2fold change|>1 and False Discovery Rate (FDR) <0.05 as threshold conditions to identify differentially expressed genes (DEGs) between papillary thyroid carcinoma tissues and adjacent normal tissues in the TCGA cohort. The obtained differentially expressed genes were intersected with the liquid-liquid phase separation-related genes to obtain the differentially expressed genes related to liquid-liquid phase separation between PTC and adjacent normal tissues in the TCGA cohort ( Figure 2 ). It contains 318 genes, of which 195 genes are upregulated in papillary thyroid carcinoma tissues and 124 genes are downregulated in papillary thyroid carcinoma tissues.
[0054] (3) Functional enrichment analysis of differentially expressed genes: The “clusterProfiler” R package in R language was used to explore the biological pathways and functions of differentially expressed genes related to liquid-liquid phase separation. First, the Kyoto Encyclopedia of Genes and Genomes (KEGG) was used to evaluate the biological pathways associated with differentially expressed genes. Then, based on Gene Ontology (GO), further functional analysis was performed on three types of biological processes regulated by DEGs, namely cellular components (CC), molecular functions (MF), and biological processes (BP). GO annotation and KEGG enrichment analysis were performed on 318 liquid-liquid phase separation-related differentially expressed genes in papillary thyroid carcinoma. Figure A shows the top 10 results of GO annotation. The results of GO annotation showed that the liquid-liquid phase separation-related genes in papillary thyroid carcinoma were mainly distributed in cellular components such as cytoskeleton, cell-basement membrane connection, contractile fiber cells, myofibrils, and sarcomeres. They mainly regulated molecular functions such as actin binding, protein kinase activity, transmembrane receptor kinase activity, growth factor binding, and cytoskeleton structure formation, and participated in biological processes such as cell migration, protein post-translational modification, tissue migration, cell development, epithelial cell development, and promotion of cell protrusion formation. Figure B KEGG enrichment analysis results showed that the phase separation-related genes differentially expressed in papillary thyroid carcinoma were mainly involved in muscle cytoskeleton, MAPK signaling pathway, PI3K-Akt signaling pathway, transcriptional disorders in cancer, nucleotide metabolism, purine metabolism, acute myeloid leukemia, bladder cancer, virions-Lassa virus and SFTS virus, axon guidance and other pathways. The results are as follows Figure 3 shown.
[0055] 3. Construction of a liquid-liquid phase separation prognostic model
[0056] (1) Obtaining candidate genes for the prognostic model: The “survival” R package was used in the training set to perform univariate Cox regression analysis on differentially expressed genes related to liquid-liquid phase separation. Progression-free survival (PFS) was used as the clinical endpoint; the follow-up time was converted to months, and genes with a hazard ratio (HR) greater than 1 and p < 0.05 were screened as candidate genes for modeling: normal tissue samples in the TCGA papillary thyroid carcinoma dataset and papillary thyroid carcinoma samples that did not meet the inclusion criteria were excluded. A total of 498 papillary thyroid carcinoma samples met the criteria. After the 498 PTC patients were randomly divided into a training set and a test set at a ratio of 7:3, univariate Cox regression analysis was performed in the training set to preliminarily screen PTC progression prognostic genes. Using P < 0.05 as the screening criterion, a total of 18 genes related to liquid-liquid phase separation were screened as candidate genes for papillary thyroid carcinoma progression. The specific information is shown in Table 1.
[0057] Table 1
[0058] Gene name P-value HR Gene name P-value HR TCEB3B 0.011 20856.203 MAST1 0.01 2.013 TRIM71 0.027 34.761 IQGAP3 0.003 1.935 GPR156 0.002 6.912 MAPK13 0.033 1.899 TNNT2 0.003 3.566 ABCA12 0.013 1.887 EN1 0.02 3.537 ANLN 0.009 1.809 SPC24 0.001 3.408 PAQR6 0.022 1.657 AZGP1 0.003 2.346 FHL1 0.016 0.788 PLXNA3 0.018 2.246 NRGN 0.048 0.66 WT1 0.049 22.039 PGM5 0.004 0.536
[0059] LASSO regression analysis further screens characteristic genes: The least absolute shrinkage and selection operator (LASSO) algorithm, also known as the Lasso algorithm, is a machine learning method based on linear regression. The LASSO algorithm can perform L1 regularization shrinkage on the variable regression coefficients. Variables with zero regression coefficients will be eliminated, leaving only variables with non-zero regression coefficients. LASSO regression is particularly suitable for dimensionality reduction and feature selection of high-dimensional, small sample data such as gene expression profiles. The "glmnet" R package was used to perform LASSO regression analysis in the training set. The minimum λ was used as the threshold to determine the best model, and further screened the LLPS risk genes related to the progression of papillary thyroid carcinoma. The results are as follows Figure 4As shown. Figure A is the coefficient path diagram of LASSO regression, where each colored curve represents a gene, the upper horizontal axis represents the number of genes with non-zero coefficients at different λ in the model, the lower horizontal axis log(λ) represents the standardized coefficient vector, the vertical axis represents the value of the regression coefficient of each gene, and the end of each gene points to its own coefficient. Figure B is the 10-fold cross-validation curve of LASSO regression analysis, where the upper horizontal axis represents the number of genes at different λ, the lower horizontal axis is the logarithm of the penalty coefficient of λ, and the vertical axis is the cross-validation error. The smaller the vertical axis, the better the fit. The left dotted line represents the λ value with the smallest error, i.e. lambda.min. At this time, the model fits the best. The right dotted line represents the λ value at 1 standard error of the smallest error, lambda.1se. At this time, the model fit is also relatively good and the least variables are included. Figure C is the ROC curve of the model. The blue and red curves are the ROC curves modeled with lambda.min and lambda.1se, respectively. Among them, the AUC (area under the ROC curve) modeled with lambda.min is 0.76. The minimum λ was selected as the threshold to construct the optimal regression model, and the corresponding number of genes was 14. The 14 genes extracted to construct the optimal regression model were "WT1", "TRIM71", "TNNT2", "TCEB3B", "SPC24", "PLXNA3", "PGM5", "NRGN", "MAST1", "MAPK13", "GPR156", "FHL1", "EN1", and "AZGP1". The results are shown in the figure. Figure 4 shown.
[0060] (4) Multivariate COX regression analysis to construct a prognostic model: In the training set, the "stringr", "survival", "survminer" and "My.stepwise" packages were used to perform multivariate COX analysis on the genes obtained after LASSO regression, and a liquid-liquid phase separation prognostic risk assessment model was constructed: After multivariate COX regression analysis, a prognostic model consisting of six genes, "SPC24", "TCEB3B", "PLXNA3", "NRGN", "AZGP1" and "GPR156", was constructed, and a forest plot was drawn as shown in Figure 2. Figure 5 shown.
[0061] 4. Training set model prediction and evaluation
[0062] (1) Calculation of LLPS-related risk score: Calculate the corresponding regression coefficients of all non-zero variables in the training set and the multi-factor COX regression coefficients, combine the regression coefficients with the mRNA expression levels of the selected genes, and calculate the LLPS-related risk score (LRS) for each patient. LRS = ∑Coef(gene i)*Expr(gene i )(Coef represents coefficient, gene represents gene variable with non-zero coefficient, and Expr represents gene expression).
[0063] The risk score of the prognostic model = 1.3929258 × SPC24 expression value + 14.9827438 × TCEB3B expression value + 0.7688884 × PLXNA3 expression value - 0.4758786 × NRGN expression value + 0.6593675 × AZGP1 expression value + 1.20117 × GPR156 expression value.
[0064] (2) Evaluation of the prognostic risk assessment model of the training set: The R packages "survival" and "survminer" were used to perform survival analysis of the progression risk of the training set: The relationship between the high and low expression of 6 core genes and PTC progression was analyzed for survival. The results showed that increased expression of SPC24, TCEB3B, PLXNA3, and AZGP1 increased the risk of PTC progression, while increased expression of NRGN decreased the risk of PTC progression ( Figure 6 ), where red represents high expression and gray represents low expression.
[0065] The optimal cutoff value (i.e., threshold) of the LRS score of the training set was calculated, and the threshold was 0.852704422136068. Samples with a risk score > the optimal cutoff value were classified as the high-risk group, and samples with a risk score < the optimal cutoff value were classified as the low-risk group. The prognostic model showed that the progression-free survival time of the high-risk group of PTC progression in the training set was significantly shortened compared with the low-risk group, and most of the progression occurred within 50 months; the 6 core genes in the prognostic model were differentially expressed between the two groups, among which SPC24, TCEB3B, PLXNA3, AZGP1, and GPR156 were highly expressed in the high-risk group, while NRGN was highly expressed in the high-risk group ( Figure 7). Among them, the figures from top to bottom are (1) the distribution of high and low risk groups in the training set (the horizontal coordinate represents the sample and the number, and the vertical coordinate represents the risk score; the red part of the curve represents the high risk, and the blue part represents the low risk) (2) the scatter plot of PTC progression event outcome and occurrence time of each patient in the high and low risk groups of the training set (the horizontal coordinate represents the sample and the number, and the vertical coordinate represents the progression-free survival time in months; event represents the progression-free survival outcome, and the red dot death represents PTC progression, and the blue dot alive represents no PTC progression) (3) the differential expression heat map of the 6 core genes in the prognosis model in the high and low risk groups of the training set (the horizontal coordinate is the sample, in which red highrisk represents the high risk group, and blue lowrisk represents the low risk; the vertical coordinate is the gene, and each column represents the expression of different genes in the same sample, and each row represents the expression of the same gene in different samples, in which red represents the increase of expression, and blue represents the decrease of expression).
[0066] Survival analysis was performed with progression-free survival as the clinical outcome, and the Kaplan-Meier curve was drawn. Figure 8 A, the horizontal coordinate represents the progression-free survival time, the upper vertical coordinate represents the outcome of progression-free survival, and the lower vertical coordinate represents the risk stratification; blue ri=highrisk represents the high risk, and yellow ri=lowrisk represents the low risk, and the number corresponding to each risk stratification represents the number of patients without PTC progression at a certain time point), the results showed that there was a significant difference in PTC progression risk between the high risk group and the low risk group, and the progression-free survival time was significantly shortened (p<0.0001); the high risk group was more likely to have PTC progression, and the progression mainly occurred in the first 5 years (60 months) after the initial treatment. The progression risk receiver operating characteristic (ROC) analysis of the training set was performed using the R package "timeROC", and the time-ROC curve was drawn. Figure 8 B, the horizontal coordinate represents specificity, and the vertical coordinate represents sensitivity, the blue curve represents the ROC curve and AUC value of 1 year, the red curve represents the ROC curve and AUC value of 3 years, and the yellow curve represents the ROC curve and AUC value of 5 years); as shown, the AUC of the prognosis model at 1 year, 3 years and 5 years was 0.71, 0.74 and 0.83, respectively, indicating that the prediction performance of the prognosis model in the training set was good.
[0067] 5. Test set model prediction and evaluation
[0068] (1) Calculate the LLPS-related risk score: calculate the non-zero variable coefficients of all factors and the corresponding regression coefficients of the multivariate COX regression, combine the regression coefficients with the mRNA expression level of the selected gene, and calculate the LLPS-related risk score (LRS) of each patient.
[0069] (2) Evaluation of the prognostic risk assessment model of the test set: The R packages “survival” and “survminer” were used to perform survival analysis of the progression risk on the test set. Figure 9 ) showed that the prognostic model significantly shortened the PFS time of the high-risk group of PTC patients in the test set compared with the low-risk group, and most of the progression occurred within 50 months; the six core genes in the prognostic model were differentially expressed between the two groups, among which SPC24, TCEB3B, PLXNA3, AZGP1 and GPR156 were highly expressed in the high-risk group, while NRGN was lowly expressed in the high-risk group.
[0070] Survival analysis was performed with progression-free survival as the clinical outcome, and Kaplan-Meier curves were drawn. Figure 10 A, the horizontal axis represents progression-free survival time, the upper vertical axis represents the outcome of progression-free survival, and the lower vertical axis represents risk stratification; blue ri=highrisk represents high risk, yellow ri=lowrisk represents low risk, and the number corresponding to each risk stratification represents the number of people who did not have PTC progression at a certain time point). The results showed that there was a significant difference in the risk of PTC progression between the high-risk group and the low-risk group, and the PFS time was significantly shortened (p=0.011); the high-risk group was more likely to have PTC progression, and the progression time mainly occurred 50 months after initial treatment (4.17 years). The R package "timeROC" was used to perform receiver operating characteristic (ROC) analysis of progression risk on the validation set. ( Figure 10 B, the horizontal axis represents specificity, the vertical axis represents sensitivity, the blue curve represents the ROC curve and AUC value of 1 year, the red curve represents the ROC curve and AUC value of 3 years, and the yellow curve represents the ROC curve and AUC value of 5 years). The results show that the AUC of the prognostic model at 1 year, 3 years, and 5 years are 0.71, 0.66, and 0.65, respectively. In summary, the predictive performance of the prognostic model in the test set is consistent with that of the training set. It is proved that the gene markers and prediction model of the present invention can accurately predict the prognosis and progression risk of papillary thyroid cancer.
[0071] 6. Statistical Analysis
[0072] All statistical analyses and graphical visualizations were performed using R software (version 4.0.1). The correlation between clinical characteristics of the training set and the test set was performed using the χ 2Chi-square test or Fisher's exact test. Wilcoxon rank-sum test was used to estimate and test the distribution of selected genes between groups. Differential analysis was performed using the Deseq2 package. GO analysis and KEGG analysis were performed using the clusterProfiler package, and LASSO models were constructed based on glmnet. ROC curves were plotted using the R package pROC. The R package survival was used to plot the Kaplan-Meier curve and Cox proportional hazards regression analysis. Forest plots were drawn using the R package forestplot. p<0.05 was considered statistically significant.
[0073] The above merely illustrates the embodiments of the present application, and the protection scope of the present application is not limited by these specific embodiments, but is determined by the claims of the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the technical idea and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A gene marker for predicting the risk of prognosis and progression of papillary thyroid cancer, characterized in that: The markers include at least one of SPC24, TCEB3B, PLXNA3, AZGP1, GRP156 and NRGN.
2. The gene marker according to claim 1, characterized in that High expression of at least one of SPC24, TCEB3B, PLXNA3, AZGP1, and GRP156 increases the risk of progression of papillary thyroid carcinoma; low expression of NRGN increases the risk of progression of papillary thyroid carcinoma.
3. Use of the gene marker according to claims 1-2 in any of the following aspects: A1) Use in the preparation of a kit for predicting the risk of prognosis and progression of papillary thyroid cancer; A2) Application in preparing a system for predicting the risk of prognosis and progression of papillary thyroid cancer.
4. A kit for predicting the risk of prognosis and progression of papillary thyroid cancer, characterized in that: The kit comprises a reagent for detecting the gene marker according to claim 1 or 2.
5. The kit according to claim 4, characterized in that The kit includes reagents for detecting the mRNA expression and / or protein expression of a gene.
6. The kit according to claim 4, wherein The kit includes any one of the following: B1) a gene or its transcript or protein, wherein the sequence of the transcript or protein is the sequence of one or more or all transcripts or proteins of the corresponding gene; B2) primer molecules and / or probe molecules that hybridize to the gene or its transcript; B3) Immune molecules of the protein expression products of the genes.
7. The kit according to claim 4, characterized in that The kit is suitable for samples containing nucleic acids. Preferably, the sample includes at least one of a blood sample, a tissue sample, a body fluid sample, an exfoliated cell sample, a stool sample, and a urine sample.
8. An evaluation system for predicting the risk of prognosis and progression of papillary thyroid cancer, characterized in that: The system includes at least the following modules: A data acquisition module, the data acquisition module being used at least to obtain the relative expression level of the gene marker according to claim 1; A data processing module, the data processing module is at least used to calculate a score according to a formula, wherein the formula is: risk score = 1.3929258 × SPC24 expression value + 14.9827438 × TCEB3B expression value + 0.7688884 × PLXNA3 expression value - 0.4758786 × NRGN expression value + 0.6593675 × AZGP1 expression value + 1.20117 × GPR156 expression value; a judgment module, the judgment module being at least configured to compare the score obtained by the data processing module with a threshold, where a score higher than the threshold is a high-risk group; a score lower than the threshold is a low-risk group, wherein the threshold is 0.852704422136068; A data output module is used at least to output a result.
9. The system according to claim 8, characterized in that The system further comprises a sample extraction module, which is at least used for extracting nucleic acid from a sample.
10. Use of the system according to any one of claims 8 to 9 in the preparation of an instrument, device or kit for predicting the prognosis and progression risk of papillary thyroid cancer.
Citation Information
Patent Citations
Circulating biomarkers for cancer
CN103782174A
Centromere / kinetochore protein genes for cancer diagnosis, prognosis and treatment selection
CN106661624A
Construction method, system and equipment applied to thyroid tumor evaluation model
CN118116476A
Set of Tumour-Markers
US20110152110A1
Compositions and methods for diagnosing thyroid cancer
US20170349951A1