A set of biomarkers, predictive models and applications for predicting the risk of developing thyroid eye disease
By screening out oxidative stress-related biomarkers KLF2, TXNRD1, CTSD, and PHC2, an OSRDGS scoring model was constructed, which addresses the shortcomings of existing technologies in early screening and risk prediction of thyroid eye disease, and realizes the potential for early risk identification and clinical translation.
Patent Information
- Application Number
- CN202511079671.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Existing technologies lack effective biomarkers for early screening and disease risk prediction of thyroid eye disease. Existing markers are difficult to reveal the pathogenesis and are not suitable for early prediction.
By integrating differentially expressed genes, weighted gene co-expression network analysis, and various machine learning algorithms, we screened out oxidative stress-related biomarkers such as KLF2, TXNRD1, CTSD, and PHC2 from peripheral blood RNA sequencing data, and constructed an OSRDGS scoring model to predict the risk of thyroid eye disease.
It provides a model with high predictive value that can identify the risk of thyroid eye disease at an early stage, raise awareness of disease prevention, and reduce the harm caused by untimely control and treatment.
Smart Images

Figure CN120581064B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical technology, and specifically relates to a group of biomarkers, prediction models and applications for predicting the risk of thyroid eye disease. Background Art
[0002] Thyroid eye disease (TED), also known as Graves' orbitopathy, is the most common orbital disease in adults and one of the most common and severe extrathyroidal manifestations of diffuse toxic goiter (Graves' disease, GD). Approximately 40% of GD patients may develop TED, with clinical manifestations including proptosis, eyelid retraction, strabismus, and diplopia, which can severely impact patients' eye health and quality of life, and may even threaten their vision.
[0003] Currently, the first-line treatment for TED primarily relies on glucocorticoids, but this regimen is often associated with numerous systemic and local adverse reactions. In recent years, a variety of biologic agents, including monoclonal antibodies targeting the insulin-like growth factor-1 receptor (IGF-1R), have provided new avenues for the treatment of TED. However, these novel treatments still present challenges such as variable response rates, significant side effects, and recurrence. Therefore, multiple TED clinical guidelines and expert consensus agree that early prediction of TED risk is necessary, with simple and effective measures to prevent its onset and progression. The development of appropriate disease risk prediction models is a potential effective approach to address this issue. Currently, biomarkers used for TED risk prediction include clinical indicators, humoral markers (such as cytokines and macromolecules in tears and blood), and various imaging markers. However, clinical indicators are often a reflection of TED onset, while protein markers in body fluids are susceptible to environmental and sampling conditions. These downstream candidate biomarkers struggle to reveal pathogenesis and are difficult to use for early prediction of disease risk.
[0004] Therefore, there is an urgent need to develop practical biomarkers for early screening of TED and disease risk prediction in this field to improve patients' quality of life and reduce medical costs. Summary of the Invention
[0005] To address the problem that there are no effective biomarkers for early screening of thyroid eye disease and disease risk prediction in the prior art, the present invention provides a set of biomarkers, prediction models and applications for predicting the risk of thyroid eye disease.
[0006] To achieve the above objectives, the present invention first provides a group of biomarkers for predicting the risk of thyroid eye disease, wherein the biomarkers include: KLF2, TXNRD1, CTSD and PHC2 .
[0007] Another aspect of the present invention provides the use of the aforementioned biomarkers in constructing a prediction model for the risk of thyroid eye disease.
[0008] Another aspect of the present invention provides a prediction model for the risk of thyroid eye disease, wherein the prediction model is OSRDGS = 0.995878 × exp ( KLF2 ) + 0.634031 × exp( TXNRD1 ) + 0.640694 × exp( CTSD ) + 0.767982 × exp( PHC2 ); Among them, exp( KLF2 )、exp( TXNRD1 )、exp( CTSD )、exp( PHC2 ) respectively represent KLF2、TXNRD1 、 CTSD 、 PHC2 The expression level is the exponential function with e as the base, and OSRDGS represents the risk score for thyroid eye disease.
[0009] Alternatively, quantitative PCR can be used to detect KLF2, TXNRD1, CTSD and PHC2 expression level.
[0010] Optionally, when the OSRDGS of the subject's sample is higher, it is predicted that the subject has a higher risk of developing thyroid eye disease; when the OSRDGS of the subject's sample is lower, it is predicted that the subject has a lower risk of developing thyroid eye disease.
[0011] Another aspect of the present invention provides a detection kit for predicting the risk of thyroid eye disease, the kit comprising: KLF2, TXNRD1, CTSD and PHC2 Expression level reagents.
[0012] Optionally, the sample is a blood sample.
[0013] Optionally, when the subject sample KLF2, TXNRD1, CTSD and PHC2 When the expression level is low, the risk of the subject developing thyroid eye disease is low; when the expression level is low, the risk of the subject developing thyroid eye disease is low. KLF2, TXNRD1, CTSD and PHC2 When the expression level is high, the evaluation subject has a higher risk of developing thyroid eye disease.
[0014] Optionally, the kit detects the KLF2, TXNRD1, CTSD and PHC2 expression level.
[0015] Optionally, the kit includes: fluorescent probe / dye; DNA polymerase; dNTPs; specific amplification KLF2, TXNRD1, CTSD and PHC2 primer sets and buffer.
[0016] Compared with the prior art, the beneficial effects of the present invention include at least:
[0017] 1. Based on peripheral blood bulk RNA sequencing data from three groups of patients with TED, healthy controls, and patients with simple GD, the present invention integrated differentially expressed genes, weighted gene co-expression network analysis results, and OS-related gene sets, and further applied multiple machine learning algorithms to screen out four oxidative stress-related biomarkers with disease risk prediction value ( KLF2, TXNRD1, CTSD, and PHC2 ), and based on this, they constructed an OS-related diagnostic gene score (OSRDGS) as a predictive model for the risk of thyroid eye disease. Through multiple evaluation strategies, they verified that the OSRDGS has a high predictive value for the risk of TED disease, and the model score level is positively correlated with disease progression, which has important clinical translation potential.
[0018] 2. The biomarkers and prediction models provided by this invention can help promote early identification and early intervention of TED patients, enhance subjects' awareness of disease risk prevention, and avoid greater harm caused by untimely control and treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of the research of the present invention.
[0020] Figure 2 This is the differential expression analysis among TED, GD and HC groups, where:
[0021] A represents the volcano plot of DEGs between TED and HC groups, red dots represent genes upregulated in TED, and blue dots represent genes downregulated in TED;
[0022] B represents the KEGG functional enrichment analysis between TED and HC groups, showing the significantly enriched pathways;
[0023] C represents the volcano plot of DEGs between TED and GD groups, red dots represent genes upregulated in TED, and yellow dots represent genes downregulated in TED;
[0024] D represents KEGG functional enrichment analysis between TED and GD groups, showing significantly enriched pathways;
[0025] E represents the Venn diagram of differentially expressed genes between TED-HC and TED-GD, showing the intersection of the two;
[0026] F represents the bubble chart display of GO BP enrichment analysis for TED-specific DEGs.
[0027] Figure 3 Identification of key module genes related to TED and oxidative stress; among them:
[0028] A represents the ssGSEA analysis results of 38 TED-related gene sets in the TED, GD, and HC groups;
[0029] B represents the heat map of the correlation between gene modules and different TED pathogenicity-related gene sets;
[0030] C represents the scatter plot of the correlation between blue module genes and OS gene set;
[0031] D represents the KEGG functional enrichment analysis of blue module genes;
[0032] E represents the intersection Venn diagram of TED-specific DEGs, blue module genes and OSRGs.
[0033] Figure 4 To screen candidate biomarkers for TED, including:
[0034] A represents the cross-validation curve of LASSO regression analysis, and the left dotted line represents the λ value with the smallest mean error in the model;
[0035] B represents the regression coefficients of different genes in LASSO regression analysis;
[0036] C represents the lollipop plot showing the gene importance of 14 candidate genes in the random forest model;
[0037] D represents a visualization of the relationship between the overall error rate and the number of decision trees in the random forest analysis;
[0038] E represents the results of univariate logistic regression analysis of 14 candidate genes;
[0039] F represents the intersection Venn diagram of the screening results of the three machine learning models; *p<0.05, **p<0.01.
[0040] Figure 5 To construct and evaluate the OSRDGS TED diagnostic model, which includes:
[0041] AB represent the results of univariate (A) and multivariate (B) logistic regression analysis integrating OSRDGS, age, and gender in the training set;
[0042] CE represents the ROC curves of the training set (C), test set (D), and TED-GD set (E);
[0043] F represents the expression levels of each OSRDGs in the high-OSRDGS group and the low-OSRDGS group;
[0044] G represents the comparison of OSRDGS in patients with mild, moderately severe and extremely severe TED (box plot representation);
[0045] H represents the comparison of OSRDGS between TED patients with abnormal and normal Schirmer test results.
[0046] Figure 6 This is a multidimensional analysis of immune infiltration and functional characteristics between high and low OSRDGS groups, including:
[0047] A represents the comparison of the estimated scores of different cell types between high and low OSRDGS groups in ssGSEA analysis (box plot display);
[0048] B shows the distribution bar graph of the 22 immune cells in each sample, and the samples are sorted according to the OSRDGS score level (high or low);
[0049] C represents the correlation bar graph between OSRDGS and the scores of each immune cell type in ssGSEA analysis, the bar length represents the correlation coefficient, and the color represents the p value;
[0050] D shows the GO functional enrichment analysis results of the differentially expressed gene sets between the high and low OSRDGS groups.
[0051] Figure 7 For the performance evaluation of OSRDGS,
[0052] A represents the correlation heat map of the four marker genes;
[0053] B represents the multivariate logistic regression analysis of OSRDGs;
[0054] CE represents the calibration curve of the training set, test set, and TED-GD set;
[0055] FH represents the DCA of the training set, test set, and TED-GD set. DETAILED DESCRIPTION
[0056] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0057] Terminology
[0058] In the present invention, “OS” refers to oxidative stress (OS), which is caused by the disruption of redox balance and is usually accompanied by the accumulation of excessive reactive oxygen species (ROS).
[0059] In the present invention, "simple GD patients" refer to GD patients who have not developed TED, have no clinical manifestations of TED including proptosis, eyelid retraction, strabismus and diplopia, and have no imaging changes of TED including thickening of extraocular muscles and increased retrobulbar fat.
[0060] Related information of the four marker genes of the present invention (based on NCBI database): KLF2 (Kruppel-like factor2), gene ID: 10365, genomic location: NC_000019.10:16324826-16328685 (+); TXNRD1 (thioredoxin reductase 1), gene ID: 7296, genomic location: NC_000012.12:104215779-104350307 (+); CTSD (Cathepsin D), gene ID: 1509, genomic location: NC_000011.10:1752755-1763927 (-); PHC2 (Polyhomeotic Homolog 2), gene ID: 1912, genome location: NC_000001.11:33323626-33431095 (-).
[0061] As previously mentioned, given the limitations of existing technologies, the present inventors have conducted extensive research and analysis to identify biomarkers for early TED screening and disease risk prediction. Studies have shown that OS is involved in multiple pathophysiological processes, including aging, tumors, and inflammation. Excessive elevation of ROS levels, leading to an enhanced OS response, is often observed in patients with TED and GD, potentially related to the regulatory effects of thyroid hormone on mitochondrial function. Furthermore, studies have found elevated levels of OS-related markers in blood, urine, and tear samples from TED patients. Furthermore, orbital fibroblasts (OFs) derived from TED patients have higher levels of 8-hydroxy-2'-deoxyguanosine (8-OHdG), malondialdehyde (MDA), and intracellular ROS compared to control OFs. Excessive ROS promotes the inflammatory phenotype and pathological differentiation of OFs, thereby triggering orbital inflammation and tissue remodeling in TED. These findings highlight the critical role of OS in the progression of TED and provide a theoretical basis for therapeutic strategies targeting OS. Several studies have shown that intervention targeting OS can effectively alleviate the clinical manifestations of TED. However, most existing studies have used OS as a therapeutic target, while research on TED diagnostic or predictive models based on the pathological mechanisms of OS is still very limited.
[0062] Furthermore, most biomarkers currently used for TED disease risk prediction are downstream candidate markers, such as clinical indicators, body fluid markers, and imaging markers. These downstream candidate markers are difficult to reveal the pathogenesis and are difficult to use for early prediction of disease risk. The transcriptome, on the other hand, can rapidly respond to ongoing physiological and pathological processes in the body and can reflect impending changes in the body, making it considered a more forward-looking disease predictor.
[0063] Based on the above, in order to obtain a set of reliable biomarkers for predicting the risk of TED, the present invention collected and analyzed peripheral blood bulk RNA sequencing data from three groups of people: TED patients, healthy controls (HC), and simple GD patients. By integrating the results of differentially expressed genes (DEGs), weighted gene co-expression network analysis (WGCNA), and OS-related gene sets (OSRGs), 14 candidate biomarkers were screened out. Furthermore, four oxidative stress-related biomarkers with disease risk prediction value were screened out using a machine learning algorithm ( KLF2, TXNRD1, CTSD, and PHC2), and based on this, constructed an OS-related diagnostic gene score (OSRDGS) as a predictive model for the risk of thyroid eye disease. Through multiple evaluation strategies, it was verified that the OSRDGS has a high predictive value for the risk of TED disease, and the model score level is positively correlated with disease progression, which has important clinical translation potential.
[0064] Figure 1 The following is a research flow chart for the present invention. The process and results of the present invention are described in detail in conjunction with experimental data. Unless otherwise indicated, the experimental methods, detection methods, and preparation methods disclosed herein utilize conventional techniques in molecular biology, biochemistry, chromatin structure and analysis, analytical chemistry, cell culture, recombinant DNA technology, and related fields.
[0065] The experimental methods used in the following examples were carried out according to conventional conditions or those recommended by the manufacturers unless otherwise specified. The materials and reagents used in the following examples were all commercially available unless otherwise specified.
[0066] Example 1 Screening and validation of biomarkers for predicting the risk of thyroid eye disease
[0067] 1. Data sources and data analysis methods
[0068] 1. Data Collection
[0069] We obtained peripheral blood bulk RNA sequencing datasets (GSE285190 and GSE280114) and clinical characteristics information previously published in the GEO database, including samples from 152 TED patients, 61 healthy controls (HCs), and 20 simple GD patients.
[0070] We obtained 32 OS-related gene sets and 784 genes from the MSigDB database (https: / / www.gsea-msigdb.org / gsea / msigdb). The gene sets were selected based on the presence of the keyword "OXIDATIVE_STRESS" in their names.
[0071] 2. Differential Expression Analysis
[0072] The DESeq2 package was used to perform differential expression analysis between the TED and HC, or TED and GD, groups. Genes with a log2 fold change (log2FC) greater than 0 and a p-value less than 0.05 were identified as upregulated differentially expressed genes (DEGs). Conversely, genes with a log2FC less than 0 and a p-value less than 0.05 were identified as downregulated DEGs. The results were visualized using volcano plots.
[0073] 3. Functional enrichment analysis
[0074] Functional enrichment analysis was performed using biological process (BP) annotations from the GO database and the KEGG database, with bidirectional bar plots displaying pathways enriched in the TED, GD, and HC groups. Furthermore, gene set enrichment analysis (GSEA) was performed using the Hallmark, Reactome, and KEGG gene sets from MSigDB, and the enrichment results were displayed in GSEA plots.
[0075] 4. Single sample gene set enrichment analysis (ssGSEA)
[0076] After obtaining 38 gene sets related to the pathogenic mechanism of TED from MSigDB, ssGSEA was performed using the GSVA package to quantitatively calculate the TED-related pathogenic phenotype enrichment scores of the 38 gene sets in each sample to explore the peripheral phenotypic differences between TED, GD and HC.
[0077] 5. Weighted gene co-expression network analysis (WGCNA)
[0078] WGCNA was used to identify key gene modules associated with the OS gene set and TED. The analysis was based on a normalized gene expression matrix, selecting the top 5,000 genes by median absolute deviation (MAD), with a final soft threshold of 4. A weighted adjacency matrix was then constructed and converted into a topological overlap matrix (TOM) to generate a hierarchical clustering tree. Genes with similar expression patterns were grouped into the same module. Correlation heatmaps were then plotted between each gene module and the enrichment scores of 38 TED-associated pathogenic pathways. Significantly enriched pathways within each module were identified through Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis and visualized using bar and bubble charts.
[0079] 6. Comprehensive machine learning algorithm screening of key OS-related diagnostic genes (oxidative stress-related diagnostic genes, OSRDG)
[0080] After excluding the influence of GD, the DEGs, most relevant module genes and OS-related genes (OSRGs) of TED and HC were intersected to obtain a set of candidate biomarkers for predicting TED disease risk.
[0081] Subsequently, 152 TED patients and 61 HC samples were divided into training and test sets using stratified sampling in a 7:3 ratio. Feature screening was performed in the training set using three methods: random forest (RF), least absolute shrinkage and selection operator (LASSO), and univariate logistic regression. The RF algorithm identified genes with above-average importance scores; the LASSO algorithm selected the lambda value that minimized the mean squared error; and the univariate logistic regression analysis identified statistically significant genes. By intersecting the results of the three algorithms, four key genes were identified and designated as key oxidative stress-related diagnostic genes (OSRDGs), biomarkers for predicting the risk of thyroid eye disease.
[0082] 7. Statistical analysis
[0083] All statistical analyses were performed in R software (version 4.4.0). Continuous variables were compared using the t-test if they followed a normal distribution; if not, the Wilcoxon rank-sum test was used. A two-sided p-value of less than 0.05 was considered statistically significant.
[0084] (2) Experimental process and results
[0085] 1. Identification of differentially expressed genes in TED
[0086] Differentially expressed genes were analyzed using the previously obtained sequencing datasets (GSE285190 and GSE280114) and clinical characteristics (including 152 TED patients and 61 HC samples). The results showed that, under the conditions of setting the threshold of |log2FC|>0 and p-value<0.05, a total of 2290 up-regulated genes and 2972 down-regulated genes were identified in the TED group compared with the HC group ( Figure 2KEGG enrichment analysis showed that the thyroid hormone signaling pathway was significantly upregulated in TED patients, as well as classic inflammation-related pathways such as the MAPK signaling pathway, PI3K-Akt signaling pathway, T cell receptor signaling pathway, and JAK-STAT signaling pathway ( Figure 2 In addition, GSEA analysis showed that pathways related to OS and inflammatory response were significantly enriched in the TED group.
[0087] Considering that TED is an ocular complication of GD, in order to reduce the impact of GD on the identified differentially expressed genes and to locate the gene regulation pattern specific to TED, the present invention further analyzed the differentially expressed genes between TED and GD data based on the peripheral blood bulk RNA sequencing data of 20 GD patients obtained above. The results showed that compared with GD, the TED group had a total of 3997 genes upregulated and 3669 genes downregulated ( Figure 2 Similarly, the TED group was enriched in thyroid hormone signaling pathways, as well as inflammation-related pathways such as MAPK signaling pathways and TNF signaling pathways ( Figure 2 GSEA analysis showed that the respiratory chain electron transport, PI3K-Akt, and MAPK signaling pathways were enriched in TED patients.
[0088] We further intersected the differentially expressed genes of the TED-HC group and the differentially expressed genes of the TED-GD group, and finally obtained 2633 genes highly correlated with TED ( Figure 2 GO enrichment analysis of these 2633 genes showed that they were significantly enriched in biological processes such as "response to oxidative stress", "response to oxygen levels" and "endogenous apoptosis signaling pathway related to oxidative stress" ( Figure 2 These results suggest that the occurrence of TED is closely related to the inflammatory process associated with oxidative stress.
[0089] 2. Identification of key module genes related to TED and OS
[0090] To screen candidate genes for constructing a TED prediction model, we also selected 38 gene sets related to the pathogenesis of TED from the MSigDB database and performed ssGSEA analysis. The results showed that compared with GD and HC, OS-related signals in TED patients were significantly changed ( Figure 3 A).
[0091] WGCNA was further used to identify the gene modules most associated with each gene set. A soft threshold of 4 was selected and the top 5,000 genes with the highest MAD were clustered to obtain 6 modules. Among them, the blue module was highly correlated with multiple OS-related gene sets (such as "response to oxidative stress" and "reactive oxygen species pathway") ( Figure 3 Further analysis showed that there was a high positive correlation between blue module members and OS phenotype (r = 0.74, p < 1e⁻²). 00 )( Figure 3 In addition, the expression levels of genes in the blue module were significantly different between TED and HC and GD. KEGG enrichment analysis showed that genes in the blue module were significantly enriched in multiple inflammation-related pathways, including neutrophil extracellular trap formation, T cell receptor signaling pathway, leukocyte transendothelial migration, Th17 cell differentiation, and NF-kappa B signaling pathway ( Figure 3 D), suggesting that this module is closely related to the TED autoimmune response.
[0092] 3. Integrate DEGs, WGCNA results and OSRGs to obtain candidate markers
[0093] Since OS is believed to be involved in the pathogenesis of TED, we extracted 784 OS-related genes (OSRGs) from the MSigDB database and intersected them with the 2633 TED-related DEGs screened previously and the 1574 genes detected in the blue module. We ultimately identified 14 candidate OS-related biomarkers that are closely related to the pathogenesis of TED, including: P4HB, SELENON, KLF2, NR4A2, PLEKHA1, CTSD, PHC2, ARHGDIA, GNB2, LYN, ACTB, QSOX1、TPP1 and TXNRD1 , and included them in subsequent analyses ( Figure 3 E).
[0094] 3. Multiple machine learning algorithms identified OS-related biomarkers with predictive value for TED disease risk
[0095] To further identify OS-related biomarkers for constructing a TED disease risk prediction model, this study developed an OS-related TED diagnostic gene signature set using a comprehensive machine learning algorithm. All TED and HC samples were randomly divided into training and test sets at a ratio of 7:3. Three independent machine learning algorithms were then applied to the training set using 14 candidate genes as input for variable screening (see Table 1 for baseline characteristics of the training and test sets). In the LASSO regression model, the selected λ value suggested retaining 13 genes for subsequent analysis ( Figure 4 The regression coefficients of related genes are shown in Figure 4 B). The RF algorithm screened the six key genes most closely related to the onset of TED based on gene importance ( Figure 4 In the univariate logistic regression model, 8 genes significantly associated with TED were screened out ( Figure 4Finally, we took the intersection of the genes obtained by the three models and identified four key genes with TED prediction potential: KLF2, TXNRD1, CTSD and PHC2 ( Figure 4 In addition, the present invention found that the correlation between these four genes was low, suggesting that the model had a small collinearity effect, which helps to improve the interpretability and stability of the prediction model ( Figure 7 A).
[0096] Table 1 Comparison of demographic characteristics between the training set and the test set
[0097]
[0098] NA means not in the above age range.
[0099] Example 2: Establishing and evaluating a prediction model for predicting TED disease risk
[0100] 1. Model construction and verification
[0101] After obtaining the corresponding regression coefficients based on the standardized gene expression levels and multivariate logistic regression analysis ( Figure 7 B), the present invention calculates an OSRDG score (OSRDG Score, OSRDGS) for each sample, and the calculation formula is:
[0102] OSRDGS = 0.995878 × exp( KLF2 ) + 0.634031 × exp( TXNRD1 ) + 0.640694 × exp( CTSD ) + 0.767982 × exp( PHC2 ).
[0103] Among them, exp( KLF2 ) indicates that KLF2 The exponential function with expression level as exponent and e as base; exp( TXNRD1 )、exp( CTSD )、exp( PHC2 ) and so on.
[0104] OSRDGS was validated in the test set. Univariate logistic regression analysis showed that higher OSRDGS significantly increased the risk of TED (OR: 2.718, 95% CI: 1.608–4.596, p<0.001) ( Figure 5After adjusting for clinical variables such as age and gender, OSRDGS was still statistically significant in multivariate logistic regression analysis (OR: 3, 95% CI: 1.765–5.471, p<0.001), suggesting that OSRDGS is an independent predictor of TED ( Figure 5 B).
[0105] In addition, the predictive performance of OSRDGS as a TED disease risk prediction model was further evaluated by receiver operating characteristic (ROC) curve, decision curve analysis (DCA), and calibration curve. The results showed that the ROC curve showed that OSRDGS had high diagnostic accuracy in the training set and test set, with AUCs of 0.704 and 0.748, respectively. Figure 5 To further verify its predictive performance, 20 TED patient samples and 20 GD patient samples were randomly selected and combined to construct a new TED-GD validation set. The ROC curve of this set also supports that OSRDGS has good predictive performance for TED risk, with an AUC value of 0.703 ( Figure 5 The calibration curves of the training and test sets showed that there was almost no deviation between the predicted probability and the actual TED incidence, further validating the good predictive performance of OSRDGS ( Figure 7 In addition, DCA showed that OSRDGS-based prediction could bring higher net benefits to TED patients compared with “all intervention” or “no intervention” ( Figure 7 The calibration curves in the TED-GD set and the DCA results also support OSRDGS as a reliable prediction tool ( Figure 7 E, H).
[0106] 2. Clinical Correlation and Immune Infiltration Analysis
[0107] The present invention further explores the correlation between the clinical manifestations and peripheral blood immune infiltration characteristics in OSRDGS and TED patients.
[0108] To explore the clinical predictive value of OSRDGS in TED patients, TED samples were divided into high OSRDGS group and low OSRDGS group according to the median OSRDGS. The four markers ( KLF2、TXNRD1 、 CTSD 、 PHC2 ) and compared the OSRDGS values of each sample in different clinical variable groups. The results were visualized by box plots. The results showed that the expression of the four markers in the high OSRDGS group was significantly increased, verifying the internal consistency of the scoring system ( Figure 5Further analysis of the relationship between OSRDGS levels and clinical features of TED showed that OSRDGS levels were significantly higher in “extremely severe” patients than in “mild” patients ( Figure 5 G), and OSRDGS showed an increasing trend from "mild" to "moderately severe" and then to "very severe" patients ( Figure 5 This suggests that OSRDGS has a certain predictive value for disease progression and prognosis. In addition, OSRDGS is also significantly increased in patients with abnormal Schirmer test results ( Figure 5 The Schirmer test is a clinical test that assesses tear secretion function. Abnormalities in this test often indicate a predisposition to dry eye in patients with TED. These results suggest that the OSRDGS is a highly predictive model for TED risk, closely associated with a higher risk of the disease and poorer clinical outcomes.
[0109] To explore the immune infiltration characteristics between high and low groups, ssGSEA and CIBERSORT algorithms were used to identify immune cell scores or proportions, and the immune cell enrichment scores obtained from ssGSEA analysis were displayed using heat maps and bidirectional bar charts respectively. KLF2、TXNRD1 、 CTSD 、 PHC2 ) and OSRDGS. In addition, GO enrichment analysis was performed on the high and low groups, including three dimensions: BP, cellular component (CC) and molecular function (MF), and the significantly enriched pathways were displayed by dot plots. The results showed that the scores of various immune cells in the high OSRDGS group were significantly increased. Among them, innate immune cells such as monocytes / macrophages, dendritic cells (DCs), natural killer cells (NK cells), mast cells and neutrophils showed significant infiltration in the high OSRDGS group, suggesting the activation of innate immune response in the early stage of TED ( Figure 6 In addition, some lymphocyte subtypes were also increased in the high OSRDGS group, including central memory CD4 and CD8 T cells, effector memory CD8 T cells, follicular helper T cells (Tfh), and type I helper T cells (type I Thelper, Th1). Most of the infiltrating immune cells were associated with four markers ( KLF2、TXNRD1 、 CTSD 、 PHC2 ) expression levels were significantly correlated. CIBERSORT analysis further depicted the proportional distribution of 22 immune cells in each sample ( Figure 6B), the proportions of monocytes, neutrophils, and resting mast cells were higher in the high OSRDGS group. In the ssGSEA results, macrophages, monocytes, and DCs were also the immune cell types most correlated with OSRDGS levels ( Figure 6 C).
[0110] GO functional enrichment analysis was further performed to elucidate the differences in biological functions between the high and low OSRDGS groups, thereby helping to explain the potential pathogenesis of TED. The results showed that in the high OSRDGS group, in addition to multiple OS-related functional items (such as "response to oxidative stress", "reactive oxygen metabolism process", and "cellular response to oxidative stress"), immune-related biological functions including "myeloid leukocyte activation", "interleukin 6 (IL-6) production", "immune response activation signaling pathway", "phagocytic vesicles", and "immune receptor activity" were also enriched. Figure 6 These results strongly support the initiation of innate immune responses in the pathogenesis of TED.
[0111] In summary, based on peripheral blood bulk RNA sequencing data from three populations: TED patients, healthy controls, and patients with simple GD, the present invention integrated differentially expressed genes, weighted gene co-expression network analysis results, and OS-related gene sets, and further applied multiple machine learning algorithms to ultimately screen out four markers with predictive value for TED disease risk. Based on this, an OS-related diagnostic gene score (OSRDGS) was constructed as a predictive model for the risk of thyroid eye disease. Experimental evaluation showed that the model had a high predictive value for the risk of TED disease, and the model score level was positively correlated with disease progression, indicating important potential for clinical translation.
[0112] Although the present invention has been described in detail through the above preferred embodiments, it should be understood that the above description is not intended to limit the present invention. After reading the above description, various modifications and substitutions of the present invention will become apparent to those skilled in the art. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. A set of biomarkers for predicting the risk of thyroid eye disease, characterized in that: The biomarkers are KLF2, TXNRD1, CTSD and PHC2 composition.
2. Use of the biomarker according to claim 1 in constructing a prediction model for the risk of thyroid eye disease.
3. A method for constructing a prediction model for the risk of thyroid eye disease, characterized in that: The prediction model is OSRDGS = 0.995878 × exp ( KLF2 ) + 0.634031 × exp( TXNRD1 ) + 0.640694 × exp( CTSD ) + 0.767982 × exp( PHC2 ); Among them, exp( KLF2 )、exp( TXNRD1 )、exp( CTSD )、exp( PHC2 ) respectively represent KLF2、TXNRD1 、 CTSD 、 PHC2 The expression level is the exponential function with e as the base, and OSRDGS represents the risk score for thyroid eye disease.
4. The method for constructing a prediction model for the risk of thyroid eye disease according to claim 3, wherein: When the OSRDGS of the subject's sample is high, it is predicted that the subject has a high risk of developing thyroid eye disease; when the OSRDGS of the subject's sample is low, it is predicted that the subject has a low risk of developing thyroid eye disease.
5. The method for constructing a prediction model for the risk of thyroid eye disease according to claim 3, wherein: Detection of samples by quantitative PCR KLF2, TXNRD1, CTSD and PHC2 expression level.
6. A detection kit for predicting the risk of thyroid eye disease, characterized in that: The kit comprises: KLF2, TXNRD1, CTSD and PHC2 Expression level reagents.
7. The detection kit according to claim 6, wherein The sample is a blood sample.
8. The detection kit according to claim 6, wherein When the sample of subjects KLF2, TXNRD1, CTSD and PHC2 When the expression level is low, the risk of the subject developing thyroid eye disease is low; when the subject's sample KLF2、TXNRD1、 CTSD and PHC2 When the expression level is high, the subject is evaluated to have a high risk of developing thyroid eye disease.
9. The detection kit according to claim 6, wherein The kit detects the KLF2, TXNRD1, CTSD and PHC2 expression level.
10. The detection kit according to claim 9, wherein The kit includes: fluorescent probe / dye; DNA polymerase; dNTPs; specific amplification KLF2, TXNRD1, CTSD and PHC2 primer sets and buffer.
Citation Information
Patent Citations
Homocysteine as Graves eye disease biomarker and application thereof
CN120084997A
KR20230081950A