Biomarker of type 2 diabetes mellitus and application thereof
Through single-cell RNA sequencing and machine learning to identify metal ion transport-related genes, biomarkers and therapeutic targets of type 2 diabetes were developed, solving the problem of distinguishing and managing type 2 diabetes in the prior art, and achieving high accuracy diagnosis and personalized treatment.
Patent Information
- Application Number
- CN202510061163.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology is difficult to effectively distinguish and manage type 2 diabetes, and the lack of in-depth research on the functions of metal ion transport-related genes in disease pathogenesis and immune cell infiltration, resulting in large differences in treatment choices and difficulty in maintaining optimal blood sugar control.
Single-cell RNA sequencing (scRNA-seq) technology is used and combined with machine learning algorithms to identify and verify genes such as ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE as biomarkers and therapeutic targets for type 2 diabetes, and develop diagnostic reagents and therapeutic drugs.
Achieve high accuracy in distinguishing between type 2 diabetes and non-diabetics, providing new diagnostic and treatment options, improving the effectiveness of blood sugar control and the possibility of personalized treatment.
Smart Images

Figure CN119979694A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of diagnosis and treatment of type 2 diabetes, and in particular to biomarkers of type 2 diabetes and applications thereof. Background Art
[0002] According to the International Diabetes Federation (IDF), approximately 537 million adults worldwide are affected by diabetes in 2021, with type 2 diabetes (T2D) accounting for approximately 90% of these cases. Characterized by insulin resistance, impaired insulin secretion, and relative insulin deficiency, the pathophysiology of T2D is influenced by genetic, environmental, and lifestyle factors, which can lead to hyperglycemia and serious complications if not properly managed. However, despite standard care measures including lifestyle changes, oral hypoglycemic agents, and insulin therapy, many patients still have difficulty maintaining optimal glycemic control, indicating an urgent need for more effective treatments. The heterogeneity of T2D has led to gaps in personalized treatment options, highlighting the need to explore new strategies to enhance metabolic control and improve patient outcomes.
[0003] Single-cell RNA sequencing (scRNA-seq) has revolutionized our ability to analyze gene expression at the cellular level, providing detailed insights into the cellular diversity and molecular changes that occur in T2D. Pancreatic islets are composed of various cell types including β-, α-, and δ-cells, which play a key role in maintaining glucose homeostasis. Disruption of these cellular functions is a key factor in the development of T2D. Understanding changes in gene expression and cell-cell interactions within these islets could reveal key pathways and regulatory networks that drive T2D.
[0004] Increasingly, research is focusing on the role of immune cell infiltration within pancreatic islets in T2D. Immune cells, such as macrophages and T cells, infiltrate pancreatic islets, creating a chronic inflammatory environment that exacerbates beta-cell dysfunction. The interaction between these immune cells and islet cells plays a key role in disease progression and may significantly affect the effectiveness of T2D management.
[0005] The role of metal ions, such as zinc, iron, and copper, in cellular function has become a key area of research in T2D. These metal ions are essential for the production, secretion, and activity of insulin, and disturbances in their balance are associated with metabolic disorders, including diabetes. However, the function and role of genes related to metal ion transport (RMITRGs) in the pathogenesis of T2D and immune cell infiltration are still lacking in-depth research and confirmation.
[0006] In short, in-depth research on the pathogenesis of T2D and its related genes is of great value and significance for a more comprehensive understanding of the pathogenesis of T2D and the discovery of new therapeutic intervention targets. Summary of the invention
[0007] The purpose of this application is to provide a new biomarker for type 2 diabetes and its application.
[0008] This application adopts the following technical solutions:
[0009] The first aspect of the present application discloses the use of at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE as a biomarker or therapeutic target for type 2 diabetes.
[0010] In one implementation of the present application, the biomarker or therapeutic target is a combination of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE.
[0011] It should be noted that the present application study found that ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE have expression differences in type 2 diabetes and non-diabetes; therefore, in principle, these genes can be used individually as biomarkers to distinguish type 2 diabetes from non-diabetes; however, in one implementation of the present application, these 13 genes are combined and specific algorithms and models are used to more accurately and effectively distinguish type 2 diabetes from non-diabetes.
[0012] The second aspect of the present application discloses the use of at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE in the preparation of a diagnostic reagent for type 2 diabetes.
[0013] It should be noted that the key to this application is that the research discovered and confirmed a group of genes (RMITRGs) related to metal ion transport, namely ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, which can be used as T2D biomarkers and therapeutic targets. In one implementation of this application, 12 machine learning algorithms were used in 108 different configuration combinations, and the AUC index in the training and validation cohorts was used to evaluate the predictive performance of these 13 genes for T2D. The results showed that the combination of Stepglm[backward] and GBM algorithms produced the best results, achieving an AUC of 0.999 in the training cohort and an AUC of 0.921 in the validation cohort, with an average AUC of 0.96, indicating that the biomarkers of this application can stably and accurately distinguish T2D from non-diabetes (ND). Moreover, through structural prediction of their encoded proteins, the results showed that the pTM scores of most proteins were higher than 0.5 and could serve as potential therapeutic targets.
[0014] It should also be noted that the present application uses one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, especially 13 combinations, as biomarkers for type 2 diabetes, and the nucleic acids or proteins encoded by the corresponding nucleic acids, for example, according to the nucleic acid expression level, or the nucleic acid encoding protein expression level, diagnose type 2 diabetes. Among them, the nucleic acid expression level can be determined by detecting DNA, RNA or mRNA; the nucleic acid encoding protein expression level can be characterized by detecting the level of mutual binding between the nucleic acid encoding protein and its specific polypeptide. In the present application, the analysis object can be a collected sample, for example, by detecting the nucleic acid expression level or the nucleic acid encoding protein expression level in the collected sample, diagnosing type 2 diabetes; or, directly analyzing the nucleic acid expression level or the nucleic acid encoding protein expression level according to the existing data, so as to diagnose type 2 diabetes. Therefore, the use of at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE in the preparation of type 2 diabetes diagnostic reagents mainly refers to designing specific detection reagents for these nucleic acids or nucleic acid-encoded proteins.
[0015] The third aspect of the present application discloses a reagent for diagnosing type 2 diabetes, which is a nucleic acid detection reagent and / or a protein detection reagent for detecting one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE.
[0016] It should be noted that the key to this application is that the research found that one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE can be used as biomarkers for type 2 diabetes; therefore, reagents for detecting these nucleic acids and / or reagents for detecting proteins encoded by these nucleic acids can be used as reagents for diagnosing T2D.
[0017] In one implementation of the present application, the nucleic acid detection reagent includes specific primers and / or probes designed for one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, and the protein detection reagent includes a polypeptide that specifically binds to or recognizes a protein encoded by one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE.
[0018] It should be noted that the specific primers of the present application refer to primers designed for one or more combinations of nucleic acids among ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, such as primers for gene sequencing, primers for polymerase chain reaction, primers for isothermal amplification reaction, or other primers capable of detecting nucleic acids; specific probes refer to probes designed for these nucleic acids, such as real-time fluorescence quantitative PCR probes, gene chip detection probes, specific probes for probe hybridization methods, or other probes capable of detecting nucleic acids.
[0019] In one implementation of the present application, the nucleic acid includes one or more genes selected from ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, a DNA fragment extracted from the gene, or RNA, mRNA, miRNA or LncRNA corresponding to the gene or its DNA fragment; the polypeptide includes one or more combinations of antibodies, antigen-binding fragments, single-chain antibodies and ligands that specifically bind to the encoded protein.
[0020] It should be noted that the key to this application is to detect the expression level of one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, and the detection object can be a nucleic acid, including various DNA or RNA, or a nucleic acid-encoded protein. When the detection object is a nucleic acid-encoded protein, it can be detected by various polypeptides that specifically bind to it, such as including but not limited to antibodies, antigen-binding fragments, single-chain antibodies, ligands that specifically bind to nucleic acid-encoded proteins, etc.
[0021] The fourth aspect of the present application discloses a kit for diagnosing type 2 diabetes, which contains the reagent for diagnosing type 2 diabetes of the present application.
[0022] In one implementation of the present application, the kit of the present application further contains a real-time fluorescent PCR reaction mixture.
[0023] It should be noted that the key to the kit of the present application lies in the specific primers and / or probes designed for nucleic acids of one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, or polypeptides that specifically bind to or recognize nucleic acid-encoded proteins; for ease of use, the reagents for detecting nucleic acids may also include conventional reagents or components corresponding to the specific detection method used, such as gene sequencing, polymerase chain reaction, isothermal amplification reaction, gene chip, probe hybridization, gel electrophoresis, Northern blotting, nucleic acid mass spectrometry or liquid chromatography. Conventional reagents or components; similarly, reagents for detecting nucleic acid-encoded proteins may also include amino acid sequencing, enzyme-linked immunosorbent assay, chemiluminescence, immunochromatography, radioimmunoassay, immunohistochemistry, immunoblotting, flow cytometry, gel electrophoresis, protein spectrometry or liquid chromatography. Conventional reagents or components. It can be understood that these conventional reagents or components corresponding to the specific detection method adopted can be purchased commercially or integrated into the kit of the present application; therefore, for the specific detection scheme of the nucleic acid-specific primer pair and the reference gene-specific primer pair of the present application, the kit of the present application also contains a real-time fluorescence PCR reaction mixture.
[0024] The fifth aspect of the present application discloses the use of an agent targeting one or a combination of more of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE in the preparation of a drug for the treatment of type 2 diabetes.
[0025] The sixth aspect of the present application discloses a drug for treating type 2 diabetes, which uses one or a combination of more of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE as therapeutic targets.
[0026] It should be noted that the key to this application is that studies have found that ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE can be used as therapeutic targets for type 2 diabetes; therefore, reagents targeting these genes can be used as drugs for the treatment of type 2 diabetes.
[0027] The seventh aspect of the present application discloses a method for evaluating or screening therapeutic drugs for type 2 diabetes, comprising detecting the nucleic acid level or protein level of one or more combinations of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE in a sample treated with the therapeutic drug.
[0028] In one implementation of the present application, the nucleic acid level refers to the expression level of DNA, RNA or mRNA, and the protein level refers to the level of binding between the nucleic acid-encoded protein serving as a marker and its specific polypeptide.
[0029] It should be noted that the present application found that there were differences in the expression of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE between T2D and ND, and therefore used them as biomarkers for T2D; on this basis, it can be understood that if the treatment of T2D patients is successful, the expression levels of these biomarkers will also be different from before the treatment. Therefore, the present application creatively proposes to use the expression levels of these biomarkers to evaluate or screen therapeutic drugs for type 2 diabetes.
[0030] The beneficial effects of this application are:
[0031] This application provides a set of new biomarkers and therapeutic targets for type 2 diabetes. By analyzing the expression levels of these biomarkers, it is possible to effectively distinguish type 2 diabetes patients from non-diabetic patients, providing a new solution and approach for the diagnosis and treatment of type 2 diabetes, which has important clinical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figures 1 to 5 is a visualization analysis diagram of gene expression in pancreatic islet cell types in the examples of the present application, wherein: Figure 1 T-SNE plot showing clustering of single-cell RNA sequencing data of 118,262 cells from 17 patients with type 2 diabetes (T2D) and 27 non-diabetic (ND) individuals, identifying seven major cell types within pancreatic islets; Figure 2 To perform subtype analysis using T-SNE, cell subpopulations were further distinguished; Figure 3 To annotate cell clusters based on known marker genes and confirm cell identities; Figure 4 Comparative analysis of cell type ratios between T2D and ND samples, shown by cell type (left) and disease state (right); Figure 5 The heat map highlights genes with significantly altered expression in different cell types in T2D patients.
[0033] Figures 6 to 8 This is a graph showing the results of using machine learning to build and verify a prediction model in an embodiment of the present application, wherein: Figure 6 1953 differentially expressed genes (DEGs) between T2D and ND samples were identified; 398 metal ion transport-related genes (RMITRGs) were identified from the GOBP_REGULATION_OF_METAL_ION_TRANSPORT gene set, of which 49 RMITRGs intersected with DEGs; Figure 7 In the performance evaluation of 108 machine learning model combinations, the Stepglm[backward]+GBM model achieved the highest accuracy in both the training (AUC=0.999) and validation (AUC=0.921) cohorts; Figure 8 Using AlphaFold 3 to predict the structure of the hub RMITRG protein, 12 of the 13 proteins were successfully predicted, except for AHNAK due to its large amino acid sequence length;
[0034] Figures 9 to 11 This is a graph showing the correlation analysis results between hub RMITRGs and immune cell types in the embodiment of this application, wherein: Fig. 9 The heat map shows the relationship between hub RMITRGs and immune cells in the training dataset; Fig.10 The corresponding heat map of the test dataset; Fig.11 Detailed correlation analysis of selected RMITRGs (AHNAK, ATF4, B2M, CYBA, GNB2, HES1) with specific immune cell types;
[0035] Fig.12 1 is a graph showing the correlation analysis results of ACTN4, PRNP, TMBIM6, TSPAN13 and VMP1 with neutrophils, mast cells, eosinophils and T helper cells, respectively, in the examples of the present application;
[0036] Fig.13 is a graph showing the results of the protein-protein interaction network analysis of the hub RMITRGs in the examples of the present application, using a PPI network constructed with 50 genes closely related to 13 hub RMITRGs, showing the interactions between these genes, it is noteworthy that GNB2, TMBIM6, ACTN4, YWHAE, CYBA, ATF4, PRNP, ATP1A1 and B2M show interconnected interactions, while TSPAN13, VMP1 and HES1 appear independent, the top ten hub genes were determined by five different algorithms: MCC, MNC, EPC, radial and closeness;
[0037] Figures 14 to 17 is a disease association analysis result diagram of hub RMITRGs under various disease conditions in the embodiment of the present application, Figures 14 to 17 The associations between the 13 hub RMITRGs and T2D, diabetes, glucose metabolism disorders, and metabolic diseases are sequentially displayed for CTD analysis, and the results are presented in bar graphs reflecting the inferred scores. DETAILED DESCRIPTION
[0038] Type 2 diabetes (T2D) is a complex metabolic disorder with a significant impact on global health. Understanding the molecular mechanisms behind T2D is crucial for developing effective treatment strategies. This application uses single-cell RNA sequencing (scRNA-seq) and machine learning methods to deeply study the pathogenesis of T2D, with a special focus on the infiltration of immune cells.
[0039] This application analyzed scRNA-seq data of pancreatic islet cells from T2D patients and non-diabetic (ND) patients to identify differentially expressed genes (DEGs), especially genes related to metal ion transport (RMITRGs). This application applied 12 machine learning algorithms to develop predictive models and evaluated the infiltration of immune cells using single sample gene set enrichment analysis (ssGSEA). The correlation between immune cells and key RMITRGs was studied, and the interactions between these genes were explored by protein-protein interaction (PPI) network analysis to reveal their potential role in T2D. In addition, this application also used machine learning to develop predictive models, explore the relationship between hub metal ion transport-related genes (RMITRGs) and various metabolic conditions, and evaluate immune cell infiltration within pancreatic islet tissue.
[0040] This application identified 1953 DEGs between T2D and ND, of which the Stepglm[backward] and GBM model combination showed high prediction accuracy and identified 13 hub RMITRGs. Protein structure was predicted using AlphaFold 3, revealing potential functional conformations. This application observed a strong correlation between hub RMITRGs and immune cells, and PPI network analysis revealed key interactions.
[0041] This application combines single-cell transcriptomics with computational modeling, immune infiltration analysis, and network analysis to provide a comprehensive framework for revealing the molecular basis of T2D. This application analysis identified genes that are critical for T2D, emphasizing and confirming ion transport, signal transduction, and immune cell interactions. These studies and discoveries provide valuable insights into disease mechanisms and provide a scientific basis for more effective diagnostic and therapeutic targets.
[0042] The present application is further described in detail below through specific examples. The following examples are only used to further illustrate the present application and should not be construed as limiting the present application.
[0043] Example
[0044] 1. Methods
[0045] 1. Single-cell RNA sequencing (ScRNA-seq) and data analysis
[0046] We collected single-cell RNA sequencing (ScRNA-seq) data of pancreatic islet cells from 17 patients with type 2 diabetes (T2D) and 27 non-diabetic (ND) individuals from the PANC-DB database. ScRNA-seq data were processed using the Seurat package (version 4.4.0) in R. To ensure data quality, we screened cells using the following thresholds: mitochondrial content less than 15%, cell count greater than 500, and gene count between 1,000 and 25,000. Highly variable genes were identified using default parameters, and the data were scaled to a maximum of 10. The data matrix was normalized (1,000 transcripts per cell), logitized, and scaled per gene. Significant dimensions were selected based on P values, and principal component (PC) analysis was performed. Graph-based clustering was performed using significant PCs. Batch effect correction was performed using the ‘RunHarmony’ function. For visualization, T-distributed stochastic neighbor embedding (T-SNE) was used to cluster the data and visualize the major and sub-cell types within the islets. Cell clusters were annotated using known marker genes, and the proportions of cell types in different disease states were calculated.
[0047] 2. Identification of genes related to metal ion transport
[0048] The Wilcoxon rank sum test was used to compare the expression differences between T2D and ND samples in different cell types to identify differentially expressed genes (DEGs), and an adjusted p-value < 0.05 was used as the significance threshold. 398 genes related to metal ion transport (RMITRGs) were identified from the GOBP_REGULATION_OF_METAL_ION_TRANSPORT gene set using the Molecular Signature Database (MSigDB) (https: / / www.gsea-msigdb.org / gsea / msigdb). These RMITRGs were intersected with the identified DEGs to obtain hub RMITRGs for further analysis.
[0049] 3. Batch RNA Data Processing
[0050] T2D datasets GSE54279 and GSE41762 were downloaded from the GEO database using the GEOquery package in R language. The dataset GSE54279 was generated using the GPL6244 platform and includes samples from 128 T2D patients. The dataset GSE41762 was also generated using the GPL6244 platform and includes 77 samples, including 20 control samples and 57 T2D samples. The microarray data of these datasets were processed using the robust multi-array average (RMA) method for background correction, normalization, and probe adjustment. The batch effect was corrected using the Combat method.
[0051] 4. Model prediction
[0052] To develop a prediction model for T2D, we explored 108 different combinations of 12 machine learning algorithms. These algorithms included LASSO, Ridge, Enet, Stepglm, SVM, GlmBoost, LDA, plsRglm, GBMs, XGBoost, naiveBayes, and RF models. These combinations were evaluated using the area under the curve (AUC) metric in the training and validation cohorts. The model construction in this case used expression data of 49 RMITRGs identified in ScRNA-seq analysis. 70% of the samples from the combined cohorts GSE54279 and GSE41762 were used for model training, and the remaining 30% were used for validation. Model performance was determined based on the AUC score, and the best performing model was selected accordingly.
[0053] 5. Protein structure prediction
[0054] To examine the structural features of hub proteins associated with T2D, we employed AlphaFold 3, an advanced protein structure prediction tool. A set of hub RMITRGs associated with T2D were selected for analysis, including ACTN4, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1, and YWHAE, except AHNAK, whose sequence length exceeded the prediction capacity. AlphaFold 3 was configured with standard parameters to ensure accurate predictions. The primary amino acid sequences of the selected proteins were submitted to AlphaFold 3, and multiple iterations were run for each protein to ensure reliable results. Confidence scores, including pLDDT and pTM, were calculated to assess the quality of the predicted structures. A pTM score above 0.5 indicates similarity to the true folded structure, while a score above 0.8 indicates a high-quality prediction.
[0055] 6. Immune cell infiltration analysis in T2D and ND groups
[0056] The infiltration fractions of immune cell types in T2D and ND groups were calculated using single sample gene set enrichment analysis (ssGSEA) in the “gsva” R package using previously published immune cell markers, as detailed in reference: Bindea G, Mlecnik B, Tosolini M, Kirilovsky A, Waldner M, Obenauf AC, Angell H, Fredriksen T, Lafontaine L, Berger A, Bruneval P, Fridman WH, Becker C, Pagès F, Speicher MR, Trajanoski Z, Galon J. Spatiotemporal dynamics of intratumoral immune cells reveal the immune landscape in human cancer. Immunity. 2013 Oct 17; 39(4): 782-95. doi:10.1016 / j.immuni.2013.10.003. PMID:24138885. Heatmap visualization was performed using the “pheatmap” package. The correlation between immune cells and related functions was assessed using Spearman correlation, and the results were visualized using the “corrplot” R package and “ggplot2.” R values greater than 0.3 were considered to be strongly positive correlations.
[0057] 7. Protein-protein interaction (PPI) network construction
[0058] To elucidate the interactions between hub RMITRGs and other related genes, we constructed a PPI network. This network was constructed using 50 genes that were closely related to the 13 hub RMITRGs we identified, which play important roles in metal ion transport and are differentially expressed in T2D and ND patients. These 50 genes were identified using the STRING database (https: / / string-db.org / ). The interactions within this network were visualized using Cytoscape (version 3.8.2). The cytoHubba plugin in Cytoscape used five algorithms—MCC, MNC, EPC, radial, and closeness—to rank the top 10 nodes in the PPI network, each of which provided a unique analytical perspective. The 50 genes used in this example include: ATP1B4, BAD, BEX3, BHLHA9, CACNA1B, CACNA1D, CREB3L2, CREB3L4, CTNNA1, CTNNA2, CTNNA3, DENND1A, DESI1, ERN1, FANCF, FXYD1, FXYD2, IAPP, KIAA0930, MAPK1, MAPK3, MARK2, MLF1, MR1, NOX3, NOX5, NOXA1, NUTM2A, NUTM2B, PRND, TP53INP2, TRARG1, CAPN12, TRIM3, MICALL2, MAGI1, PLCB3, RPS6KA6, KCNJ5, TAPBPL, KIR2DL3, IGHV3-43D, TRDV1, ENSP00000497443, POLDIP2, NOXO1, PDCL, GNG14, HTR7 and SPRN.
[0059] 8. Disease association by comparing toxicogenomics databases
[0060] The Comparative Toxicogenomic Database (CTD) was used to explore the interactions between the 13 hub RMITRGs and various diseases, including diabetes, T2D, glucose metabolism disorders, and metabolic diseases. The inferred scores of these interactions were calculated and displayed as bar graphs.
[0061] 9. Statistical Analysis
[0062] All analyses were performed in R (version 4.2.1). P values less than 0.05 were considered statistically significant.
[0063] 2. Results
[0064] 1. Overview of the research workflow
[0065] This study started with single-cell RNA sequencing (scRNA-seq) analysis of pancreatic islet cells from type 2 diabetes (T2D) patients and non-diabetic (ND) individuals, revealing different differentially expressed genes (DEGs) and cell type distributions. Figures 1 to 5 Subsequent steps involved developing machine learning models based on these DEGs, leading to the identification of key hub genes and prediction of their protein structures, as shown in Figures 6 to 8 Single-sample gene set enrichment analysis (ssGSEA) was then used to calculate the infiltration scores of 22 immune cell types in the T2D and ND groups and to examine the correlation between these immune cells and the identified hub genes, as shown in Figures 9 to 11 A protein-protein interaction (PPI) network was constructed to explore the interactions between these hub genes, as shown in Fig.13 As shown in the figure, the data of Comparative Toxicogenomic Database (CTD) were used to analyze their association with various metabolic diseases including T2D. Figures 14 to 17 shown.
[0066] 2. Gene Expression Mapping of Islet Cell Types
[0067] Building on our comprehensive approach to understanding T2D, we used scRNA-seq archives of 118,262 cells from 17 T2D patients and 27 ND individuals to map gene expression in various islet cell types, such as Figures 1 to 5 T-SNE analysis of single-cell data identified seven major cell types in the islet samples, providing a comprehensive view of the cell type composition, as shown in Figure 2. Figure 1 The cell types identified included endocrine cells, astrocytes, endothelial cells, mast cells, ductal cells, acinar cells, and macrophages. Further T-SNE analysis provided a more detailed resolution of cell identity and heterogeneity within the islets, as shown in the following figure. Figure 2 Clusters were annotated based on known marker genes to ensure accurate identification of cell types, including β cells, α cells, δ cells, and other pancreatic islet cell types. Figure 3 Analysis of cell proportions by disease state highlighted differences in cellular composition between T2D and ND samples. Figure 4 The heat map shows the genes that are significantly upregulated or downregulated in T2D patients in different cell types compared with healthy controls. Figure 5 shown.
[0068] 3. Identification of DEGs and their functional roles
[0069] Furthermore, we analyzed the gene expression profiles of pancreatic islet cell types in detail and delved into the molecular features that distinguish T2D and ND states. DEGs were identified, which are crucial for understanding the pathophysiology of T2D. A total of 1953 DEGs were identified between β cells of T2D and ND samples, such as Figure 6 As shown. Among them, 398 genes related to metal ion transport (RMITRGs) were identified from the GOBP_REGULATION_OF_METAL_ION_TRANSPORT gene set. The intersection of DEGs and these RMITRGs produced 49 RMITRGs. It is crucial to study these genes because metal ions play key roles in various cellular processes, including enzyme activity, signal transduction, and maintaining cellular homeostasis. Disruption of metal ion transport may affect insulin secretion and sensitivity and may contribute to the pathogenesis of T2D. By identifying and analyzing these RMITRGs, we can gain a deeper understanding of the molecular mechanisms behind T2D and discover possible therapeutic targets to improve disease management.
[0070] 4. Development and validation of machine learning models
[0071] Having established the differential gene expression profile and the importance of RMITRGs in T2D, we continued to take these findings in a new direction - using machine learning to predict T2D with greater accuracy. This approach allowed us to harness the power of computational algorithms to parse the complex patterns in the expression data of the 49 RMITRGs we identified through scRNA-seq analysis.
[0072] To predict T2D, we explored 12 machine learning algorithms in 108 different configurations and evaluated their performance using the AUC metric in the training and validation cohorts. Expression data of 49 RMITRGs identified in scRNA-seq analysis were used to build the models. The combination of Stepglm[backward] and GBM algorithms produced the best results, achieving an AUC of 0.999 in the training cohort and 0.921 in the validation cohort, with an average AUC of 0.96, as shown in Figure 2. Figure 7 This high level of accuracy highlights the robustness of the model in distinguishing T2D from ND status. The Stepglm[backward]+GBM model incorporated 13 genes, including ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1, and YWHAE, which can serve as T2D biomarkers.
[0073] The protein structures of 12 of the 13 hub RMITRGs were successfully predicted using AlphaFold 3, except for AHNAK, whose amino acid sequence length exceeded the prediction capability of AlphaFold 3, e.g. Figure 8 As shown. The structural predictions provided valuable insights into the potential functional conformations of these proteins. The overall predicted folds of most proteins, including YWHAE, GNB2, B2M, ATP1A1, TSPAN13, VMP1, ACTN4, CYBA, TMBIM6, and PRNP, were considered reliable, with pTM scores above 0.5, indicating a high accuracy of their predicted structures. However, the pTM scores of HES1 and ATF4 were below 0.5, indicating that the predictions were less accurate and require further experimental validation. These results highlight the potential of combining machine learning and structural biology approaches to identify and validate key molecular players in T2D. The identified hub genes and their predicted structures provide a basis for future functional studies and therapeutic targeting, with implications for improving T2D management.
[0074] 5. Correlation between hub RMITRGs and immune cells
[0075] Building on the above studies, we shifted our focus to the complex relationship between hub RMITRGs and immune cell infiltration, a key aspect of disease pathogenesis. This shift allowed us to bridge the gap between gene expression and immunological implications, providing a more comprehensive view of the complex dynamics of T2D.
[0076] To further investigate the immune cell infiltration and function between T2D patients and ND controls, ssGSEA was used to evaluate the enrichment scores of various immune cell subsets and functions. Fig. 9 ) and the test set ( Fig.10 ) analyzed the correlation between 13 hub RMITRGs and immune cells, revealing strong positive correlations between multiple genes and immune cells, including ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1, and YWHAE. In particular, AHNAK, ATF4, B2M, CYBA, GNB2, and HES1 showed strong positive correlations with at least two immune cell types, such as Fig.11As shown. For example, AHNAK is strongly associated with Th1 cells and neutrophils; ATF4 is strongly associated with CD8 T cells and cytotoxic cells; B2M is strongly associated with macrophages and T helper cells; CYBA is strongly associated with macrophages and NK CD56dim cells; GNB2 is strongly associated with cytotoxic cells and Th17 cells; HES1 is strongly associated with T helper cells, NK CD56dim cells, neutrophils and macrophages. In addition, ACTN4, PRNP, TMBIM6, TSPAN13 and VMP1 are positively correlated with neutrophils, mast cells, eosinophils, mast cells and T helper cells, respectively, as shown Fig.12 Only ATP1A1 and YWHAE were not associated with immune cells.
[0077] These correlations provide valuable insights into the immunological aspects of T2D and suggest potential avenues for therapeutic intervention. They also highlight the need for further investigation of the precise mechanisms by which these hub RMITRGs influence immune cell behavior and promote disease progression.
[0078] 6. PPI Network Analysis of Hub RMITRGs
[0079] Building on our understanding of the relevance of hub RMITRGs to immune cells, we extended our analysis to investigate the complex interaction networks that these genes may participate in in the context of T2D cells. We therefore performed PPI network analysis, which is critical to deciphering how these genes may collaboratively or independently influence disease outcomes.
[0080] The PPI network constructed for 50 genes closely associated with 13 hub RMITRGs provided insights into the molecular interactions and regulatory mechanisms involved in T2D. In this network, GNB2, TMBIM6, ATP1A1, ACTN4, YWHAE, CYBA, ATF4, PRNP, and B2M showed significant mutual interactions, suggesting their central roles and potential collaborative functions in disease progression.
[0081] In contrast, TSPAN13, VMP1, and HES1 did not interact with other hub RMITRGs, suggesting that they may function independently or in different molecular pathways, such as Fig.13As shown. To identify the most critical nodes in the PPI network, five analysis algorithms were used—MCC, MNC, EPC, radial, and closeness. These algorithms helped identify hub genes based on different aspects of network topology, such as connectivity and centrality. The top ten hub genes identified by these algorithms intersected, revealing seven genes that appeared consistently across all methods: MAPK1, MAPK3, ATP1A1, ATF4, ATP1B4, FXYD2, and CYBA. This consensus highlights the importance of these genes in the network and their potential as key regulators of T2D.
[0082] These findings highlight the complexity of T2D at the molecular level and the potential for targeted therapies to disrupt disease progression by modulating these key interaction networks. PPI network analysis not only enhances our understanding of T2D pathophysiology, but also helps to unravel the detailed mechanisms of how these hub genes work in the future.
[0083] 7. Associations between hub RMITRGs and disease conditions
[0084] After we explored the PPI network and identified key genes with significant interactions, we went on to investigate the broader impact of these hub RMITRGs in various disease conditions. This step is critical to understanding the scope of their impact beyond T2D and identifying their potential roles in other metabolic disorders.
[0085] Analysis using the Comparative Toxicogenomics Database (CTD) highlighted relationships between the 13 hub RMITRGs and various disease conditions. Inference scores were calculated for type 2 diabetes, diabetes, glucose metabolism disorders, and metabolic diseases, and significant associations are shown in bar graphs. Results for type 2 diabetes are shown in Fig.14 As shown, the results of diabetes mellitus are Fig.15 As shown, the consequences of glucose metabolism disorders are as follows Fig.16 As shown, the results of metabolic diseases are Fig.17 These results provide valuable insights into the involvement of these genes in various metabolic diseases, highlighting their importance in T2D pathology.
[0086] Specifically, five genes—ATF4, ATP1A1, B2M, CYBA, and PRNP—emerged with the highest inferred scores in the context of T2D and diabetes, suggesting that these genes play key roles in the molecular mechanisms of these conditions. For example, beta-2-microglobulin (B2M), as a component of MHC class I molecules, is essential for immune responses. Elevated levels of B2M are associated with inflammation and metabolic disorders, suggesting its role in inflammatory processes that promote insulin resistance and beta cell dysfunction in T2D.
[0087] In addition, the prion protein (PRNP), traditionally associated with prion diseases, also plays a role in cellular processes such as signal transduction, cell adhesion and resistance to oxidative stress. Its association with T2D suggests a broader function in maintaining cellular health under conditions of metabolic stress. Extension of the analysis to glucose metabolism disorders and a wider range of metabolic diseases reinforces the importance of these hub genes in a variety of metabolic contexts. The consistent identification in multiple disease conditions emphasizes their central role in metabolic regulation and their potential as therapeutic targets.
[0088] Other hub RMITRGs, such as HES1, TMBIM6, TSPAN13, and VMP1, also showed significant associations with metabolic diseases, such as Fig.17 As shown, multiple molecular pathways involved in T2D are highlighted. For example, HES1 is involved in regulating developmental processes and cell differentiation, and its association with metabolic diseases suggests a potential role in regulating pancreatic β-cell function and insulin secretion. TMBIM6 is known for its ability to inhibit apoptosis and regulate calcium homeostasis, which may contribute to its protective effects on β-cells and its impact on cellular stress responses in T2D. TSPAN13 and VMP1 are involved in processes such as cell adhesion, signal transduction, and autophagy, which are critical for cell maintenance and stress responses, further implicating their roles in the pathogenesis of T2D.
[0089] These findings highlight the multifaceted nature of hub RMITRGs and their potential impact in a range of metabolic conditions, providing a foundation for studies to elucidate their specific roles and develop targeted therapeutic strategies.
[0090] Discussion
[0091] Unraveling the complexities of type 2 diabetes requires a deep understanding of the molecular mechanisms behind it. In this study, we took a comprehensive approach to unravel these complexities by integrating single-cell RNA sequencing (scRNA-seq), machine learning, and protein-protein interaction (PPI) network analysis. Our findings reveal a complex gene expression landscape and key regulatory pathways in pancreatic islet cells from patients with T2D.
[0092] ScRNA-seq analysis provided important insights, revealing that the gene expression profiles of pancreatic islet cells in T2D patients were significantly altered compared with non-diabetic (ND) controls. T-SNE clustering exposed different cell populations and highlighted changes in the proportions of major cell types, especially β cells, which are critical for insulin secretion. Differential gene expression analysis identified 1,953 differentially expressed genes (DEGs), providing a list of genes widely implicated in the pathogenesis of T2D. These results highlight the critical role of β cell dysfunction in T2D and suggest new molecular targets for therapeutic intervention.
[0093] By integrating machine learning algorithms with scRNA-seq data, we developed a predictive model for T2D. Among the 108 combinations of 12 machine learning algorithms tested, the Stepglm[backward] and GBM model combination achieved the highest predictive accuracy. The model identified 13 key metal ion transport-related genes (RMITRGs), 12 of which had their protein structures successfully predicted using AlphaFold 3. AHNAK could not be modeled due to its large sequence length, reflecting the limitations of current computational tools in handling very large proteins. These 13 identified hub genes could serve as biomarkers for T2D diagnosis and as therapeutic targets for new therapies.
[0094] Construction of the PPI network revealed key interactions between hub RMITRGs, highlighting key regulatory nodes such as ATP1A1, GNB2, TMBIM6, and ACTN4. Top hub genes identified by multiple analysis algorithms underscored their central role in the network. Further analysis using the Comparative Toxicogenomics Database (CTD) highlighted strong associations between these hub genes and various metabolic disorders including T2D, glucose metabolism disorders, and general metabolic diseases. This emphasizes the broad relevance of these genes and their potential as therapeutic targets.
[0095] Our study provides insights into the association between hub RMITRGs and immune cells, revealing important interactions that provide insights into the pathogenesis of T2D. The integration of these analyses provides a comprehensive understanding of the potential role of RMITRGs in regulating immune cell behavior and their infiltration in pancreatic islet cells.
[0096] 1. Biological basis: The role of metal ions in immune cell function is well known, and ions such as zinc and iron are essential for cell signaling, redox balance, and inflammation regulation. Given the central role of these ions, changes in RMITRGs expression may affect the local microenvironment and affect the infiltration and activity of immune cells within the islets.
[0097] 2. Correlation analysis: Our findings showed a strong correlation between hub RMITRGs and immune cells, suggesting that differential expression of these genes may regulate inflammatory responses in T2D. This correlation highlights the potential of RMITRGs as regulatory nodes in the inflammatory process mediated by immune cells within the pancreatic islets.
[0098] 3. Mechanistic relationship: The mechanistic relationship between RMITRG expression and immune cell infiltration is multifaceted: (1) Altered metal ion homeostasis: Altered RMITRG expression may disrupt metal ion homeostasis, affect redox balance, and regulate inflammatory responses within the islets. (2) Inflammatory signaling pathways: Differential expression of RMITRGs may activate inflammatory signaling pathways, such as NF-κB, which is critical for immune cell activation and cytokine production. (3) Impact on T2D pathogenesis: Immune cell infiltration affected by RMITRG expression plays a key role in β-cell function and T2D progression. Our findings suggest that altered RMITRG expression may directly affect β-cell function by regulating the local islet microenvironment and influencing cytotoxic molecules released by immune cells. This direct effect on β-cells may lead to dysfunction and apoptosis, which are key hallmarks of T2D progression. Immune cell infiltration regulated by RMITRG expression not only disrupts β-cell function, but also contributes to the chronic inflammatory state within the islets, further exacerbating insulin resistance and β-cell failure.
[0099] Understanding the complex relationship between RMITRG expression, immune cell infiltration, and beta cell function is critical for developing targeted therapies designed to preserve beta cell function and reduce inflammation in T2D. By targeting the molecular pathways that link RMITRGs to immune cell activity, we may be able to mitigate the deleterious effects of inflammation on beta cells and slow the progression of T2D. This understanding also highlights the potential for therapeutic interventions that stabilize or even reverse the inflammatory processes that lead to beta cell dysfunction.
[0100] Our study revealed unique expression patterns of DEGs between T2D and non-diabetic patients, revealing molecular signatures that distinguish T2D at the cellular level. Identification of these DEGs, particularly those involved in metal ion transport, provided key insights into pathways dysregulated in the pathogenesis of T2D. Metal ion transport genes, such as those encoding zinc and iron transporters, play critical roles in insulin production, secretion, and activity. Our findings suggest that perturbations in the normal function of these transporters may contribute to impaired insulin secretion and action, hallmarks of T2D. Dysregulation of these genes affects not only insulin signaling but also redox balance and inflammatory responses within islet cells, exacerbating β-cell dysfunction and insulin resistance.
[0101] The biological significance of these DEGs extends to the development of personalized treatment strategies for T2D. Identifying reliable biomarkers among these DEGs could aid early diagnosis and patient stratification, enabling more targeted and effective therapeutic interventions. For example, understanding the specific role of metal ion transporters in β-cell function could lead to the development of therapies aimed at normalizing their expression or function, thereby improving insulin secretion and overall glycemic control.
[0102] In our study exploring the molecular mechanisms of T2D, we used an integrated scRNA-seq and machine learning approach to reveal the role of RMITRGs. Beyond traditional bulk RNA sequencing or single gene studies, this approach allowed us to identify a set of hub RMITRGs that were previously unrelated to T2D. These genes were integrated into prediction models and showed high accuracy in distinguishing T2D from non-diabetic states, thus they can serve as biomarkers and therapeutic targets. Using AlphaFold 3, we predicted the protein structures of these hub RMITRGs, providing new insights into their functional roles and laying the foundation for future studies.
[0103] Our findings also revealed a correlation between hub RMITRGs and immune cell infiltration, providing a new perspective on the immune landscape of T2D. Furthermore, the construction of the PPI network provides a systems-level understanding of molecular interactions in T2D, enhancing our knowledge of the disease pathophysiology. These insights provide new approaches for targeted and effective therapeutic strategies.
[0104] Overall, our study has established a robust framework for understanding the genetic and immune basis of T2D. By continuing to investigate these hub genes and their pathways, researchers can advance personalized therapies and optimize clinical outcomes for T2D management. Although the prediction accuracy of HES1 and ATF4 is low, and AHNAK cannot be structurally predicted, the above studies have shown that ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1, and YWHAE, these 13 hub RMITRGs can serve as biomarkers for T2D diagnosis and can serve as therapeutic targets. These studies and findings provide a new solution and approach for the diagnosis and treatment of T2D, which has important clinical application value.
[0105] The above contents are further detailed descriptions of the present application in combination with specific implementation methods, and it cannot be determined that the specific implementation of the present application is limited to these descriptions. For ordinary technicians in the technical field to which the present application belongs, several simple deductions or substitutions can be made without departing from the concept of the present application.
Claims
1. Use of at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE as a biomarker or therapeutic target for type 2 diabetes.
2. The use according to claim 1, characterized in that: The biomarkers or therapeutic targets are a combination of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE.
3. Use of at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE in the preparation of diagnostic reagents for type 2 diabetes.
4. A reagent for diagnosing type 2 diabetes, characterized in that: The reagent is a nucleic acid detection reagent for detecting at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE and / or a protein detection reagent for detecting a protein encoded by the nucleic acid.
5. The reagent according to claim 4, characterized in that: The reagent is a nucleic acid detection reagent for detecting a combination of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE and / or a protein detection reagent for a protein encoded by the nucleic acid; Preferably, the nucleic acid detection reagent comprises specific primers and / or probes designed for ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, and the protein detection reagent comprises a polypeptide that specifically binds to or recognizes proteins encoded by ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE; Preferably, the nucleic acid comprises genes of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE, DNA fragments extracted from the genes, RNA, mRNA, miRNA or LncRNA corresponding to the genes or their DNA fragments; The polypeptide includes at least one of an antibody, an antigen binding fragment, a single-chain antibody, and a ligand that specifically binds to the encoded protein.
6. A kit for diagnosing type 2 diabetes, characterized in that: Containing the reagent according to claim 4 or 5.
7. The kit according to claim 6, characterized in that: It also contains real-time fluorescence PCR reaction mixture.
8. Use of an agent targeting at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE in the preparation of a drug for treating type 2 diabetes.
9. A medicine for treating type 2 diabetes, characterized in that: The drug uses at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE as a therapeutic target.
10. A method for evaluating or screening a drug for treating type 2 diabetes, characterized in that: The method comprises detecting the nucleic acid level or protein level of at least one of ACTN4, AHNAK, ATF4, ATP1A1, B2M, CYBA, GNB2, HES1, PRNP, TMBIM6, TSPAN13, VMP1 and YWHAE in a sample treated with a therapeutic drug.
Citation Information
Cited By
Marker combination for diagnosis or prognosis evaluation of diabetes and application
CN122017253A