Biomarker composition for diabetic foot ulcer and application of biomarker composition

By screening DDIT4 and PKM genes as biomarkers and combining transcriptome and single-cell sequencing data, diagnostic kits and predictive models were developed, solving the diagnostic and treatment challenges of diabetic foot ulcers and improving diagnostic accuracy and treatment efficacy.

CN120866518APending Publication Date: 2025-10-31JIAXING NO 1 HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511179724.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies are insufficient for the effective diagnosis and treatment of diabetic foot ulcers. The lack of reliable biomarkers and targets leads to poor treatment outcomes and may result in serious consequences.

Method used

By integrating transcriptome sequencing data and single-cell sequencing data, DDIT4 and PKM genes were screened as biomarkers, and diagnostic kits were developed. The expression levels of these genes were detected using specific primers, probes, or antibodies. A predictive model was constructed by combining PPI networks and Lasso regression models to improve diagnostic accuracy.

Benefits of technology

It provides reliable molecular targets and a predictive model with high diagnostic accuracy, reveals the abnormal energy metabolism and immune microenvironment regulation mechanism of diabetic foot ulcers, provides a theoretical basis for treatment, and improves the scientific value of early diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120866518A_ABST
    Figure CN120866518A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomarkers, in particular to a biomarker composition for diabetic foot ulcer and application of the biomarker composition. The invention provides a biomarker combination for diabetic foot ulcer. The biomarker combination comprises a DDIT4 gene and a PKM gene. According to the invention, by integrating transcriptome sequencing data and single cell sequencing data, biomarkers DDIT4 and PKM closely related to diabetic foot ulcer (DFU) are successfully screened out. The expression of the two genes in a disease group is remarkably up-regulated, and the two genes are stably expressed in different data sets, so that a reliable molecular target is provided for early diagnosis of DFU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomarker technology, and in particular to a combination of biomarkers for diabetic foot ulcers and their applications. Background Technology

[0002] Diabetic foot ulcers (DFU) are a global health problem, estimated to affect approximately 18.6 million people worldwide each year. These ulcers not only impair physical function and significantly reduce quality of life, but also significantly increase the burden on healthcare resources. If left untreated, foot ulcers can worsen, developing into soft tissue infections, gangrene, and even potentially leading to amputation. Notably, about half of diabetic foot ulcer patients also have peripheral artery disease in the lower extremities. Furthermore, approximately 50% of ulcer cases become infected, with up to 20% of infected patients requiring hospitalization. Even more seriously, 15% to 20% of patients with moderate to severe infections may ultimately face amputation, with a 5-year mortality rate as high as 30%.

[0003] Transcriptome sequencing is a genomics research method based on high-throughput sequencing technology, designed to obtain the sequence and expression information of almost all transcripts in a specific cell or tissue under a given state. Single-cell sequencing (scRNA-seq) is a technique for high-throughput sequencing and analysis of the transcriptome at the single-cell level. It can obtain intracellular gene expression information by detecting single-cell sequences to further study their related functions.

[0004] In summary, this invention explores biomarkers related to DFU based on transcriptome sequencing data and single-cell analysis from public databases, understands the molecular characteristics of these biomarkers and the pathophysiological mechanisms mediating DFU, and uses scRNA-seq technology to investigate the expression distribution of these biomarkers among different cells, providing a reference for the subsequent diagnosis and treatment of DFU. Summary of the Invention

[0005] The purpose of this invention is to provide a combination of biomarkers for diabetic foot ulcers and their applications.

[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution:

[0007] This invention provides a combination of biomarkers for diabetic foot ulcers, characterized in that the combination of biomarkers includes the DDIT4 gene and the PKM gene.

[0008] Preferably, the DDIT4 and PKM genes are significantly upregulated in the tissues of patients with diabetic foot ulcers.

[0009] The present invention also provides the application of the aforementioned biomarker combination in the preparation of diagnostic reagents for diabetic foot ulcers.

[0010] Preferably, the reagent is a kit.

[0011] The present invention also provides a diagnostic kit for diabetic foot ulcers, the kit comprising reagents for detecting the expression levels of the combination of said biomarkers.

[0012] Preferably, the reagent includes specific primers, probes, or antibodies for quantitative detection of the mRNA or protein expression levels of the DDIT4 and PKM genes.

[0013] Compared with existing technologies, this invention has the following beneficial effects: First, by integrating transcriptome sequencing data and single-cell sequencing data, this invention successfully screened out biomarkers DDIT4 and PKM closely related to diabetic foot ulcer (DFU). These two genes showed significantly upregulated expression in the disease group and remained stable across different datasets, providing reliable molecular targets for the early diagnosis of DFU. Second, enrichment analysis revealed that the candidate genes were significantly enriched in key pathways such as glycolysis and pyruvate metabolism, revealing the important role of abnormal energy metabolism in the pathogenesis of DFU. Furthermore, the constructed PPI network and Lasso regression model further validated the reliability of the biomarkers, while nomogram and ROC curve analyses showed that the prediction model based on DDIT4 and PKM had high diagnostic accuracy (AUC = 1), providing a powerful tool for clinical practice.

[0014] Furthermore, this invention delves into the association between biomarkers and immune cell infiltration, discovering significant correlations between DDIT4 and PKM and various immune cells (such as activated B cells and effector memory CD4 T cells), providing a new perspective for understanding the regulatory mechanisms of the immune microenvironment in DFU. Through GSEA enrichment analysis and GeneMANIA database prediction, the biological functions of these genes, such as glucose catabolism and NADH regeneration, were further clarified, laying a theoretical foundation for developing targeted therapeutic strategies. In summary, this invention not only provides new biomarkers and potential targets for the diagnosis and treatment of DFU but also reveals its molecular mechanisms through multi-omics analysis, possessing significant scientific value and promising clinical applications. Attached Figure Description

[0015] Figure 1 This is a volcano plot of differentially expressed genes; the vertical axis is -log10, and the horizontal axis represents the fold change log2FC.

[0016] Figure 2A heatmap of differentially expressed genes was created; the top 10 upregulated genes and the top 10 downregulated genes were plotted based on log2FC fold change.

[0017] Figure 3 The distribution of nFeature_RNA, nCount_RNA, and percent.mt before quality control is shown.

[0018] Figure 4 This is a distribution chart of nFeature_RNA, nCount_RNA, and percent.mt after quality control.

[0019] Figure 5 For screening of highly variable genes.

[0020] Figure 6 The images show a linear dimension plot (top) and a scree plot (bottom). In the top plot, the horizontal axis represents the principal components (PCs), ordered from highest to lowest according to their explanatory power of the dataset variance. Each point or band on the horizontal axis corresponds to a specific principal component. The vertical axis represents a statistical significance or importance indicator, expressed as a p-value. This measures the statistical significance of each principal component in capturing the true structure of the dataset.

[0021] In the graph below, the horizontal axis represents the principal component numbers, sorted from highest to lowest according to their explanatory power for the dataset variance. Each point or marker on the horizontal axis corresponds to a specific principal component. The vertical axis represents the proportion of dataset variance explained by each principal component.

[0022] Figure 7 This is a UMAP clustering diagram; the horizontal axis represents the first dimension of the data after UMAP dimensionality reduction in the low-dimensional space. The vertical axis represents the second dimension of the data after UMAP dimensionality reduction in the low-dimensional space. The positions of different cells on the horizontal and vertical axes reflect their relative positions in the reduced-dimensional space.

[0023] Figure 8 This is a PCA plot of the sample; the horizontal axis represents the first principal component (PC1), which has the largest contribution to variance among all principal components, reflecting the projection of cells onto the first principal component; the vertical axis represents the second principal component (PC2), which has the second largest contribution to variance after PC1, reflecting the projection of cells onto the second principal component.

[0024] Figure 9 For Marker Gene expression; Note: The horizontal axis represents cell type, the vertical axis represents Marker genes, the size of the bubbles represents the percentage of Marker gene expression in the cell, and the closer the bubble color is to orange, the greater the expression level.

[0025] Figure 10 This is the annotated UMAP clustering diagram.

[0026] Figure 11 The image shows a UMAP visualization after annotation. The visualization is based on the proportion of 12 cell types in the DFU group and the control group. As shown in the figure, smooth muscle cells (SMCs) account for the largest proportion in both the control group and the DFU group.

[0027] Figure 12 Visualize the proportion of cells between groups; Note: The horizontal axis represents the proportion of cells, and the vertical axis represents different cell types.

[0028] Figure 13 Box plot for differential cell identification; the top plot shows the expression percentage of 12 cell types; the bottom plot shows the expression percentage of B lymphocytes and lymphoendothelial cells; note: the horizontal axis represents cell type, and the vertical axis represents the expression percentage of cells.

[0029] Figure 14 Manhattan plot for differential gene identification; Note: Horizontal axis represents cell type, and vertical axis represents log2FC.

[0030] Figure 15 Candidate genes were identified for the Venn diagram.

[0031] Figure 16 The graph shows the GO enrichment results; the horizontal axis represents the number of genes located in the pathway, and the vertical axis represents the GO entry name.

[0032] Figure 17 This is a GO enrichment result image. The concentric circles from the outside to the inside represent: the first layer, the ID of the GO function, showing the three functions: BP, CC, and MF; the second layer, the color intensity represents the significance level, and the length, width, and number represent the number of genes enriched in this function; the third layer represents the number of genes upregulated or downregulated in this function, and the color is used to distinguish between gene upregulation and downregulation; the innermost layer of color blocks represents different functions, and the size represents the Rich Factor of this pathway. The larger the Rich Factor, the greater the degree of enrichment.

[0033] Figure 18 The image shows the KEGG enrichment results; the horizontal axis represents the number of genes located in the pathway, and the vertical axis represents the pathway name.

[0034] Figure 19This is a KEGG enrichment result diagram. The layers from the outside to the inside represent: the first layer, the ID of the KEGG function; the second layer, the color intensity represents the significance, and the length, width, and number represent the number of genes enriched in that function; the third layer represents the number of genes upregulated or downregulated in that function, and the color is used to distinguish between gene upregulation and downregulation; the innermost layer's color blocks represent different functions, and the size represents the Rich Factor of that pathway (the Rich Factor is the ratio of the number of differentially expressed transcripts located at that KEGG entry to the total number of transcripts located at that KEGG entry in all annotated transcripts; the larger the Rich Factor, the greater the degree of enrichment).

[0035] Figure 20 This is a PPI network; Note: Each node represents a gene, and the lines represent interactions at the protein level.

[0036] Figure 21 This is a Lasso regression analysis; the top graph shows the Lasso regression coefficient path: the horizontal axis is the logarithm of lambdas, and the vertical axis is the variable coefficients (regression coefficients). As lambdas increases, the variable coefficients converge and approach 0. When the optimal lambda is reached, variables with coefficients equal to 0 are removed. The bottom graph shows the Lasso cross-validation curve: the horizontal axis is the logarithm of lambdas, and the vertical axis is the model error. As can be seen from the graph above, the optimal lambda value is at the point of lowest error, indicated by the dashed line on the left.

[0037] Figure 22 Box plots for expression levels; Note: The top plot shows the expression levels of candidate key genes in the training set; the bottom plot shows the expression levels of candidate key genes in the validation set; In the violin plots, **** represents p<0.0001, *** represents p<0.001, ** represents p<0.01, * represents p<0.05, and ns represents p>0.05.

[0038] Figure 23 This is a nomogram; note: each variable's line segment is labeled with a scale representing the range of that variable, and the length of the line segment reflects the contribution of that factor to the outcome event. The individual scores (Points) in the graph represent the individual scores for each variable at different values. The total score (Total Points) represents the total score, summed up after considering all variable values.

[0039] Figure 24 For calibration curves; Note: The horizontal axis represents the prediction effect of the nomogram model, and the vertical axis represents the actual prediction effect.

[0040] Figure 25This is the ROC curve; Note: the closer the AUC is to 1, the more accurate the diagnosis. When 0.7 > AUC > 0.5, the diagnostic accuracy is low; when 0.9 > AUC > 0.7, the diagnostic effect has some reference value; when AUC > 0.9, the diagnostic accuracy is high. When AUC = 0.5, there is no diagnostic value.

[0041] Figure 26 This is a GSEA enrichment analysis of the DDIT4 gene. Note: The above figure can be divided into two parts. Part 1: The top five lines represent the gene enrichment score. The vertical axis represents the corresponding Running Estimate (ES). A peak in the line graph represents the enrichment score of this gene set; genes before the peak are the core genes in this set. The horizontal axis represents each gene in this set, corresponding to the barcode-like vertical lines in Part 2. Part 2: The barcode-like section represents Hits, with each vertical line corresponding to one gene in the set. Part 3: This is a distribution of the rank values ​​for all genes. The vertical axis represents the ranked list metric, i.e., the ranking value of the gene. The values ​​in the lower figure are correlation coefficients, and the bubble size represents the significance p-value.

[0042] Figure 27 This is a GSEA enrichment analysis of PKM genes. Note: The above figure can be divided into two parts. Part 1: The top five lines represent the gene enrichment score. The vertical axis represents the corresponding Running Estimate (ES). A peak in the line graph represents the enrichment score of this gene set; genes before the peak are the core genes in this set. The horizontal axis represents each gene in this set, corresponding to the barcode-like vertical lines in Part 2. Part 2: The barcode-like section represents Hits, with each vertical line corresponding to one gene in this set. The following figure shows the distribution of rank values ​​for all genes. The vertical axis represents the ranked list metric, i.e., the ranking value of the gene. The values ​​in the following figure are correlation coefficients, and the bubble size represents the significance p-value.

[0043] Figure 28 This is a network diagram for analyzing the GeneMANIA database. Note: The large circle in the middle represents biomarkers, and the small circles on the outside represent genes associated with the biomarkers. The network lines on the right, from top to bottom, represent: physical interactions, co-expression, prediction, co-localization, genetic interactions, pathways, and shared protein domains. Detailed Implementation

[0044] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.

[0045] Example 1

[0046] Training set 1 (GSE199939)

[0047] The GSE199939 dataset was downloaded from the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ). Ten foot skin tissue samples from diabetic foot ulcer patients were selected as the DFU patient group, and 11 normal foot skin tissue samples were selected as the control group. Sequencing was performed using high-throughput sequencing, and the annotation platform was GPL24676.

[0048] 2. Validation set (GSE134431)

[0049] The GSE134431 dataset was downloaded from the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ). Thirteen foot skin tissue samples from diabetic foot ulcer patients were selected as the DFU patient group, and eight normal foot skin tissue samples were selected as the control group. Sequencing was performed using high-throughput sequencing, and the annotation platform was GPL18573.

[0050] 3. Single-cell dataset (GSE165816)

[0051] The GSE165816 dataset was downloaded from the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ). Peripheral blood and forearm skin samples were removed from the 54 samples. Five foot skin tissue samples from diabetic foot ulcer patients were selected as the DFU patient group, and 11 normal foot skin tissue samples were selected as the control group. Sequencing method: high-throughput sequencing; annotation platform: GPL24676.

[0052] 4 Glycolysis-associated related genes (GRGs) [PMID:39118022]:

[0053] Three gene sets related to glycolysis were obtained from the Molecular Feature Database (MSigDB) (https: / / www.gsea-msigdb.org / gsea / msigdb): KEGG GLYCOLYSIS GLUCONEOGENESIS, HALLMARK GLYCOLYSIS, and REACTOME GLYCOLYSIS. After merging and deduplicating the genes in the gene sets, a total of 290 glycolysis-related genes were obtained, denoted as GRGs.

[0054] 5. Differential Expression Analysis

[0055] Differential expression analysis aims to identify genes with significantly different expression levels among different samples. Grouping may involve different biological states, such as drug treatment versus control, diseased versus healthy individuals, different tissues, or different developmental stages.

[0056] This invention uses the "DEseq2" package in R to analyze the differential expression of genes in the training set between the DFU disease group and the control group (DFU vs control). The screening criteria were set as: adj.pvalue < 0.05 & |log2FC| > 1. To understand the overall distribution of differentially expressed genes, the R package "ggplot2" was used to create a volcano plot to visualize the differentially expressed genes. The top 5 upregulated genes and top 5 downregulated genes were marked in the volcano plot according to the fold change log2FC. At the same time, the R package "ComplexHeatmap" was used to create an expression heatmap to display the top 10 upregulated genes and top 10 downregulated genes according to the fold change log2FC.

[0057] A total of 5397 differentially expressed genes were obtained through the above threshold screening, denoted as DEGs1 (Differentially Expressed Genes, DEGs). Among them, 1689 genes were upregulated and 3708 genes were downregulated in the disease group. Volcano plots and heatmaps were generated based on the results, such as... Figure 1 and 2 As shown.

[0058] 6 Single-cell analysis

[0059] Single-cell sequencing technology has provided unprecedented opportunities to deepen our understanding of transcriptomics, genomics, proteomics, epigenomics, and metabolomics information of individual cells. In recent years, it has become feasible and popular to use scRNA-seq data to identify immune phenotypes and novel immune cell-related functional biomarkers in the tumor microenvironment.

[0060] To investigate gene expression in DFU at the single-cell level, confirm that the analyzed cells belong to different cell populations, and identify cell types in DFU patient samples, we conducted a series of single-cell analyses.

[0061] 7 Quality Control

[0062] First, gene expression in DFU patient tissue samples was investigated at the single-cell level. The single-cell data were then used to create Seurat objects from the R package "Seurat". For example... Figure 3The image shows a violin plot illustrating the distribution of nFeature_RNA (number of genes detected per cell), nCount_RNA (sum of expression levels of all genes detected per cell), and percent.mt (proportion of mitochondrial genes detected per cell) in the pre-quality control sample. The plot shows that the number of genes detected in the cells is concentrated below 7500, the sum of expression levels of all genes is concentrated below 10000, and mitochondrial genes are concentrated below 20%, totaling 53237 cells. To obtain higher quality cells, we performed quality control on them.

[0063] The quality control criteria were set to select high-quality cells with a gene count of 200–6000 per cell, a total gene expression level of less than 20,000, and a mitochondrial ratio of <15%, ultimately yielding 45,746 cells and 22,590 genes. Figure 4 The image shows a violin plot illustrating the distribution of nFeature_RNA (the number of genes detected in each cell), nCount_RNA (the sum of the expression levels of all genes detected in each cell), and percent.mt (the proportion of mitochondrial genes detected in each cell) in the quality control sample.

[0064] 8UMAP dimensionality reduction

[0065] First, a global scaling normalization method (LogNormalize) is used to normalize the characteristic expression measurement value of each cell according to the total expression, multiply it by a scaling factor (default is 10000), and then perform a logarithmic transformation on the result.

[0066] Secondly, the "FindVariableFeatures" function of the "Seurat" package was used for filtering. The top 2000 highly variable genes were selected from the quality-controlled data as highly variable genes. Figure 5 As shown.

[0067] Then, linear dimensionality reduction is used to obtain the optimal linear dimensionality value for cell clustering, such as... Figure 6 As shown in the figure above, principal components located above the dashed line (p.value = 0.05) are considered important, and their corresponding principal components are statistically more significant, making them more important for dimensionality reduction and subsequent analysis. Figure 6 As shown in the figure below, find the "elbow" point, which is where the curve begins to flatten. The principal components before this point are usually considered to contribute the most to the variation of the dataset, so they should be retained in subsequent analysis.

[0068] Combining the two images, the optimal latitude value of 30 was selected for subsequent cell group division. The dataset was further divided into 27 cell subclusters ranging from 0 to 26, and the UMAP map was plotted as follows. Figure 7As shown, a PCA plot is drawn between samples to determine whether there are clear boundaries between samples, such as... Figure 8 As shown, there are no clear boundaries between samples.

[0069] 9-cell annotation

[0070] First, the "FindAll Markers" function was used to identify important marker genes for different clusters. Cell types were named according to the expression of marker genes in cells. Marker genes were obtained from the literature [PMID:35013299], and different cell types were identified using these genes. A total of 12 cell types were annotated, such as... Figure 9 As shown, they are respectively:

[0071] 1) Smooth muscle cells (SMCs) (TAGLN, ACTA2)

[0072] 2) Fibroblasts (DCN, CFD)

[0073] 3) Endo (ACKR1) endothelial cells

[0074] 4) T lymphocytes (T-lympho (CD3D, CCL5))

[0075] 5) Macrophages (IL1B, CD163, FCGR3A)

[0076] 6) Kera cells (KRT5, KRT14, KRT1),

[0077] 7) Mast cells (TPSAB1),

[0078] 8) B lymphocytes B-lympho (MS4A1),

[0079] 9) Lymphoendothelial cells (CCL21)

[0080] 10) Melano / Schwann cells (MLANA, CDH19), 11) Sweat gland cells / Eccrine sweat gland cells (DCD)

[0081] 12) Plasma cells (MZB1)

[0082] Next, the cell types are mapped onto the UMAP diagram according to the annotations, showing the naming of each cell group in the UMAP clustering method, such as... Figure 10 and 11 As shown.

[0083] The results were visualized based on the proportion of 12 cell types in the DFU group and the control group, such as... Figure 12 As shown, smooth muscle cells (SMCs) accounted for the largest proportion in both the control group and the DFU group.

[0084] 10 Differential Gene Identification

[0085] To obtain differentially expressed genes between DFU patient tissue samples and normal control samples in the scRNA-seq dataset, based on the proportion of different cells in each sample, the Wilcoxon test in the "rstatix" package was used to screen for cell types with significant differences between the DFU group samples and normal control samples. Box plots of the differential test results were then plotted using the R package "ggplot2" to display the results. Figure 13 As shown, B lymphocytes (B-lympho) and lymphoendothelial cells (LymphEndo) showed significant differences between groups, and were therefore defined as differentially expressed cells.

[0086] Next, in the single-cell sequencing dataset, the FindMarkers function in the "Seurat" R package was used to screen for genes with significant differential expression between differentially expressed cell types in the DFU group and the normal control group. The threshold was set to |avg_log2FC|>0.25 and p.value<0.05. After screening and deduplication based on the above thresholds, a total of 2746 differentially expressed genes were obtained, denoted as DEGs2. The Manhattan plot of differentially expressed genes between differentially expressed cell types was drawn using the "scRNAtoolVis" package, as shown below. Figure 14 As shown.

[0087] Identification of 11 candidate genes

[0088] To obtain candidate genes related to glycolysis in DFU, the intersection of differentially expressed DFU-related genes DEGs1 (obtained at the transcriptome level) and DEGs2 (obtained at the single-cell level) obtained in the above steps, and glycolysis-related genes was performed to obtain genes that are related to both disease-related DFU and glycolysis. Among them, DEGs1 contained 5397 differentially expressed genes, DEGs2 contained 2746, and GRGs contained 290 glycolysis-related genes. Figure 15 As shown, using the "VennDiagram" package to draw a Venn diagram, a total of 14 intersecting genes were obtained, namely: PKM, SLC16A3, TGFBI, GAPDH, HSPA5, NOL3, ME2, DDIT4, MXI1, MIF, P4HA2, STC1, TFF3, and CXCR4. The intersecting genes are recorded as candidate genes.

[0089] 12 Enrichment Analysis

[0090] To investigate the biological functions and signaling pathways involved in the candidate genes, we used the R package “clusterProfiler” to perform GO and KEGG enrichment analyses on 14 candidate genes.

[0091] Gene Ontology (GO) enrichment analysis is an analytical method that uses the GO database to annotate genes functionally, including three parts: biological process (BP), cellular component (CC), and molecular function (MF).

[0092] KEGG pathway analysis uses the KEGG database to annotate pathways for all identified protein sets or screened differentially expressed proteins, and analyzes the most important metabolic and signal transduction pathways in which these proteins or genes are involved.

[0093] 13. GO enrichment analysis

[0094] Based on the screening criterion of p.value < 0.05, a total of 387 results were obtained, including 329 biological processes, 15 cellular components, and 43 molecular functions. The top 5 p-values ​​for each component were selected for visualization, as shown below. Figure 16 and 17 As shown.

[0095] In terms of biological processes (BP): the candidate genes were significantly enriched in GO items such as pyruvate metabolic process, pyridine nucleotide metabolic process, nicotinamide nucleotide metabolic process, pyridine-containing compound metabolic process, and glycolytic process.

[0096] Regarding cellular components (CC): candidate genes were significantly enriched in GO entries such as ficolin-1-rich granule lumen, ficolin-1-rich granule, endoplasmic reticulum chaperone complex, endoplasmic reticulum lumen, and nuclear membrane;

[0097] In terms of molecular function (MF): the candidate genes were significantly enriched in GO items such as NAD binding, electron transfer activity, protease binding, peptidyl-proline dioxygenase activity, and potassium ion binding.

[0098] 15. KEGG enrichment analysis

[0099] Based on the screening criterion of p.value < 0.05, the candidate genes were enriched into 10 KEGG signaling pathways. For example... Figure 18 and Figure 19 As shown, the Top 10 pathways, sorted by p-value and visualized, significantly enriched pathways including Carbon metabolism, Pyruvate metabolism, Glycolysis / Gluconeogenesis, Central carbon metabolism in cancer, Biosynthesis of amino acids, Virion-Human immunodeficiency virus, Phenylalanine metabolism, Protein export, Tyrosine metabolism, and Type II diabetes mellitus.

[0100] 16. PPI Network Construction

[0101] To understand the interactions between candidate genes and their corresponding proteins, 14 candidate genes were input into the STRING database (https: / / string-db.org / ). An interaction score > 0.4 was set, and four discrete proteins were included. Their PPI network was then visualized using Cytoscape, as shown below. Figure 20 As shown, the network includes 10 nodes, and database analysis revealed 14 edges, an average node degree of 2, a local clustering coefficient of 0.553, and an enrichment p-value of 0.000668.

[0102] 17. Lasso Machine Learning Algorithm

[0103] Lasso regression, short for Least Absolute Shrinkage and Selection Operator, is a regression analysis method. Its core idea is to minimize the sum of squared residuals under the constraint that the sum of the absolute values ​​of the regression coefficients is less than a constant. In this process, Lasso regression introduces an L1 regularization term to compress the coefficients of some less important variables to zero, thereby achieving feature selection and model simplification. The penalty function = original penalty function (e.g., mean squared error) + L1 regularization term (i.e., the sum of the absolute values ​​of all regression coefficients multiplied by a regularization parameter). In optimizing this penalty function, Lasso regression tends to select features that have a significant impact on the target variable, while compressing the coefficients of less important features to zero. This results in a more concise model with stronger generalization ability.

[0104] Here, the R package "glmnet" is used to perform Lasso regression analysis on 14 candidate genes using the training set. Figure 21 As shown, when Optimallambda = 0.00048 and log(optimallambda) = -7.634705, there are 5 feature genes whose regression coefficients are not penalized to 0. That is, when these 5 candidate genes are retained, the prediction accuracy for the disease is the highest. They are PKM, NOL3, ME2, DDIT4, and TFF3, which are denoted as candidate key genes.

[0105] 18. Expression level verification analysis

[0106] To clarify the expression of candidate key genes in DFU patient and control samples, the expression level differences of candidate genes between the two groups were analyzed (p<0.05). The expression levels of five candidate key genes (DDIT4, ME2, NOL3, PKM, and TFF3) in the training set were obtained, as follows: Figure 22As shown, genes NOL3 and TFF3 were significantly downregulated in the disease group, while DDIT4, ME2, and PKM were significantly upregulated in the disease group.

[0107] In the validation set, the expression of candidate key genes DDIT4 and PKM was significantly upregulated in the disease group, while the expression of other genes was not significantly different. This means that the expression trends of the two candidate key genes (DDIT4 and PKM) are consistent with those of the corresponding genes in the training set. This indicates that the expression of these two candidate key genes is more stable in different datasets between disease and control samples. Genes with significant differences and consistent trends are recorded as biomarkers, namely DDIT4 and PKM.

[0108] 19. Construction of biomarker nomograms

[0109] A nomogram, also known as a nomogram, integrates multiple predictive indicators and plots them on a plane using scaled line segments to represent the relationships between variables in a predictive model.

[0110] The basic principle is to assign a score to each value level of each influencing factor based on the contribution of each influencing factor to the outcome variable (the magnitude of the regression coefficient), then add up the scores to obtain the total score, and finally calculate the predicted value of the individual's outcome event by using the functional transformation relationship between the total score and the probability of the outcome event.

[0111] In the training set, the R package "rms" is used to construct nomograms using biomarkers (DDIT4, PKM). For example... Figure 23 As shown, each gene corresponds to a single score Points, and the scores of each factor are added together to form a total score Total Points. Then, the probability of DFU patients is predicted based on the total points. The higher the total score, the greater the probability of developing the disease.

[0112] The "rms" package was used to plot calibration curves, and the "pROC" package was used to plot ROC curves. The calibration curves were used to evaluate the predictive power of the nomogram, and the AUC values ​​were used to assess the ability of the biomarker in the nomogram to diagnose DFU. Figure 24 , Figure 25 As shown, the nomogram curve plotted in the calibration curve has a high degree of overlap with the reference line and AUC = 1 (AUC > 0.7), which proves that the nomogram has high diagnostic value for the disease.

[0113] 20. GSEA enrichment analysis

[0114] To further understand the biological functions and pathways involved by the screened biomarkers in DFU, the gene set "c2.cp.kegg.v7.4.symbols.gmt" from the molecular feature database was used as the reference gene set based on all samples in the training set. Spearman correlation analysis was performed on each biomarker and all other genes using the R package "psych" to obtain the correlation coefficients. These coefficients were then sorted, and using the descending sort results, GSEA (Gene Set Enrichment Analysis) pathway enrichment analysis was performed on the biomarkers (DDIT4, PKM) using the R package "clusterProfiler".

[0115] Based on a p.value < 0.05 & |NES| > 1, DDIT4 was significantly enriched in 88 pathways, and PKM was significantly enriched in 94 pathways, including signaling pathways such as KEGG PROTEASOME, KEGG DNA REPLICATION, KEGG OXIDATIVE PHOSPHORYLATION, KEGG CELL CYCLE, KEGG SPLICEOSOME, and KEGG RETINOL METABOLISM. Pathways related to carbohydrates were selected for visualization, and correlation analysis was performed to create bubble charts, such as... Figure 26 and 27 As shown.

[0116] 21. GeneMANIA Database Analysis

[0117] To identify other genes associated with biomarker function, GeneMANIA (http: / / www.genemania.org / ) was used to predict genes associated with biomarker function and the functions they participate in.

[0118] Based on the database, 20 functionally related genes were predicted using two biomarkers (DDIT4 and PKM): DDIT4L, PKLR, SREBF1, ARAF, LDHA, ENO3, LDHB, ENO2, ENO1, JMJD8, TPI1, MPC1, PFKL, FGG, CLYBL, TSC1, PFKP, TSC2, MPC2, and HSD17B10. The connections between these genes represent different interactions, such as... Figure 28 As shown.

[0119] The database predicted 37 functions involving biomarkers, including: glucose catabolic process to pyruvate, glycolytic process through fructose-6-phosphate, glycolytic process through glucose-6-phosphate, NADH regeneration, glucose catabolic process, NADH metabolic process, ATP generation from ADP, ADP metabolic process, nucleoside diphosphate phosphorylation, and NAD metabolic process. (For pathway information, see genemania-functions.txt) In exploring the molecular mechanisms of diabetic foot ulcers, we first employed a differential expression analysis strategy to screen differentially expressed genes between the disease group and the control group from the training set. Following a rigorous screening process, we identified 5397 differentially expressed genes (DEGs1), which may provide key clues to the pathogenesis of diabetic foot ulcers. To further refine the function and role of these differentially expressed genes, we introduced a single-cell dataset. Through quality control, dimensionality reduction, and cell type annotation of the single-cell data, we successfully identified 12 different cell types. Based on this, we further analyzed the differentially expressed genes (DEGs2) in these cells, which reflect specific changes in disease state at the single-cell level.

[0120] Next, we performed an intersection analysis on differentially expressed genes DEGs1, DEGs2, and glycolysis-related genes (GRGs) to identify common candidate genes. This step identified 14 candidate genes that showed significant differential expression between the disease and control groups. Enrichment analysis of these candidate genes revealed that they were significantly enriched in glycolysis-related pathways, including pyruvate metabolism, pyridine nucleotide metabolism, nicotinamide nucleotide metabolism, and glycolysis, which may play important roles in the pathogenesis of diabetes and its complications.

[0121] To gain a deeper understanding of the functions and interactions of these candidate genes, we constructed a protein-protein interaction (PPI) network and further screened it using machine learning algorithms. By introducing an external validation set for expression level validation, we ultimately identified two genes, DDIT4 and PKM, with significant and consistent expression trends across groups as biomarkers. These two genes may play a crucial role in the pathogenesis of diabetic foot ulcers.

[0122] In addition, we performed immune infiltration analysis to investigate the infiltration of immune cells in the training set and its differences between the disease and normal groups. We found significant differences in 17 types of immune cells between groups, including activated B cells, activated CD8 T cells, and effector memory CD4 T cells. These changes in immune cells may be closely related to the pathogenesis and progression of diabetic foot ulcers. Correlation analysis further confirmed the significant correlation between biomarkers and immune cells. To gain a more comprehensive understanding of the regulatory mechanisms of biomarkers, we constructed a molecular regulatory network, demonstrating the miRNAs and lncRNAs corresponding to the biomarkers. These non-coding RNAs may influence the pathogenesis of diabetic foot ulcers by regulating the expression of biomarkers.

[0123] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A biomarker combination for diabetic foot ulcers, characterized in that, The biomarker combination includes the DDIT4 gene and the PKM gene.

2. The biomarker combination according to claim 1, characterized in that, The expression of the DDIT4 and PKM genes was significantly upregulated in the tissues of patients with diabetic foot ulcers.

3. The use of the biomarker combination according to claim 1 or 2 in the preparation of diagnostic reagents for diabetic foot ulcers.

4. The application according to claim 3, characterized in that, The reagent is a kit.

5. A diagnostic kit for diabetic foot ulcers, characterized in that, The kit contains reagents for detecting the expression levels of the combination of biomarkers as described in claim 1 or 2.

6. The diagnostic kit according to claim 5, characterized in that, The reagents include specific primers, probes, or antibodies, used to quantitatively detect the mRNA or protein expression levels of the DDIT4 and PKM genes.

Citation Information

Cited By

  • A marker combination, a prediction model and application for predicting the risk of diabetic foot ulcer

    CN122651909A