Screening method and application of sepsis marker based on gene co-expression network

Through hierarchical clustering and GO functional analysis, the clustering preset parameters of the gene co-expression network were dynamically adjusted to solve the problem of imprecise gene module division and improve the accuracy and precision of sepsis marker screening.

CN120690286APending Publication Date: 2025-09-23THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510596014.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The division of gene modules in the existing technology is not precise, resulting in low accuracy in gene module identification. Existing biomarkers lack the ability to dynamically adjust clustering preset parameters based on GO functional results.

Method used

By obtaining the gene expression data of the experimental group and the control group, preprocessing and screening of differentially expressed genes, hierarchical clustering analysis was performed, combined with GO functional analysis, and dynamic adjustment of clustering preset parameters, including correction of step size and segmentation length, to screen out hub genes.

Benefits of technology

The screening accuracy of gene modules was improved, abnormal genes unrelated to the pathophysiological process of sepsis were avoided, the division of gene modules was optimized, and more refined marker screening was achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690286A_ABST
    Figure CN120690286A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biology, in particular to a gene co-expression network-based sepsis marker screening method and application. The screening method comprises the following steps: screening differential expression genes; performing hierarchical clustering analysis on each differential expression gene to combine each differential expression gene into a plurality of target gene modules; and screening out hub genes from each target gene module. According to the method, gene expression data and clinical feature data are combined, the accuracy of screening sepsis marker genes is improved, possible abnormal genes irrelevant to the pathophysiological process of sepsis are avoided, gene modules are divided more finely through primary clustering analysis and secondary clustering analysis, GO function analysis is combined, and the accuracy of screening the sepsis marker genes is improved. According to an analysis result, selection of a GO term subset is optimized, and a step length set by primary clustering is corrected, so that the fine granularity and accuracy of analysis are improved, division of gene modules is optimized, and more refined treatment and screening of markers of sepsis are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to a screening method for sepsis markers based on a gene co-expression network and its application. Background Art

[0002] Sepsis is a dysregulated host response to infection, leading to organ dysfunction and posing a serious threat to human health. Sepsis biomarkers are measurable physiological, biochemical, immunological, and genetic indicators expressed by pathogenic microorganisms after invasion. They are used to quantitatively assess the state of the host's defense response and pathophysiological processes and can be used for early screening of patients with sepsis in emergency departments. Although significant progress has been made in understanding the pathophysiological mechanisms of sepsis, early identification and clinical treatment of sepsis remain unsatisfactory. Gene co-expression network analysis is a bioinformatics method that analyzes coordinated changes in gene expression to reveal underlying regulatory mechanisms and functional modules. It has shown great potential in biomarker discovery and disease pathogenesis exploration. By constructing gene co-expression networks, key gene modules and markers associated with sepsis can be identified, providing new insights and approaches for early diagnosis and precise treatment of sepsis.

[0003] A Chinese patent document with authorization announcement number CN109872776B discloses a method for screening potential biomarkers for gastric cancer based on weighted gene co-expression network analysis and its application. The technical point is that FERMT2 is the potential biomarker screened out by using weighted gene co-expression network analysis (WGCNA) and KEGG pathway, GO enrichment analysis and other analysis methods; this shows that existing biomarkers lack the ability to dynamically adjust clustering preset parameters based on GO functional results, resulting in low accuracy in the identified gene modules. Summary of the Invention

[0004] To this end, the present invention provides a method for screening sepsis markers based on a gene co-expression network, so as to overcome the problem in the prior art that the division of gene modules is not precise, resulting in low accuracy of the identified gene modules.

[0005] To achieve the above objectives, the present invention provides a method for screening sepsis markers based on gene co-expression network, comprising:

[0006] Obtaining gene expression data of the experimental group and the control group, preprocessing the gene expression data, and screening a number of differentially expressed genes based on the first feature data;

[0007] performing hierarchical cluster analysis on the differentially expressed genes to combine the differentially expressed genes into a number of target gene modules;

[0008] The hierarchical cluster analysis of each differentially expressed gene includes: performing a primary clustering on each differentially expressed gene based on a preset step size to obtain an initial gene cluster, and performing a secondary clustering on the initial gene cluster based on the second feature data to obtain a plurality of target gene clusters, performing GO function analysis on each target gene cluster, and determining whether to adjust the clustering preset parameters according to the analysis results;

[0009] The clustering preset parameters are a preset step size and a preset segmentation length;

[0010] Adjusting the clustering preset parameters includes correcting the preset step length to a corrected step length, and correcting the preset segmentation length to a corrected step length;

[0011] The first feature is gene expression level, and the second feature data is clinical feature data;

[0012] Hub genes were screened from each target gene module.

[0013] Furthermore, GO functional analysis was performed on each target gene cluster, including:

[0014] Obtain the relationship between the GO term subsets corresponding to each target gene cluster, link the target gene cluster with the GO term subset to obtain the corresponding gene cluster chain;

[0015] Obtaining the actual number of gene cluster chains of the gene cluster chain and comparing it with the standard number of gene cluster chains, determining the clustering effect based on the comparison result, and adjusting the clustering preset parameters according to the similarity of the gene cluster chain pairs if the clustering effect does not meet the standard;

[0016] Wherein, the gene cluster chain pair is composed of two of the gene cluster chains.

[0017] Furthermore, determining the clustering effect based on the comparison results includes:

[0018] If the actual number of gene cluster chains is less than the standard number of gene cluster chains, the clustering effect is judged to be unsatisfactory;

[0019] If the actual number of gene cluster chains is greater than or equal to the standard number of gene cluster chains, the clustering effect is determined to be up to standard.

[0020] Furthermore, the clustering preset parameters are adjusted according to the similarity of the gene cluster chain pairs, including:

[0021] Calculating the similarity value between each pair of gene cluster chains, and performing a similarity comparison between the similarity value and a similarity threshold, and determining GO term subset similar chain pairs based on the comparison results;

[0022] Obtain the number of similar chain pairs in the GO term subset, record it as the number of similar chain pairs, calculate the percentage of the number of similar chain pairs to the total number of gene cluster chain pairs, record it as the real-time similarity ratio, and compare the standard similarity ratio with the real-time similarity ratio:

[0023] If the real-time similarity ratio is less than the standard similarity ratio, the preset step size is corrected to the corrected step size;

[0024] If the real-time similarity ratio is greater than or equal to the standard similarity ratio, the preset segmentation length is corrected to the corrected segmentation length;

[0025] The revised step length is the preset step length-1; and the revised segmentation length is 90% of the preset segmentation length.

[0026] Furthermore, based on the comparison results, similar chain pairs of GO term subsets were determined to include:

[0027] Obtaining a first comparison result and a second comparison result;

[0028] When the first comparison result is obtained, the corresponding gene cluster chain pair is not marked;

[0029] When the second comparison result is obtained, the corresponding gene cluster chain pair is marked as a GO term subset similar chain pair;

[0030] If the similarity value is less than the similarity threshold, a first comparison result is obtained; if the similarity value is greater than or equal to the similarity threshold, a second comparison result is obtained.

[0031] Further, the relationship between the GO term subsets corresponding to each target data cluster is obtained, and the target data cluster is linked to the GO term subset, including:

[0032] Calculate the real-time similarity between each target data cluster and the GO term subset, and compare the standard similarity with each real-time similarity.

[0033] If the real-time similarity is less than the standard similarity, the relationship is determined to be irrelevant, and the target data cluster is not linked to the GO term subset;

[0034] If the real-time similarity is greater than the standard similarity, the relationship is determined to be related, and the target data cluster is linked to the GO term subset.

[0035] Furthermore, the GO terms are segmented into a number of GO term subsets using a preset segmentation length.

[0036] Furthermore, the initial clustering of the differentially expressed genes based on the preset step size includes:

[0037] Adjacent differentially expressed genes are combined in sequence with a preset step size to obtain the initial gene cluster.

[0038] Furthermore, the hub genes are CTSB, CTSD, ATP6V0D1, UBE2D1 and ATP6V0C.

[0039] On the other hand, the present invention also provides a diagnostic kit for sepsis marker genes screened out according to the screening method for sepsis markers based on gene co-expression network, comprising:

[0040] The kit is used for detecting sepsis.

[0041] Compared with the prior art, the beneficial effect of the present invention is that, by combining gene expression data with clinical characteristic data, the accuracy of screening sepsis marker genes is improved, and possible abnormal genes unrelated to the pathophysiological process of sepsis are avoided. Through primary clustering and secondary clustering analysis, gene modules are divided more finely, and combined with GO functional analysis, the selection of GO term subsets is optimized and the step size set for the primary clustering is corrected according to the analysis results, avoiding poor clustering effect caused by the number or range of GO terms included in the GO term subset, and avoiding the formation of gene clusters of genes with low correlation due to a large step size, thereby improving the granularity and accuracy of the analysis, optimizing the division of gene modules, and achieving more refined processing and screening of sepsis markers.

[0042] Furthermore, when it is determined that the real-time similarity ratio is greater than or equal to the standard similarity ratio, it means that the GO term subsets linked to multiple gene cluster chains are too similar. This situation is due to inappropriate segmentation length of GO terms, resulting in insufficient diversity of functional annotations. In this case, the preset segmentation length needs to be corrected, that is, the preset segmentation length needs to be reduced to improve the accuracy of functional annotations of gene clusters. When it is determined that the real-time similarity ratio is greater than or equal to the standard similarity ratio, it means that the number of gene cluster chains linked to the same GO term subset is small, which may indicate that the step size is set too large, resulting in unrelated genes being clustered together. In this case, the preset step size needs to be reduced to improve the clustering effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 Schematic diagram of the process of screening sepsis markers based on gene co-expression network according to an embodiment of the present invention;

[0044] Figure 2 Schematic diagram of the process of performing GO function analysis on each target gene cluster according to an embodiment of the present invention;

[0045] Figure 3 A schematic diagram of a process for adjusting clustering preset parameters according to an embodiment of the present invention;

[0046] Figure 4 A schematic diagram of a process for determining similar chain pairs of GO term subsets based on comparison results according to an embodiment of the present invention;

[0047] Figure 5 For sample clustering and difference analysis of the embodiment of the present invention;

[0048] Figure 6 is the co-expression module of sepsis identified by WGCNA in an embodiment of the present invention;

[0049] Figure 7 The hub gene detected by the embodiment of the present invention;

[0050] Figure 8 This is the enrichment result of GO Molecular function in the embodiment of the present invention;

[0051] Figure 9 Correlation analysis between hub genes and immune cells in the embodiment of the present invention;

[0052] Figure 10 This is the scRNA-seq analysis of the present invention;

[0053] Figure 11 The expression of the hub gene of the embodiment of the present invention and its clinical significance;

[0054] Figure 12 This is the expression of CTSB and ATP6V0D1 proteins in circulation according to the present invention. DETAILED DESCRIPTION

[0055] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0056] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0057] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0058] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0059] See also Figure 1 As shown, it is a schematic structural diagram of a method for screening sepsis markers based on a gene co-expression network according to an embodiment of the present invention. The present invention provides a method for screening sepsis markers based on a gene co-expression network, comprising:

[0060] Step S1, obtaining gene expression data of the experimental group and the control group, preprocessing the gene expression data, and screening a number of differentially expressed genes based on the first feature data;

[0061] Step S2, performing hierarchical cluster analysis on each of the differentially expressed genes to combine each of the differentially expressed genes into a plurality of target gene modules;

[0062] The hierarchical cluster analysis of each differentially expressed gene includes: performing a primary clustering on each differentially expressed gene based on a preset step size to obtain an initial gene cluster, and performing a secondary clustering on the initial gene cluster based on the second feature data to obtain a plurality of target gene clusters, performing GO function analysis on each target gene cluster, and determining whether to adjust the clustering preset parameters according to the analysis results;

[0063] The clustering preset parameters are a preset step size and a preset segmentation length;

[0064] Adjusting the clustering preset parameters includes correcting the preset step length to a corrected step length, and correcting the preset segmentation length to a corrected step length;

[0065] The first feature is gene expression level, and the second feature data is clinical feature data;

[0066] Step S3, selecting hub genes from each target gene module;

[0067] Among them, several target gene clusters with the highest gene correlation in each target gene module are selected as hub genes.

[0068] In this embodiment, the experimental group consists of samples from patients with sepsis, and the control group consists of samples from healthy controls. Preprocessing of the gene expression data includes normalizing and filtering the gene expression data to remove low-quality genes and samples to obtain differentially expressed genes. The differentially expressed genes represent genes screened for sepsis-related diseases. A gene module refers to a set of functionally related genes that may play a role in the same biological pathway. By dividing the gene modules, gene groups that participate in specific biological processes in sepsis can be identified. The second characteristic data may also be age or gender. Because some biomarkers may be abnormally expressed in sepsis patients but have little correlation with the pathophysiological process of sepsis, such as procalcitonin (PCT) levels, which are significantly elevated in infected patients and decrease after the infection is resolved or after adequate antibiotic treatment, the accuracy of the screening results is improved by associating the differentially expressed genes with the clinical characteristic data, that is, performing secondary clustering on the initial gene clusters. The data normalization method is TPM.

[0069] By combining gene expression data with clinical characteristic data, the accuracy of screening sepsis marker genes was improved, and possible abnormal genes unrelated to the pathophysiological process of sepsis were avoided. Through primary and secondary clustering analysis, gene modules were divided more finely. Combined with GO functional analysis, the selection of GO term subsets was optimized and the step size set for primary clustering was corrected according to the analysis results. This avoided poor clustering effect caused by the number or range of GO terms included in the GO term subset, and avoided the formation of gene clusters of genes with low correlation due to a large step size. This improved the granularity and accuracy of the analysis, optimized the division of gene modules, and achieved more refined processing and screening of sepsis markers.

[0070] Specifically, the GO terms are segmented into a number of GO term subsets using a preset segmentation length.

[0071] The preset segmentation length in this embodiment represents the number or range of GO terms selected when constructing a GO term subset. A larger segmentation length includes more GO terms, while a smaller segmentation length includes only a few related GO terms. If the segmentation length is too small, important GO terms will be excluded, thereby losing valuable information. If the segmentation length is too large, redundant information will be introduced, making the analysis results complicated and difficult to interpret.

[0072] See Figure 2 As shown, it is a schematic diagram of the process of performing GO function analysis on each target gene cluster according to an embodiment of the present invention;

[0073] Specifically, GO functional analysis of each target gene cluster includes:

[0074] Step S2001, obtaining the relationship between the GO term subsets corresponding to each target gene cluster, linking the target gene cluster with the GO term subset to obtain the corresponding gene cluster chain;

[0075] Step S2002: obtaining the actual number of gene cluster chains of the gene cluster chain and comparing it with the standard number of gene cluster chains, determining the clustering effect based on the comparison result, and adjusting the clustering preset parameters according to the similarity of the gene cluster chain pairs if the clustering effect does not meet the standard;

[0076] Wherein, the gene cluster chain pair is composed of two of the gene cluster chains.

[0077] In this embodiment, the relationship between the GO term subsets corresponding to each target gene cluster is obtained by comparing the gene sequence of each target gene cluster with an existing GO annotation database, such as UniProt, Ensembl, etc., and adding the matching GO term subset to the corresponding gene, that is, linking the target gene cluster with the GO term subset to form a gene cluster chain. The length of the GO term subset affects the matching result between the GO term subset and the gene, and determines whether each gene will be assigned one or more GO terms. Therefore, the number of gene cluster chains is statistically analyzed to measure the clustering effect, and the length of the GO term subset is adjusted based on the feedback of the clustering effect, thereby ensuring the accuracy of the GO functional analysis.

[0078] Specifically, the clustering effect is determined based on the comparison results, including:

[0079] If the actual number of gene cluster chains is less than the standard number of gene cluster chains, the clustering effect is judged to be unsatisfactory;

[0080] If the actual number of gene cluster chains is greater than or equal to the standard number of gene cluster chains, the clustering effect is determined to be up to standard.

[0081] The clustering effect is evaluated by comparing the actual number of gene cluster chains with the standard number of gene cluster chains. If the number of gene cluster chains linked by the GO term subset is large, it means that the current step size setting may be appropriate and the gene clustering effect is good. If the number of gene cluster chains linked by the GO term subset is small, it may indicate that the step size setting is too large or the GO term subset segmentation is unreasonable, resulting in unrelated genes being clustered together. Further analysis of the reasons is needed to select dynamically adjusted parameters.

[0082] See Figure 3 As shown, it is a schematic diagram of the process of adjusting the clustering preset parameters according to an embodiment of the present invention;

[0083] Specifically, the clustering preset parameters are adjusted according to the similarity of the gene cluster chain pairs, including:

[0084] Step S2012, calculating the similarity value between each pair of gene cluster chains, and performing a similarity comparison between the similarity value and a similarity threshold, and determining a GO term subset similar chain pair based on the comparison result;

[0085] Step S2022, obtaining the number of similar chain pairs in the GO term subset, recorded as the number of similar chain pairs, and calculating the percentage of the number of similar chain pairs to the total number of gene cluster chain pairs, recorded as the real-time similarity ratio;

[0086] Step S2032: Compare the standard similarity ratio with the real-time similarity ratio:

[0087] If the real-time similarity ratio is less than the standard similarity ratio, the preset step size is corrected to the corrected step size;

[0088] If the real-time similarity ratio is greater than or equal to the standard similarity ratio, the preset segmentation length is corrected to the corrected segmentation length;

[0089] The revised step length is the preset step length-1; and the revised segmentation length is 90% of the preset segmentation length.

[0090] When the real-time similarity ratio is determined to be greater than or equal to the standard similarity ratio, it means that the GO term subsets linked to multiple gene cluster chains are too similar. This situation is due to inappropriate segmentation length of GO terms, resulting in insufficient diversity of functional annotations. In this case, the preset segmentation length needs to be corrected, that is, the preset segmentation length needs to be reduced to improve the accuracy of functional annotations of gene clusters. When the real-time similarity ratio is determined to be greater than or equal to the standard similarity ratio, it means that the number of gene cluster chains linked to the same GO term subset is small, which may indicate that the step size is set too large, resulting in unrelated genes being clustered together. In this case, the preset step size needs to be reduced to improve the clustering effect.

[0091] See Figure 4 As shown, it is a schematic diagram of a process for determining similar chain pairs of GO term subsets based on comparison results according to an embodiment of the present invention;

[0092] Specifically, based on the comparison results, similar chain pairs of GO term subsets were determined to include:

[0093] Step S2112, obtaining a first comparison result and a second comparison result;

[0094] Step S2212: when the first comparison result is obtained, the corresponding gene cluster chain pair is not marked;

[0095] When the second comparison result is obtained, the corresponding gene cluster chain pair is marked as a GO term subset similar chain pair;

[0096] If the similarity value is less than the similarity threshold, a first comparison result is obtained; if the similarity value is greater than or equal to the similarity threshold, a second comparison result is obtained;

[0097] In this embodiment, the similarity value represents the similarity of GO term subsets between different gene cluster chains; the similarity threshold is set to 0.7; the similarity value is calculated using the Jaccard similarity coefficient, and the two chains in the GO term subset similar chain are represented by sets respectively. By measuring the similarity of the two sets, the similarity value is obtained by calculating the size of the intersection of the two sets and dividing it by the size of their union. This will not be repeated here.

[0098] Specifically, the relationship between the GO term subsets corresponding to each target data cluster is obtained, and the target data cluster is linked to the GO term subset, including:

[0099] Calculate the real-time similarity between each target data cluster and the GO term subset, and compare the standard similarity with each real-time similarity.

[0100] If the real-time similarity is less than the standard similarity, the relationship is determined to be irrelevant, and the target data cluster is not linked to the GO term subset;

[0101] If the real-time similarity is greater than the standard similarity, the relationship is determined to be related, and the target data cluster is linked to the GO term subset.

[0102] In this embodiment, the real-time similarity is calculated using the Pearson correlation coefficient, which ranges from -1 to 1, where 1 indicates a complete positive correlation, -1 indicates a complete negative correlation, and 0 indicates no linear correlation. For example, a target data cluster and a subset of GO terms are obtained, each subset is represented as a numerical vector, and the Pearson correlation coefficient between the two vectors is calculated to obtain the real-time similarity, which will not be repeated here. The standard similarity is set to 0.6.

[0103] Specifically, the present invention also provides a diagnostic kit for sepsis marker genes screened out according to the above-mentioned method for screening sepsis markers based on gene co-expression network, comprising:

[0104] The kit is used for detecting sepsis.

[0105] See Figure 5 As shown, it is the sample clustering and difference analysis of the embodiment of the present invention;

[0106] Specifically, in order to verify the feasibility of the screened markers in actual clinical application, such as Figure 5 As shown in Figure 3, this heatmap shows the expression of the top 20 target gene modules, and we found that the five hub genes were all upregulated in sepsis patients compared with healthy controls.

[0107] See Figure 6 As shown, it is the co-expression module of sepsis identified by WGCNA in an embodiment of the present invention;

[0108] Figure 6 A is the principal component analysis of bulk RNA-seq on PBMCs; Figure 6 B is the volcano plot of differentially expressed genes. The screening criteria for differentially expressed genes were |log2(fold change)|>1 and adjusted P value <0.05; Figure 6 C is a heat map showing the expression of the top 20 DEGs.

[0109] See Figure 7 As shown, it is the hub gene detected by the embodiment of the present invention;

[0110] Figure 7 A The upper panel is a Venn diagram, the lower panel is a PPI network, and the lines represent protein-protein interactions; Figure 7 B is the key subnetwork detected by the MCODE application and the hub genes identified by the CytoHubba application; Figure 7 C is a heat map showing the expression of hub genes.

[0111] See Figure 8 As shown, it is the enrichment result of GO Molecular function in an embodiment of the present invention;

[0112] Figure 8 A is a bar graph of the top 20 GO terms, colored by P value; Figure 8 B is the network enriched with GO terms. Nodes are colored by cluster name, and nodes with the same cluster name are close to each other.

[0113] See Figure 9 As shown, it is the correlation analysis between the hub genes and immune cells in the embodiment of the present invention;

[0114] Figure 9 A shows the comparison of immune cell abundance between sepsis patients and healthy individuals by ssGSEA analysis. The Mann-Whitney test was used to compare the differences between the two groups. Figure 9 B is a lollipop chart showing the correlation between immune cells and hub genes. The color of the point reflects the P value, and the size of the point reflects the absolute value of the correlation coefficient.

[0115] See Figure 10 As shown, it is the scRNA-seq analysis of the embodiment of the present invention;

[0116] Figure 10 A is the UMAP image of 19 cell clusters from HC and sepsis samples; Figure 10 B. Dot plot shows the marker genes for each cluster. The size of the dot indicates the proportion of cells expressing the gene, and the depth of the color indicates the average transcription level of the corresponding cluster. Figure 10C is the UMAP map of 10 cell clusters identified by cell marker genes; Figure 10 D is the relative proportion of each cell type in HC and sepsis samples; Figure 10 E shows hub genes (CTSB, CTSD, ATP6V0D1, UBE2D1, and ATP6V0C), where the color depth represents the expression intensity of individual cells; Figure 10 F represents the transcription level of hub genes in monocytes, and HC represents healthy controls.

[0117] See Figure 11 As shown, it is the expression of the hub gene of the embodiment of the present invention and its clinical significance;

[0118] Figure 11 A shows the comparison of hub gene expression (CTSB, CTSD, ATP6V0D1, UBE2D1, and ATP6V0C) between healthy individuals and sepsis patients, and the Mann-Whitney test was used to compare the differences between the two. Figure 11 B shows the receiver operating characteristic curve showing the effect of sepsis diagnosis; Figure 11 C is a Kaplan-Meier curve showing the difference in 28-day mortality among sepsis patients with different hub gene expression. The survival difference was compared using the log-rank test.

[0119] See Figure 12 As shown, it is the expression of CTSB and ATP6V0D1 proteins in the circulation according to the embodiment of the present invention;

[0120] Figure 12 A is a representative fluorescent immunocytochemistry image showing the co-expression of CD14 (green) and CTSB (red); Figure 12 B, Expression of CD14 (green) and ATP6V0D1 (red) in PBMCs of septic patients and healthy individuals; cell nuclei were stained with DAPI (blue).

[0121] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

[0122] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for screening sepsis markers based on gene co-expression networks, characterized in that: include, Obtaining gene expression data of the experimental group and the control group, preprocessing the gene expression data, and screening a number of differentially expressed genes based on the first feature data; performing hierarchical cluster analysis on the differentially expressed genes to combine the differentially expressed genes into several target gene modules; The hierarchical cluster analysis of each differentially expressed gene includes: performing a primary clustering on each differentially expressed gene based on a preset step size to obtain an initial gene cluster, and performing a secondary clustering on the initial gene cluster based on the second feature data to obtain a plurality of target gene clusters, performing GO function analysis on each target gene cluster, and determining whether to adjust the clustering preset parameters according to the analysis results; The clustering preset parameters are a preset step size and a preset segmentation length; Adjusting the clustering preset parameters includes correcting the preset step length to a corrected step length, and correcting the preset segmentation length to a corrected step length; The first feature is gene expression level, and the second feature data is clinical feature data; Hub genes were screened from each target gene module.

2. The method for screening sepsis markers based on gene co-expression network according to claim 1, characterized in that: GO functional analysis of each target gene cluster includes: Obtain the relationship between the GO term subsets corresponding to each target gene cluster, link the target gene cluster with the GO term subset to obtain the corresponding gene cluster chain; Obtaining the actual number of gene cluster chains of the gene cluster chain and comparing it with the standard number of gene cluster chains, determining the clustering effect based on the comparison result, and adjusting the clustering preset parameters according to the similarity of the gene cluster chain pairs if the clustering effect does not meet the standard; Wherein, the gene cluster chain pair is composed of two of the gene cluster chains.

3. The method for screening sepsis markers based on gene co-expression network according to claim 2, characterized in that: The clustering effect is determined based on the comparison results, including: If the actual number of gene cluster chains is less than the standard number of gene cluster chains, the clustering effect is judged to be unsatisfactory; If the actual number of gene cluster chains is greater than or equal to the standard number of gene cluster chains, the clustering effect is determined to be up to standard.

4. The method for screening sepsis markers based on gene co-expression network according to claim 2, characterized in that: According to the similarity of gene cluster chain pairs, the clustering preset parameters are adjusted, including: Calculating the similarity value between each pair of gene cluster chains, and performing a similarity comparison between the similarity value and a similarity threshold, and determining GO term subset similar chain pairs based on the comparison results; Obtain the number of similar chain pairs in the GO term subset, record it as the number of similar chain pairs, calculate the percentage of the number of similar chain pairs to the total number of gene cluster chain pairs, record it as the real-time similarity ratio, and compare the standard similarity ratio with the real-time similarity ratio: If the real-time similarity ratio is less than the standard similarity ratio, the preset step size is corrected to the corrected step size; If the real-time similarity ratio is greater than or equal to the standard similarity ratio, the preset segmentation length is corrected to the corrected segmentation length; Among them, the correction step size is the preset step size - 1; The corrected split length is 90% of the preset split length.

5. The method for screening sepsis markers based on gene co-expression network according to claim 4, characterized in that: Based on the comparison results, similar chain pairs of GO term subsets were determined, including: Obtaining a first comparison result and a second comparison result; When the first comparison result is obtained, the corresponding gene cluster chain pair is not marked; When the second comparison result is obtained, the corresponding gene cluster chain pair is marked as a GO term subset similar chain pair; If the similarity value is less than the similarity threshold, a first comparison result is obtained; if the similarity value is greater than or equal to the similarity threshold, a second comparison result is obtained.

6. The method for screening sepsis markers based on gene co-expression network according to claim 2, characterized in that: Obtain the relationship between the GO term subsets corresponding to each target data cluster, and link the target data cluster with the GO term subset, including: Calculate the real-time similarity between each target data cluster and the GO term subset, and compare the standard similarity with each real-time similarity. If the real-time similarity is less than the standard similarity, the relationship is determined to be irrelevant, and the target data cluster is not linked to the GO term subset; If the real-time similarity is greater than the standard similarity, the relationship is determined to be related, and the target data cluster is linked to the GO term subset.

7. The method for screening sepsis markers based on gene co-expression network according to claim 6, characterized in that: The GO terms are segmented into several GO term subsets using a preset segmentation length.

8. The method for screening sepsis markers based on gene co-expression network according to claim 1, characterized in that: The initial clustering of differentially expressed genes based on the preset step size includes: Adjacent differentially expressed genes are combined in sequence with a preset step size to obtain the initial gene cluster.

9. The method for screening sepsis markers based on gene co-expression network according to claim 1, characterized in that: The hub genes are CTSB, CTSD, ATP6V0D1, UBE2D1 and ATP6V0C.

10. A diagnostic kit for sepsis marker genes screened out by the method for screening sepsis markers based on gene co-expression network according to any one of claims 1 to 9, characterized in that: include, The kit is used for detecting sepsis.

Citation Information

Patent Citations

  • A method for screening potential biomarkers for gastric cancer based on weighted gene co-expression network analysis and its application.

    CN109872776B