Weighted gene co-expression network analysis-based gastric cancer and inflammatory cancer transformation marker analysis method and system

Through the weighted gene co-expression network analysis method, key modules and genes in the transformation process of gastric precancerous lesions to gastric cancer are identified, and the analysis system for transformation markers of gastric cancer is constructed, which solves the problems of early prediction and intervention of gastric cancer and improves the effectiveness of cancer prevention.

CN120356512APending Publication Date: 2025-07-22INST OF BASIC RES & CLINICAL MEDICINE CHINA ACAD OF CHINESE MEDICAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510492542.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and predict key modules and genes in the transformation process of gastric precancerous lesions to gastric cancer, and lacks effective early intervention methods.

Method used

Weighted gene co-expression network analysis method is adopted to construct a transformation marker analysis system for gastric cancer cancer through differential gene analysis, weighted gene co-expression network analysis and survival analysis, and identify key modules and genes, providing a reference for advance intervention in gastric cancer.

Benefits of technology

Important key modules and genes in the progression of gastric cancer were identified, providing early prediction and intervention reference for gastric cancer, and improving the effectiveness of cancer prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356512A_ABST
    Figure CN120356512A_ABST
Patent Text Reader

Abstract

The invention discloses a weighted gene co-expression network analysis-based gastric cancer and inflammatory cancer transformation marker analysis method and system, and relates to the technical field of bioinformatics. The method comprises the following steps: acquiring a gastritis cancer transformation related data set; carrying out difference analysis on the gastritis cancer transformation related data set; carrying out weighted gene co-expression network analysis, and constructing a plurality of corresponding weighted co-expression network modules; carrying out GO and KEGG enrichment analysis; selecting a key gene, and carrying out survival analysis on the key gene to obtain a survival analysis result; and carrying out normality and variance homogeneity test to determine the final gastric cancer and inflammatory cancer transformation marker. According to the method, data of precancerous lesions and early gastric cancer are analyzed by utilizing differential gene analysis, weighted gene co-expression network analysis, survival analysis and other methods, important key modules and key genes in the disease progress process are determined, and effective reference is provided for advanced intervention of gastric cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and more particularly, to a method and system for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis. Background Art

[0002] Studies have shown that in Asia, the incidence ratio of intestinal-type gastric cancer to diffuse-type gastric cancer is 1.42:1, and the incidence of intestinal-type gastric cancer is the highest, accounting for about 46.3%. The occurrence of gastric cancer usually starts from chronic gastritis (CG), and gradually progresses to precancerous lesions of gastric cancer (PLGC), early gastric cancer (EGC), and finally develops to advanced gastric cancer (AGC). Among them, precancerous lesions of gastric cancer include low-grade intraepithelial neoplasia (LGIN) and high-grade intraepithelial neoplasia (HGIN).

[0003] Precancerous lesions will increase the risk of cancer occurrence, and timely treatment of precancerous lesions can effectively reduce the cancer incidence rate. A multi-center population cohort study confirmed that screening and treating precancerous lesions of upper gastrointestinal cancer can reduce the population incidence of upper gastrointestinal cancer. Therefore, it is very necessary to predict and analyze early gastric cancer lesions.

[0004] Weighted gene co-expression network (WGCNA) is based on gene expression profile data, considering the expression correlation among multiple genes, regulating the correlation through soft threshold, using the topological overlap matrix (TOM) method to achieve the integration of direct and indirect correlations, and using hierarchical clustering and cut tree methods to divide genes with similar expression patterns into a group, and a group of genes is a module. Some studies have found that genes within the same module may have the same regulatory mechanism or biological function, so the module information can be used to predict genes with unknown functions. According to the intra-module connectivity and the association between the module and the phenotype, it can be used to identify candidate biomarker genes or therapeutic targets.

[0005] Exploring the key modules and targets in the process of gastric cancer precancerous lesion to cancer transformation based on the weighted gene co-expression network plays a crucial role in providing a reference for the early intervention of gastric cancer. Summary of the Invention

[0006] To overcome the above problems or at least partially solve the above problems, the present invention provides a method and system for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis, which analyzes data of precancerous lesions and early gastric cancer by methods such as differential gene analysis, weighted gene co-expression network analysis, and survival analysis, determines important key modules and key genes during the disease progression, and provides an effective reference for the early intervention of gastric cancer.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0008] In the first aspect, the present invention provides a method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis, including the following steps:

[0009] Obtain a dataset related to gastric inflammation-cancer transformation;

[0010] Perform differential analysis on the dataset related to gastric inflammation-cancer transformation to obtain differential expression gene data under various disease states; the disease states include LGIN, HGIN, and EGC;

[0011] Perform weighted gene co-expression network analysis based on the differential expression gene data under various disease states to construct corresponding multiple weighted co-expression network modules;

[0012] Perform GO and KEGG enrichment analysis based on the multiple weighted co-expression network modules to obtain gene set enrichment results;

[0013] Select key genes based on the multiple weighted co-expression network modules and the differential expression gene data under various disease states, and perform survival analysis on the key genes to obtain survival analysis results;

[0014] Perform normality and homogeneity of variance tests on the expression of the key genes based on the dataset related to gastric inflammation-cancer transformation to determine the final gastric cancer inflammation-cancer transformation markers.

[0015] The present invention obtains a gene expression dataset related to gastric cancer transformation from a corresponding database, performs differential gene analysis on it, uses the limma package to screen genes differentially expressed in LGIN, HGIN, and EGC, and observes the changes in gene expression under different disease states; then performs WGCNA analysis, constructs a gene co-expression network to capture gene modules related to LGIN, HGIN, and EGC, and then performs KEGG pathway enrichment analysis on these modules; at the same time, selects key genes with significant differences in expression for survival analysis to explore the possible effects of these key genes on the transformation of precancerous lesions of gastric cancer to gastric cancer. The present invention uses methods such as differential gene analysis, weighted gene co-expression network analysis, and survival analysis to analyze the data of precancerous lesions and early gastric cancer, determine important key modules and key genes during the disease progression process, and provide an effective reference for the early intervention of gastric cancer.

[0016] Based on the first aspect, further, the above-mentioned gene expression dataset related to gastric cancer transformation includes the GSE130823 dataset and the GSE55696 dataset. Among them, the GSE130823 dataset is used as the training set, and the GSE55696 dataset is used as the validation set.

[0017] Based on the first aspect, further, the method for analyzing gastric cancer-gastritis transformation markers based on weighted gene co-expression network analysis further includes the following steps:

[0018] Import the GSE130823 dataset and the GSE55696 dataset into the Sangerbox3.0 platform for ID annotation.

[0019] Based on the first aspect, further, the method for performing differential analysis on the gene expression dataset related to gastric cancer transformation includes the following steps:

[0020] Based on the Sangerbox3.0 platform, use the limma package to perform differential analysis on the training set.

[0021] Based on the first aspect, further, the method for performing differential analysis on the training set using the limma package includes the following steps:

[0022] Set the standard value for screening differential genes;

[0023] According to the standard value for screening differential genes, using CG as the control group, obtain the differential expression gene data under various disease states of LGIN, HGIN, and EGC respectively, and perform visualization operations on the obtained differential expression gene data.

[0024] Based on the first aspect, further, the method for performing weighted gene co-expression network analysis according to the differential expression gene data under various disease states includes the following steps:

[0025] Import the differentially expressed gene data under multiple disease states into the Sangerbox 3.0 platform for weighted gene co-expression network analysis, screen out genes with an average standard deviation ratio greater than 50%, and after obtaining and filtering outlier samples, construct weighted co-expression network modules.

[0026] Based on the first aspect, further, the method for performing GO and KEGG enrichment analysis based on multiple weighted co-expression network modules includes the following steps:

[0027] Import the gene sets included in the weighted co-expression network modules related to the three disease states of LGIN, HGIN, and EGC into the Sangerbox 3.0 platform for GO and KEGG enrichment analysis.

[0028] Based on the first aspect, further, the method for selecting key genes based on multiple weighted co-expression network modules and differentially expressed gene data under multiple disease states includes the following steps:

[0029] Select the weighted co-expression network modules most relevant to the three disease states of LGIN, HGIN, and EGC and the hub genes in the weighted co-expression network modules most relevant to the three disease states of LGIN, HGIN, and EGC;

[0030] Based on the differentially expressed gene data, select the genes that are differentially expressed in all three disease states of LGIN, HGIN, and EGC among the hub genes as key genes.

[0031] Based on the first aspect, further, the method for performing survival analysis of key genes includes the following steps:

[0032] Import the key genes into the gene expression prognosis analysis module in the Sangerbox 3.0 platform for survival analysis.

[0033] In a second aspect, the present invention provides a gastric cancer inflammation-cancer transformation biomarker analysis system based on weighted gene co-expression network analysis, including a data set acquisition module, a differential analysis module, a WGCNA analysis module, an enrichment analysis module, a key gene analysis module, and a gene testing module, wherein:

[0034] The data set acquisition module is used to acquire a data set related to gastric cancer inflammation-cancer transformation;

[0035] The differential analysis module is used to perform differential analysis on the data set related to gastric cancer inflammation-cancer transformation to obtain differentially expressed gene data under multiple disease states; the disease states include LGIN, HGIN, and EGC;

[0036] The WGCNA analysis module is used to perform weighted gene co-expression network analysis based on differential expression gene data under multiple disease states, and construct corresponding multiple weighted co-expression network modules;

[0037] The enrichment analysis module is used to perform GO and KEGG enrichment analysis based on multiple weighted co-expression network modules to obtain gene set enrichment results;

[0038] The key gene analysis module is used to select key genes based on multiple weighted co-expression network modules and differential expression gene data under multiple disease states, and perform key gene survival analysis to obtain survival analysis results;

[0039] The gene test module is used to perform normality and homogeneity of variance tests on the expression of key genes based on the gastritis-carcinoma transformation-related data set to determine the final gastric cancer gastritis-carcinoma transformation markers.

[0040] The present invention has at least the following advantages or beneficial effects:

[0041] The present invention provides a method and system for analyzing gastric cancer gastritis-carcinoma transformation markers based on weighted gene co-expression network analysis. The gene expression data set related to gastritis-carcinoma transformation is obtained from the corresponding database, and differential gene analysis is performed on it. The limma package is used to screen genes differentially expressed in LGIN, HGIN, and EGC, and observe the changes in gene expression under different disease states; then WGCNA analysis is adopted. By constructing a gene co-expression network, gene modules related to LGIN, HGIN, and EGC are captured, and then KEGG pathway enrichment analysis is performed on these modules; at the same time, key genes with significant differences in expression are selected for survival analysis to explore the possible effects of these key genes on the transformation process from gastric precancerous lesions to gastric cancer. The present invention uses methods such as differential gene analysis, weighted gene co-expression network analysis, and survival analysis to analyze the data of precancerous lesions and early gastric cancer, determine important key modules and key genes during the disease progression, and provide an effective reference for the early intervention of gastric cancer. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a flowchart of a method for analyzing gastric cancer gastritis-carcinoma transformation markers based on weighted gene co-expression network analysis according to an embodiment of the present invention;

[0044] Figure 2 Schematic diagram of the distribution of differential genes under three disease states in the embodiments of the present invention;

[0045] Figure 3 Volcano plot and heat map of differentially expressed genes under three groups of disease states in the embodiments of the present invention;

[0046] Figure 4 Schematic diagram of the number of module genes in the embodiments of the present invention;

[0047] Figure 5 Schematic diagram of the results of weighted co-expression network module analysis in the embodiments of the present invention;

[0048] Figure 6 Partial display diagram of the results of enrichment analysis of key modules in the embodiments of the present invention;

[0049] Figure 7 Schematic diagram of the results of adverse factors in the survival analysis of key genes in the embodiments of the present invention;

[0050] Figure 8 Schematic diagram of the expression of key genes in the GSE130823 data set in the embodiments of the present invention;

[0051] Figure 9 Schematic diagram of the expression of key genes in the GSE55696 data set in the embodiments of the present invention. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0053] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without any creative efforts shall fall within the protection scope of the present invention.

[0054] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0055] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0056] In the description of the embodiments of the present invention, "a plurality" represents at least two.

[0057] Embodiment:

[0058] As Figure 1 shown, in a first aspect, an embodiment of the present invention provides a method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis, including the following steps:

[0059] S1. Obtain a dataset related to gastric cancer inflammation-cancer transformation; the above dataset related to gastric cancer inflammation-cancer transformation includes the GSE130823 dataset and the GSE55696 dataset, where the GSE130823 dataset is used as the training set and the GSE55696 dataset is used as the validation set.

[0060] Furthermore, it also includes: importing the GSE130823 dataset and the GSE55696 dataset into the Sangerbox 3.0 platform for ID annotation.

[0061] In some embodiments of the present invention, the "GEOquery" R package is used to obtain a dataset related to gastric cancer inflammation-cancer transformation (GSE130823, GSE55696) from GEO. Among them, GSE130823 is used as the training set and GSE55696 is used as the validation set. The basic information of the two datasets is shown in Table 1. The expression matrices of the two datasets are imported into the Sangerbox 3.0 platform (http: / / sangerbox.com / home.html) for ID annotation.

[0062] Table 1 Dataset Table Related to Gastric Cancer Inflammation-Cancer Transformation

[0063]

[0064] S2. Perform differential analysis on the gastric cancer transformation-related dataset to obtain differential expression gene data under multiple disease states; the disease states include LGIN, HGIN, and EGC.

[0065] Further, it includes: Based on the Sangerbox 3.0 platform, use the limma package to perform differential analysis on the training set.

[0066] Further, it includes: Set the standard value for screening differential genes; according to the standard value for screening differential genes, using CG as the control group, obtain differential expression gene data under multiple disease states of LGIN, HGIN, and EGC respectively, and perform visualization operations on the obtained differential expression gene data.

[0067] In some embodiments of the present invention, based on the Sangerbox 3.0 platform, use limma to perform differential analysis on the training set, set fold change = 1.5 and p < 0.05 as the criteria for screening differential genes, use CG as the control group, obtain the differential expression genes of LGIN, HGIN, and EGC respectively, and perform visualization operations on the obtained differential gene results. At the same time, import the differential expression genes under the three disease states into the microbioinformatics platform (https: / / www.bioinformatics.com.cn / ) to perform visualization processing on the distribution of the differential expression genes of LGIN, HGIN, and EGC. The results of differential expression gene analysis are as Figure 2 and Figure 3 shown. Compared with the CG group, a total of 4065 differential expression genes were obtained in the LGIN group, of which 2030 were down-regulated genes and 2035 were up-regulated genes; 3428 differential expression genes were obtained in the HGIN group, of which 1618 were down-regulated genes and 1810 were up-regulated genes; 2706 differential expression genes were obtained in the EGC group, of which 1075 were down-regulated genes and 1631 were up-regulated genes, as shown in the volcano plot ( Figure 2 ). In the figure, A and B are the volcano plot and heat map of the differential expression genes of LGIN; C and D are the volcano plot and heat map of the differential expression genes of HGIN; E and F are the volcano plot and heat map of the differential expression genes of EGC. Take the intersection of the differential genes of LGIN, HGIN, and EGC ( Figure 1 ), and 1439 genes common to the three disease states were found.

[0068] S3. Perform weighted gene co-expression network analysis based on the differential expression gene data under multiple disease states to construct corresponding multiple weighted co-expression network modules.

[0069] Furthermore, the differentially expressed gene data under various disease states were imported into the Sangerbox3.0 platform for weighted gene co-expression network analysis, and genes with an average standard deviation ratio greater than 50% were screened out. After outlier samples were obtained and filtered, a weighted co-expression network module was constructed.

[0070] In some embodiments of the present invention, all differential gene expression profile data under three disease states in the training set were imported into the Sangerbox 3.0 platform for WGCNA analysis, the network type was selected as "unsigned", genes with a mean standard deviation ratio greater than 50% were screened, and after filtering outlier samples, a weighted co-expression network module was constructed.

[0071] Module Eigengene (ME) is the first principal component (Principal Component 1, PC1) of the gene expression data within each module, and can be regarded as the representative of the gene expression pattern of the module. The Pearson correlation coefficient between ME and the corresponding variable represents the correlation between the module and the corresponding disease state. Correlation coefficient>0 indicates positive correlation, correlation coefficient<0 indicates negative correlation, and 0<|correlation coefficient|<1 indicates that there is a certain degree of linear correlation between the two variables. The closer to 1, the closer the linear relationship.

[0072] Gene significance (GS) is used to measure the correlation between genes and phenotypes. The higher the GS value, the stronger the correlation between the gene and the trait, which may mean that the gene plays a more important role in the formation of the trait or the occurrence of the disease. Module membership (MM) is used to measure the correlation between a gene and its module. If MM is close to 1, it means that the gene is highly correlated with the module; if the MM value is close to 0, it means that the gene has little relationship with the module. The thresholds of |MM|>0.8 and |GS|>0.2 are used to screen hub genes related to phenotypes in the module.

[0073] The results of weighted co-expression network module analysis are as follows Figure 5 As shown, A: sample clustering tree diagram; B: soft threshold determination; C: average connectivity; D: dynamic tree cutting to classify the gene clustering tree and merge modules with high similarity; E: module feature vector clustering diagram; F: module and phenotype correlation heat map. Taking 0.85 as the correlation coefficient threshold, the soft threshold β=16( Figure 5 B). The minimum module size is set to 30 and the sensitivity is 3, resulting in 9 modules ( Figure 5D), namely cyan, purple, magenta, pink, red, blue, green, brown, and grey, where the grey module is considered to be a gene set that cannot be assigned to any module, and the number of all genes contained in the other 8 modules is as Figure 3 shown; among them, there are 45 hub genes of LGIN in the pink module, 35 hub genes of HGIN in the purple module, 85 hub genes of EGC in the red module, and 76, 49, and 59 hub genes of LGIN, HGIN, and EGC in the magenta module respectively.

[0074] According to the results of the correlation between modules and phenotypes ( Figure 5 F), it can be seen that multiple modules show unique correlations with the occurrence and development of EGC. Among them, the modules most relevant to LGIN, HGIN, and EGC are pink (r = -0.42, p < 0.05), purple (r = -0.26, p < 0.05), and red (r = 0.43, p < 0.05) respectively. The magenta module is the most relevant to these three disease states, and its correlations are -0.34 (p < 0.05), -0.24 (p < 0.05), and 0.26 (p < 0.05) respectively. Therefore, the pink module, purple module, red module, and magenta module are selected as the key modules for subsequent analysis in this study.

[0075] S4. Conduct GO and KEGG enrichment analysis based on multiple weighted co-expression network modules to obtain gene set enrichment results;

[0076] Furthermore, it includes: importing the gene sets contained in the weighted co-expression network modules related to the three disease states of LGIN, HGIN, and EGC into the Sangerbox 3.0 platform for GO and KEGG enrichment analysis.

[0077] In some embodiments of the present invention, the gene sets contained in some modules related to the three disease states are imported into the Sangerbox 3.0 platform for GO and KEGG enrichment analysis. This platform uses the GO annotation of genes in the R software package org.Hs.eg.db (version 3.1.0) as the background and maps the genes to the background set; at the same time, the gene annotation of the latest KEGG Pathway is obtained as the background and the genes are mapped to the background set, and the R software package clusterProfiler (version 3.14.3) is used for enrichment analysis to obtain the gene set enrichment results. The minimum gene set is set to 5, the maximum gene set is set to 5000, and P < 0.05 is set as the statistically significant entry.

[0078] Enrichment analysis was performed on the modules most relevant to LGIN, HGIN, and EGC respectively, and the modules most relevant to all three. Part of the results of the key module enrichment analysis is shown as Figure 6 follows. Among them, A, B: Partial results of GO and KEGG of the magenta module are shown; C, D: Partial results of GO and KEGG of the pink module are shown; E, F: Partial results of GO and KEGG of the purple module are shown; G, H: Partial results of GO and KEGG of the red module are shown. Analysis found that the pink module is related to biological processes such as cell adhesion, angiogenesis, and inflammatory response, and is involved in pathways such as basal cell carcinoma and proteoglycans in cancer. Since pink is the module most relevant to LGIN, there are biological processes such as inflammatory response in LGIN; the purple module most relevant to HGIN is related to biological processes such as doxorubicin metabolism, positive regulation of tumor necrosis factor secretion, and cell production of reactive oxygen species, and this module is involved in pathways such as chemical carcinogenesis, drug metabolism, and β-alanine metabolism; the red module most relevant to EGC is related to immune response and inflammatory response, and is involved in pathways such as NF-kappa B signaling pathway and Toll-like receptor signaling pathway; and the above three disease states are all related to the magenta module, which is related to complement activation and immune response, and is involved in pathways such as cytokine-cytokine receptor interaction and chemokine signaling pathway.

[0079] S5. Select key genes based on the differential expression gene data of multiple weighted co-expression network modules and multiple disease states, and perform survival analysis of the key genes to obtain the survival analysis results;

[0080] Furthermore, it includes: selecting the weighted co-expression network modules most relevant to the three disease states of LGIN, HGIN, and EGC and the hub genes in the weighted co-expression network modules most relevant to the three disease states of LGIN, HGIN, and EGC; based on the differential expression gene data, selecting the genes that are differentially expressed in all three disease states of LGIN, HGIN, and EGC among the hub genes as key genes.

[0081] Furthermore, it includes: importing the key genes into the gene expression prognosis analysis module in the Sangerbox 3.0 platform for survival analysis.

[0082] In some embodiments of the present invention, to explore the possible effects of changes in gene expression levels during the process of gastric cancer transformation from gastritis, the modules most relevant to the three disease states and the hub genes in the modules most relevant to the three disease states were selected. Genes that were differentially expressed in all three disease states among the hub genes were selected as key genes. The key genes were imported into the gene expression prognosis analysis module in the sangerbox3.0 platform for survival analysis. The sample source and data source were both selected as TCGA. The data transformation was selected as log(x + 0.001). Samples with an expression level of 0 were removed, and samples with a minimum time shorter than 30 days were excluded. The survival data was selected as Overall servial. In downstream analysis, TCGA-STAD (gastric cancer) was selected to plot the KMplot.

[0083] The intersection genes of the hub genes and the differentially expressed genes common to the three disease states were taken as key genes, and a total of 20 key genes were obtained (Table 2). Through survival analysis of the 20 key genes, it was found that 13 key genes (ANKRD29, CYP2AB1P, EFEMP1, FAM20A, FZD8, MAL, MAP7D2, PHYHD1, PRKCB, RDH12, REP15, STOX2, ZNF662) had significant prognostic differences ( Figure 7 ) p < 0.05, HR > 1, all of which were unfavorable factors. Statistical analysis of the expression of all key genes at different disease stages found that IGHG3, FCRL3, and PRKCB were highly expressed in EGC, while the remaining genes were highly expressed in CG. The results are as Figure 8 shown. Five key genes were not annotated or expressed in the validation set, namely IGHG3, CYP2AB1P, REP15, SLC26A9, and ZNF662. The results of statistical analysis of the expression of the remaining 15 key genes at different disease stages are as Figure 9 shown, among which FCRL3 and PRKCB were highly expressed in EGC, and the remaining genes were highly expressed in CG.

[0084] Table 2 Key genes and their corresponding modules

[0085]

[0086] S6. Perform normality and homogeneity of variance tests on the expression of key genes based on the gastritis-gastric cancer transformation-related dataset to determine the final gastric cancer gastritis-cancer transformation markers.

[0087] In some embodiments of the present invention, the data is analyzed using the SPSSPRO online platform (https: / / www.spsspro.com / ), and visualized using GraphPad Prism 8.0.2 software. The expression of key genes in different groups is tested for normality and homogeneity of variance. For genes with expression that satisfies the normal distribution, one-way ANOVA is used for testing; for key genes that do not satisfy normality but satisfy homogeneity of variance, Welch's variance test is adopted. LSD is used for multiple comparisons between groups, and p < 0.05 indicates that the difference is statistically significant.

[0088] The present invention obtains a gene expression dataset related to gastric cancer transformation from the corresponding database, performs differential gene analysis on it, screens for genes differentially expressed in LGIN, HGIN, and EGC using the limma package, and observes the changes in gene expression under different disease states; limma, as an efficient differential expression analysis tool, is particularly suitable for processing small-sample data and can effectively identify gene expression changes related to disease states. Then, WGCNA analysis is adopted. By constructing a gene co-expression network, gene modules related to LGIN, HGIN, and EGC are captured, and then KEGG pathway enrichment analysis is performed on these modules; the advantage of WGCNA lies in its ability to systematically analyze the co-regulatory relationships between genes and is particularly suitable for studying the molecular mechanisms of complex diseases. Through WGCNA analysis, gene modules significantly related to the three disease states are identified, and hub genes are screened from them. Hub genes have a high connectivity in the module and are usually considered the core of the regulatory network and may play a key role in the occurrence and development of diseases. At the same time, key genes with significant differences in expression are selected for survival analysis to explore the possible effects of these key genes on the transformation of precancerous gastric lesions to gastric cancer. Twenty hub genes that are significantly differentially expressed in LGIN, HGIN, and EGC are selected as key genes, and these genes may play an important role in the entire gastric cancer transformation process.

[0089] To evaluate the clinical significance of these key genes, survival analysis was performed on them. The results showed that high expression of 13 out of the 20 key genes was significantly associated with reduced patient survival rate. These genes were ANKRD29, CYP2AB1P, EFEMP1, FAM20A, FZD8, MAL, MAP7D2, PHYHD1, PRKCB, RDH12, REP15, STOX2, and ZNF662, suggesting that these genes might be adverse factors, and inhibiting the high expression of these genes had positive significance for the prognosis of gastric cancer patients. Through one-way ANOVA on these 20 key genes and combining their expression in the validation set, it was found that 4 genes (such as FCRL3, EFEMP1, ANKRD29, and STOX2) showed an expression trend highly consistent with the original data set, supporting their key role in gastric cancer development; while 11 genes (such as ABCC5, ALDH3A1, SOSTDC1, etc.) had an expression trend not completely consistent with the original data set, possibly affected by sample heterogeneity.

[0090] The protein encoded by the FCRL3 gene is mainly expressed in B lymphocytes and is involved in regulating the B cell receptor signaling pathway and immune tolerance. Its polymorphism may lead to abnormal B cell function, promote the production of autoantibodies against gastric parietal cells or intrinsic factor, and thus induce or exacerbate chronic inflammatory damage of the gastric mucosa. At the EGC stage, FCRL3+ B cells may promote a regulatory phenotype by secreting IL-10 or expressing PD-L1, thereby inhibiting the anti-tumor immune response.

[0091] EFEMP1 (EGF containing fibulin extracellular matrix protein 1) belongs to the fibulin family and is a fibronectin-like protein containing epidermal growth factor. It plays a role in the extracellular matrix and has different functions in various cancer types. It can act as both an oncogene and a tumor suppressor. In lung cancer cells, downregulation of EFEMP1 expression is associated with tumor growth and invasion ability. EFEMP1 can inhibit the migration ability of liver cancer cells. Downregulation of EFEMP1 increases the migration ability of liver cancer cells and is related to the expression levels of ERK1 / 2, MMP2, and MMP9. MMPs are a group of enzymes that can degrade the extracellular matrix and play an important role in tumor invasion and metastasis. The expression levels of MMP2 and MMP9 are negatively correlated with the expression level of EFEMP1, indicating that at the precancerous stage (LGIN, HGIN), EFEMP1 may promote the transformation to gastric cancer by affecting the expression of MMPs. Since EFEMP1 may play a role in inhibiting tumor growth and invasion in different cancers, treatment targeting EFEMP1 may have clinical application value in future gastric cancer treatment.

[0092] ANKRD29 belongs to the ankyrin repeat domain (ANKRD) protein family, which is widely involved in protein interactions and signal transduction in eukaryotes. It may affect the treatment response of patients by regulating the immune microenvironment and drug sensitivity, suggesting its potential application value in immunotherapy and chemotherapy. ANKRD29 is a tumor suppressor gene and plays an important role in the tumorigenesis of non-small cell lung cancer in particular. ANKRD29 inhibits the progression of non-small cell lung cancer by regulating key biological processes such as cell proliferation, migration and apoptosis. A decrease in its expression level will lead to a significant increase in the proliferation and migration abilities of non-small cell lung cancer cells, and the restored expression can inhibit tumor growth by inhibiting the cell cycle process and regulating related signaling pathways.

[0093] STOX2 is a gene encoding a transcription factor. After being inhibited by miR-30a, it can activate multiple signaling pathways related to tumor progression, including the ERK, AKT and P38 pathways. The activation of these pathways can further promote cell survival, proliferation and metastasis.

[0094] In this invention, methods such as differential gene analysis, weighted gene co-expression network analysis, and survival analysis are used to analyze the data of precancerous lesions and early gastric cancer, to determine important key modules and key genes during the disease progression, and to provide an effective reference for the early intervention of gastric cancer.

[0095] In a second aspect, the embodiments of this invention provide a gastric cancer inflammation-cancer transformation biomarker analysis system based on weighted gene co-expression network analysis, including a data set acquisition module, a differential analysis module, a WGCNA analysis module, an enrichment analysis module, a key gene analysis module, and a gene testing module, where:

[0096] The data set acquisition module is used to acquire a data set related to gastric inflammation-cancer transformation;

[0097] The differential analysis module is used to perform differential analysis on the data set related to gastric inflammation-cancer transformation to obtain differential expression gene data under multiple disease states; the disease states include LGIN, HGIN, and EGC;

[0098] The WGCNA analysis module is used to perform weighted gene co-expression network analysis according to the differential expression gene data under multiple disease states to construct corresponding multiple weighted co-expression network modules;

[0099] The enrichment analysis module is used to perform GO and KEGG enrichment analysis based on multiple weighted co-expression network modules to obtain gene set enrichment results;

[0100] A key gene analysis module, which is used to select key genes based on multiple weighted co-expression network modules and differentially expressed gene data under multiple disease states, and perform survival analysis on the key genes to obtain survival analysis results;

[0101] A gene testing module, which is used to perform normality and homoscedasticity tests on the expression of key genes based on a dataset related to gastric cancer transformation from gastritis, so as to determine the final gastric cancer transformation marker from gastritis.

[0102] Through the cooperation of multiple modules such as a dataset acquisition module, a differential analysis module, a WGCNA analysis module, an enrichment analysis module, a key gene analysis module, and a gene testing module, this system obtains a gene expression dataset related to gastric cancer transformation from gastritis from the corresponding database, performs differential gene analysis on it, uses the limma package to screen genes differentially expressed in LGIN, HGIN, and EGC, and observes the changes in gene expression under different disease states; then performs WGCNA analysis, constructs a gene co-expression network to capture gene modules related to LGIN, HGIN, and EGC, and then performs KEGG pathway enrichment analysis on these modules; at the same time, selects key genes with significant differences in expression for survival analysis to explore the possible impacts of these key genes on the transformation from precancerous lesions of gastric cancer to gastric cancer. This invention uses methods such as differential gene analysis, weighted gene co-expression network analysis, and survival analysis to analyze the data of precancerous lesions and early gastric cancer, determines important key modules and key genes during the disease progression process, and provides an effective reference for the early intervention of gastric cancer.

[0103] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can also be implemented in other ways. The method and system embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the methods and systems, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0104] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0105] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0106] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claimed rights.

Claims

1. A method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis, characterized in that, It includes the following steps: Obtain a dataset related to gastric cancer transformation from gastritis; Perform differential analysis on the dataset related to gastric cancer transformation from gastritis to obtain differential expression gene data under multiple disease states; the disease states include LGIN, HGIN, and EGC; Perform weighted gene co-expression network analysis based on the differential expression gene data under multiple disease states to construct corresponding multiple weighted co-expression network modules; Perform GO and KEGG enrichment analysis based on multiple weighted co-expression network modules to obtain gene set enrichment results; Select key genes based on multiple weighted co-expression network modules and differential expression gene data under multiple disease states, and perform survival analysis of key genes to obtain survival analysis results; Perform normality and homogeneity of variance tests on the expression of key genes based on the dataset related to gastric cancer transformation from gastritis to determine the final gastric cancer transformation biomarker from gastritis.

2. The method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis according to claim 1, wherein The dataset related to gastric cancer transformation from gastritis includes the GSE130823 dataset and the GSE55696 dataset. Among them, the GSE130823 dataset is used as the training set, and the GSE55696 dataset is used as the validation set.

3. The method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis according to claim 2, wherein It also includes the following steps: Import the GSE130823 dataset and the GSE55696 dataset into the Sangerbox3.0 platform for ID annotation.

4. A method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis according to claim 2, characterized in that The method for performing differential analysis on the dataset related to gastric cancer transformation from gastritis includes the following steps: Based on the Sangerbox3.0 platform, use the limma package to perform differential analysis on the training set.

5. A method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis according to claim 4, characterized in that, The method for using the limma package to perform differential analysis on the training set includes the following steps: Set the standard value for screening differential genes; According to the standard value for screening differential genes, using CG as the control group, obtain differential expression gene data under multiple disease states of LGIN, HGIN, and EGC respectively, and perform visualization operations on the obtained differential expression gene data.

6. The gastric cancer inflammation-cancer transformation biomarker analysis method based on weighted gene co-expression network analysis according to claim 1, characterized in that The method for performing weighted gene co-expression network analysis based on the differential expression gene data under multiple disease states includes the following steps: Import the differential expression gene data under multiple disease states into the Sangerbox3.0 platform for weighted gene co-expression network analysis, screen out genes with an average standard deviation ratio greater than 50%, and after filtering outlier samples, construct weighted co-expression network modules.

7. The gastric cancer inflammation-cancer transformation biomarker analysis method based on weighted gene co-expression network analysis according to claim 1, characterized in that The method for performing GO and KEGG enrichment analysis based on multiple weighted co-expression network modules includes the following steps: Import the gene sets included in the weighted co-expression network modules related to the three disease states of LGIN, HGIN, and EGC into the Sangerbox3.0 platform for GO and KEGG enrichment analysis.

8. A method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis according to claim 1, characterized in that, The method for selecting key genes based on multiple weighted co-expression network modules and differential expression gene data under multiple disease states includes the following steps: Select the weighted co-expression network modules most relevant to the three disease states of LGIN, HGIN, and EGC and the hub genes in the weighted co-expression network modules most relevant to the three disease states of LGIN, HGIN, and EGC; Based on the differentially expressed gene data, genes that are differentially expressed in all three disease states of LGIN, HGIN, and EGC among the hub genes are selected as key genes.

9. A method for analyzing gastric cancer inflammation-cancer transformation markers based on weighted gene co-expression network analysis according to claim 1, characterized in that The method for performing survival analysis of key genes includes the following steps: Import the key genes into the gene expression prognosis analysis module in the Sangerbox 3.0 platform for survival analysis.

10. A gastric cancer inflammation-cancer transformation biomarker analysis system based on weighted gene co-expression network analysis, characterized in that, It includes a dataset acquisition module, a differential analysis module, a WGCNA analysis module, an enrichment analysis module, a key gene analysis module, and a gene testing module, where: The dataset acquisition module is used to obtain datasets related to gastric cancer transformation. The differential analysis module is used to perform differential analysis on the datasets related to gastric cancer transformation to obtain differentially expressed gene data under various disease states; the disease states include LGIN, HGIN, and EGC. The WGCNA analysis module is used to perform weighted gene co-expression network analysis based on the differentially expressed gene data under various disease states to construct corresponding multiple weighted co-expression network modules. The enrichment analysis module is used to perform GO and KEGG enrichment analysis based on multiple weighted co-expression network modules to obtain gene set enrichment results. The key gene analysis module is used to select key genes based on multiple weighted co-expression network modules and differentially expressed gene data under various disease states, and perform survival analysis of key genes to obtain survival analysis results. The gene testing module is used to perform normality and homogeneity of variance tests on the expression of key genes based on the datasets related to gastric cancer transformation to determine the final gastric cancer transformation markers.