A functional module analysis method for multi-center source gene expression data

CN122551901APending Publication Date: 2026-08-11XIN HUA HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

针对现有技术的不足,本发明提供了一种多中心来源基因表达数据的功能模块分析方法,具备可重复性提高、可复现性提高、生物学意义提高的优点,解决了单个中心来源模块分析结果可信度较低的问题

Benefits of technology

1、该多中心来源基因表达数据的功能模块分析方法,可重复性提高:仅保留3个中心以上均存在高可信度模块归属的基因,说明该基因的模块关系具备稳健性。根据基因与模块的相关性推断基因在模块内的重要性,提示基因在单个中心中的重要性越高,那么在所有中心中的重要性也应该越高,即中心之间应具有一致性。所以我们利用单中心的基因-模块相关性与多个中心的中位基因-模块相关性进行相关性分析,反应多个中心之间的一致性。仅保留显著相关的模块,说明该单中心模块在多个中心间的一致性较高。通过上述高的稳健性和一致性保障了单中心模块分析结果的可重复性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551901A_ABST
    Figure CN122551901A_ABST
Patent Text Reader

Abstract

This invention relates to the field of biology and discloses a method for functional module analysis of gene expression data from multiple centers. The method includes single-center module identification, single-center robust module identification, single-center overlapping robust module merging, multi-center consensus module identification, sub-module identification of the consensus module, and functional annotation of the sub-modules. Specifically, it includes the following steps: S1, identifying genes with high-confidence module attribution in at least three centers; S2, identifying single-center robust modules; and S3, obtaining the identification results of the multi-center consensus module and its sub-modules. This method improves reproducibility by ensuring the reproducibility of single-center module analysis results through high robustness and consistency. It also improves reproducibility by extracting sub-modules from the consensus module results based on the robust module partitioning results of each center, preserving the differences in module partitioning between centers, and further improving the reproducibility of our consensus module and sub-module results in any single center.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, specifically to a method for functional module analysis of gene expression data from multiple centers. Background Technology

[0002] Gene module analysis algorithms such as WGCNA and EPIG only support module analysis using data from a single center with weak batch effects. Strong batch effects from multiple center sources render directly merging analysis results unreliable, and the matching degree between analysis results from different center sources is low. There is a lack of systematic analysis methods to effectively integrate module analysis results from multiple center sources. These techniques are all based on the correlation or adjacency relationships between genes, dividing genes into different modules according to the algorithm's correlation threshold. Therefore, strong batch effects lead to module results primarily reflecting batch-related gene expression differences, rather than biologically significant results. Furthermore, module analysis results from a single center source often inevitably contain certain technical components and errors due to the limited technology platform and inherent errors in population composition, resulting in results that are not entirely biologically significant and cannot be corroborated by results from other center sources. This leads to weak evidence from single-center data, i.e., low reliability of module analysis results.

[0003] Based on the single-center module analysis results obtained by the EPIG algorithm, this application constructs a complete multi-center integrated analysis process, which can realize the screening and identification of robust functional modules and obtain system gene module analysis results with multi-center evidence, high robustness and high functional purity. Therefore, a functional module analysis method for gene expression data from multiple centers is proposed. Summary of the Invention

[0004] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a functional module analysis method for gene expression data from multiple centers, which has the advantages of improved reproducibility, reproducibility, and biological significance, and solves the problem of low reliability of analysis results from single-center sources.

[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: a method for functional module analysis of gene expression data from multiple centers, including single-center module identification, single-center robust module identification, single-center overlapping robust module merging, multi-center consensus module identification, sub-module identification of consensus modules, and functional annotation of sub-modules, specifically including the following steps: S1, genes with high-confidence module attribution in at least 3 centers; S2, Single-center robust module identification; S3. Obtain the identification results of the multi-center consensus module and sub-modules.

[0006] Preferably, the single-center module identification specifically involves Z-score normalization of the log2 (RMA) data from gene chips and the log2 (TPM) data from next-generation sequencing, categorized by gene. If a pre-defined grouping exists, it is used as the biological repeat annotation according to the pre-defined grouping. If no pre-defined grouping exists, Spearman correlation analysis is performed on the normalized gene expression data, and unsupervised hierarchical clustering is performed on the Spearman correlation matrix between samples using the Complete method. The clustering results are then grouped as biological repeat annotations. The processed data is used as input, and the output of the EPIG algorithm is obtained under default parameters, including the gene module attribution results and the gene module attribution relationship confidence results.

[0007] Preferably, the single-center robust module identification involves the following steps: the reliability results of the gene module attribution relationship are categorized into three levels (low, medium, and high) by the algorithm's default parameters. Only genes and their module attribution results with high reliability at least in three centers are retained. If a preset grouping exists, based on the Spearman correlation analysis results between gene-group correlation and gene-module correlation, modules with significant correlation (P < 0.05) are retained as single-center modules. Genes within unretained single-center modules are reassigned to the retained single-center modules with the highest correlation based on gene-module correlation. The median gene-module correlation of all centers is calculated based on the gene-module correlation of the single centers. The Spearman correlation analysis results between the single-center gene-module correlation and the median gene-module correlation are calculated again, and modules with significant correlation are retained as single-center robust modules. The threshold for significant correlation is P < 0.05. Simultaneously, the correlation coefficient must be higher than the threshold, which is determined by the intersection of the cumulative density curves of the correlation coefficients of modules with P ≥ 0.05 and modules with P < 0.05. Genes that are highly confidently assigned to robust modules in at least one center are retained. If a gene is assigned to a non-robust module in other centers, it is reassigned to the most relevant robust module based on gene-module correlation.

[0008] Preferably, the single-center overlapping robust module merging is as follows: based on the module attribution relationship with the second highest gene-module correlation in each single-center robust module, the number of genes in the module with the second highest correlation within that module is divided by the total number of genes in that module as the overlap ratio. The average of the overlap ratios between the two modules is taken as the module overlap coefficient. Modules with overlap coefficients higher than the threshold of 0.6-0.9 are merged to obtain the final single-gene robust module result.

[0009] Preferably, the multi-center consensus module identification involves pairwise comparison of the results of multiple single-center robust modules. The center with fewer robust modules is designated as the reference module, and the genes from the robust module of the other center are designated as the target module. The target module is matched to the reference module with the highest enrichment score, calculated by dividing the number of overlapping genes between the target module and the reference module by the total number of genes in the reference module. After matching, the matching rate between the two centers is calculated by dividing the number of overlapping genes between the matched modules by the total number of genes in all modules. The average matching rate between each center and all other centers is calculated, and the center with the highest average is designated as the reference center. Finally, the robust module results from other centers that are all matched to the reference center are considered the matched module results. Genes repeatedly assigned to the same matched module in more than half of the centers, along with their module affiliation, are retained as the multi-center consensus module results.

[0010] Preferably, the sub-module identification of the consensus module is as follows: A multi-center consensus module can be divided into different robust modules in different single centers, indicating the existence of sub-modules. The n robust modules in the reference center are sorted from largest to smallest based on the number of genes, and assigned values ​​from 1 to n. The m robust modules in other centers are first divided into n parts based on the matching module results and sorted according to the matching module assignments of the reference center. Then, within each part, several modules are sorted from largest to smallest based on their enrichment scores with the matching modules. Finally, values ​​from 1 to m are assigned according to the overall order. The same assigned value represents both the consensus module and each robust module in a single center. To avoid the assignment scale affecting the sub-module partitioning results, the above assignments are Z-score normalized by center. For each consensus module's assignment matrix, if the average value of the assignment matrices differs significantly between single centers, it indicates the continued influence of the assignment scale. In this case, Z-score normalization by center is performed again. Then, unsupervised hierarchical clustering is performed based on the normalized assignment matrix, dividing the consensus module into a varying number of sub-modules according to the differences in the height of the clustering trees. These consensus modules and sub-modules are the results of robust and mutually corroborating gene modules among multiple centers.

[0011] Preferably, the functional annotation of the sub-modules involves functional enrichment analysis of the genes within the sub-modules using the KEGG and GO databases. This allows direct observation that the sub-modules within a consensus module exhibit high functional purity, meaning they possess similar biological functions. For example, genes within a sub-module may significantly and collectively perform the same biological function, and sub-modules within a consensus module may also perform similar biological functions in a highly consistent manner. In contrast, different consensus modules may primarily function different biological functions, such as cell cycle, extracellular matrix, and immune response.

[0012] Compared with existing technologies, this invention provides a method for functional module analysis of gene expression data from multiple centers, which has the following advantages: 1. The functional module analysis method for gene expression data from multiple centers improves reproducibility: Only genes with high-confidence module affiliations from three or more centers are retained, indicating robustness in the module relationships of these genes. Inferring gene importance within a module based on gene-module correlation suggests that higher gene importance in a single center should imply higher importance across all centers, indicating consistency among centers. Therefore, we use single-center gene-module correlation and median gene-module correlation from multiple centers for correlation analysis to reflect consistency across multiple centers. Retaining only significantly correlated modules indicates high consistency of the single-center module across multiple centers. This high robustness and consistency ensure the reproducibility of single-center module analysis results.

[0013] 2. Improved reproducibility of the functional module analysis method for gene expression data from multiple centers: Due to significant differences and varying numbers of robust module divisions across different centers, we proposed a reproducible method to match and integrate module results from different centers. This method matches module results from all centers to the same reference module, resulting in reproducible consensus module results across multiple centers with different experimental conditions, experimental times, and technical platforms. Furthermore, based on the robust module division results from each center, sub-modules were extracted from the consensus module results, preserving the differences in module division across centers, further improving the reproducibility of our consensus module and sub-module results in any single center.

[0014] 3. The functional module analysis method for multicenter-sourced gene expression data enhances biological significance: Conventional analysis methods often yield single-center gene modules that perform a wide range of different biological functions, making it difficult to simply describe the overall function of the gene module. Our consensus modules and sub-modules, however, possess highly pure biological functions, with genes within each module performing similar biological functions, allowing for a more intuitive and straightforward understanding of the biological significance behind the module. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the genes in this invention that are at least 3 centers and belong to a certain module with high confidence; Figure 2 This is a schematic diagram of the single-center robust module identification of the present invention; Figure 3 This is a schematic diagram illustrating the identification and functional annotations of the consensus module and sub-modules of the present invention. Figure 4 This is a schematic diagram illustrating the module analysis of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Please see Figure 1-4 A method for functional module analysis of gene expression data from multiple centers includes single-center module identification, single-center robust module identification, single-center overlapping robust module merging, multi-center consensus module identification, sub-module identification of consensus modules, and functional annotation of sub-modules. Specifically, it includes the following steps: S1, genes with high-confidence module attribution in at least 3 centers; S2, Single-center robust module identification; S3. Obtain the identification results of the multi-center consensus module and sub-modules.

[0018] Furthermore, the single-center module identification specifically involves Z-score normalization of the log2 (RMA) data from gene chips and the log2 (TPM) data from next-generation sequencing, categorized by gene. If a pre-defined grouping exists, it is used as the biological repeat annotation according to the pre-defined grouping. If not, Spearman correlation analysis is performed on the normalized gene expression data, and unsupervised hierarchical clustering is performed on the Spearman correlation matrix between samples using the Complete method. The clustering results are then grouped as biological repeat annotations. Using the above-processed data as input, the output of the EPIG algorithm is obtained under default parameters, including the gene module attribution results and the gene module attribution relationship confidence results.

[0019] Furthermore, the single-center robust module identification involves classifying the gene module attribution reliability results into three levels—low, medium, and high—based on the algorithm's default parameters. Only genes and their module attribution results with high reliability at least in three centers are retained. If a preset grouping exists, based on the Spearman correlation analysis results between gene-group correlation and gene-module correlation, modules with significant correlation (P < 0.05) are retained as single-center modules. Genes within unretained single-center modules are reassigned to the retained single-center modules with the highest correlation based on gene-module correlation. The median gene-module correlation for all centers is calculated based on the gene-module correlation of the single centers. The Spearman correlation analysis results between single-center gene-module correlation and median gene-module correlation are calculated again, and modules with significant correlation are retained as single-center robust modules. The threshold for significant correlation is P < 0.05. Simultaneously, the correlation coefficient must be higher than the threshold, which is determined by the intersection of the cumulative density curves of the correlation coefficients of modules with P ≥ 0.05 and modules with P < 0.05. Genes that are highly confidently assigned to robust modules in at least one center are retained. If a gene is assigned to a non-robust module in other centers, it is reassigned to the most relevant robust module based on gene-module correlation.

[0020] Furthermore, the single-center overlapping robust module merging is as follows: based on the module attribution relationship with the second highest gene-module correlation in each single-center robust module, the number of genes in the module with the second highest correlation within that module is divided by the total number of genes in that module as the overlap ratio. The average of the overlap ratios between the two modules is taken as the module overlap coefficient. Modules with an overlap coefficient higher than the threshold of 0.6-0.9 are merged to obtain the final single-gene robust module result.

[0021] Furthermore, the multi-center consensus module identification involves pairwise comparisons of the results from multiple single-center robust modules. The center with fewer robust modules is designated as the reference module, and the genes from the other center's robust module are designated as the target module. The target module is matched to the reference module with the highest enrichment score, calculated by dividing the number of overlapping genes between the target and reference modules by the total number of genes in the reference module. After matching, the matching rate between the two centers is calculated by dividing the number of overlapping genes between the matched modules by the total number of genes in all modules. The average matching rate between each center and all other centers is calculated, and the center with the highest average is designated as the reference center. Finally, the robust module results from other centers that are all matched to the reference center are considered the matched module results. Genes repeatedly assigned to the same matched module in more than half of the centers, along with their module affiliation, are retained as the multi-center consensus module results.

[0022] Furthermore, the sub-module identification of the consensus module: A multi-center consensus module can be divided into different robust modules in different single centers, indicating the existence of sub-modules. The n robust modules in the reference center are sorted from largest to smallest based on the number of genes, and assigned values ​​from 1 to n. The m robust modules in other centers are first divided into n parts based on the matching module results and sorted according to the matching module assignments of the reference center. Then, within each part, several modules are sorted from largest to smallest based on their enrichment scores with the matching module. Finally, values ​​from 1 to m are assigned based on the overall order. The same assigned value represents both the consensus module and each robust module in a single center. To avoid the assignment scale affecting the sub-module partitioning results, the above assignments are Z-score normalized by center. For each consensus module's assignment matrix, if the average value of the assignment matrices differs significantly between single centers, it indicates the continued influence of the assignment scale. Therefore, Z-score normalization by center is performed again. Then, unsupervised hierarchical clustering is performed based on the normalized assignment matrix, dividing the consensus module into a varying number of sub-modules according to the differences in the height of the clustering trees. These consensus modules and sub-modules are the results of robust and mutually corroborating gene modules among multiple centers.

[0023] Furthermore, the functional annotation of the submodules involves functional enrichment analysis of the genes within the KEGG and GO databases. This allows direct observation that the submodules within the consensus module exhibit high functional purity, meaning they possess similar biological functions. For example, genes within a submodule significantly and collectively perform the same biological function, and submodules within the consensus module also exhibit highly consistent and similar biological functions. Different consensus modules, however, primarily function in different biological areas, such as cell cycle, extracellular matrix, and immune response.

[0024] Example: A method for functional module analysis of gene expression data from multiple centers includes single-center module identification, single-center robust module identification, merging of overlapping single-center robust modules, multi-center consensus module identification, sub-module identification of consensus modules, and functional annotation of sub-modules. Specifically, it includes the following steps: S1, genes with high-confidence module attribution in at least 3 centers; S2, Single-center robust module identification; S3. Obtain the identification results of the multi-center consensus module and sub-modules; See Figure 1 In step S1, genes with high-confidence module attribution were identified from at least three centers. This case study included data from nine centers, with each center yielding anywhere from a dozen to several dozen modules. See Figure 2In step S2, robust modules are identified in a single center. Using P < 0.05 and correlation coefficient > 0.3 as thresholds, robust modules from 9 centers are screened. Genes with high confidence belonging to robust modules within each center are retained. For genes belonging to non-robust modules in other centers, the genes are reassigned to robust modules, ultimately resulting in only robust modules being retained. (A) The cumulative density curves of correlation coefficients between modules with P < 0.05 and modules with P ≥ 0.05 intersect at 0.3 as the screening threshold for robust modules; (B) The screening results of robust modules in a single center; (C) The attribution of the aforementioned high-confidence genes to robust modules within a single center. See Figure 3 Step S3: Identification results of multi-center consensus modules and sub-modules. Nine centers were matched to four reference modules of the GSE14333 center to obtain four consensus modules. Sub-modules containing differences in the single-center module division were identified again for each of the four consensus modules. Functional enrichment analysis results show that the consensus modules and sub-modules have high functional purity. In this case, consensus module CP1 mainly consists of components of different parenchymal cells such as extracellular matrix, epithelial cells, and myocytes; CP2 mainly involves RNA splicing modification and cell cycle related to nucleic acid processing; CP3 mainly involves immune-related functions; and CP4 mainly involves signal transduction. Among them, (A) are the matching results of consensus modules in each center; and (B) are the identification and functional enrichment analysis of sub-modules, annotating the function of sub-module genes.

[0025] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for functional module analysis of gene expression data from multiple centers, characterized in that: This includes single-center module identification, single-center robust module identification, merging of overlapping single-center robust modules, multi-center consensus module identification, sub-module identification of consensus modules, and functional annotation of sub-modules, specifically including the following steps: S1, genes with high-confidence module attribution in at least 3 centers; S2, Single-center robust module identification; S3. Obtain the identification results of the multi-center consensus module and sub-modules.

2. The method for functional module analysis of gene expression data from multiple centers according to claim 1, characterized in that: The single-center module identification specifically involves Z-score normalization of the log2 (RMA) data of gene chips and the log2 (TPM) data of next-generation sequencing by gene.

3. The functional module analysis method of multi-centric source gene expression data according to claim 1, characterized in that: The single-center robust module identification: The reliability results of the gene module attribution relationship are divided into three levels: low, medium, and high by the algorithm's default parameters. Only genes and their module attribution results with high reliability attribution relationships in at least three centers are retained.

4. The functional module analysis method of multi-centric source gene expression data according to claim 1, characterized in that: The single-center overlapping robust module merging is as follows: based on the module with the second highest gene-module correlation in each single-center robust module, the number of genes in the module with the second highest correlation within that module is divided by the total number of genes in that module as the overlap ratio. The average of the overlap ratios between the two modules is taken as the module overlap coefficient. Modules with overlap coefficients higher than the threshold of 0.6-0.9 are merged to obtain the final single-gene robust module result.

5. The functional module analysis method of multi-centric source gene expression data according to claim 1, characterized in that: The multi-center consensus module identification involves comparing the results of multiple single-center robust modules pairwise. The center with fewer robust modules is used as the reference module, and the genes of the robust module in the other center are used as the target module. The target module is matched to the reference module with the highest enrichment score based on the enrichment score obtained by dividing the number of overlaps between the target module and the reference module by the total number of genes in the reference module.

6. The functional module analysis method of multi-centric source gene expression data according to claim 1, wherein: Sub-module identification of the consensus module: The multi-center consensus module can be divided into different robust modules in different single centers, indicating the existence of sub-modules.

7. The functional module analysis method of multi-centric source gene expression data according to claim 1, characterized in that: Functional annotation of the submodules: Functional enrichment analysis of the genes in the submodules was performed based on the KEGG and GO databases.