Differential gene regulation and control network reconstruction method based on mutual information and redundancy regulation and control filtering system
By employing a differential gene regulatory network inference method based on mutual information, combined with conditional mutual information and redundant edge removal mechanisms, the problem of insufficient accuracy and interpretability in the differential identification of gene regulatory networks in existing technologies is solved, and gene regulatory network reconstruction with high accuracy and interpretability under multiple conditions is achieved.
Patent Information
- Application Number
- CN202510997428.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies struggle to accurately identify differences in gene regulatory networks under multiple conditions, especially in small-sample, high-dimensional data where it is difficult to capture nonlinear relationships and eliminate redundant regulatory edges, resulting in insufficient accuracy and interpretability of differential gene regulatory network reconstruction.
A differential gene regulatory network inference method based on mutual information is adopted, which combines conditional mutual information and redundant edge removal mechanism. Through mutual information calculation, conditional mutual information analysis and redundancy determination, significantly changed regulatory relationships are identified. Granger causality and data-driven regulatory direction determination are used to construct differential gene regulatory networks.
It improves the accuracy and interpretability of differential gene regulatory network reconstruction, is applicable to multiple gene expression platforms, and the output results can be used for downstream analysis and pathway modeling. It also supports the export and visualization of multiple data formats.
Smart Images

Figure CN120895100A_ABST
Abstract
Description
I. TECHNICAL FIELD
[0001] The present application belongs to the cross field of bioinformatics, systems biology and artificial intelligence, and particularly relates to a technology for modeling, analyzing and reconstructing a differential regulatory network based on an information theory method of mutual information, and more particularly relates to a method and system for high-confidence network reconstruction using conditional mutual information and redundant edge elimination mechanism. The present application is suitable for gene regulation mechanism research, disease biomarker mining, drug target identification and transcription factor prediction, etc. in biomedical and precision medicine scenarios. II. BACKGROUND
[0002] With the rapid development of high-throughput sequencing technologies (such as RNA-Seq, scRNA-seq), researchers can comprehensively obtain gene expression information at the cellular or tissue level. Building gene regulatory networks (GRN) based on these data is of great significance for understanding transcriptional regulation mechanisms, identifying key regulatory factors, and assisting disease mechanism research and drug development.
[0003] In actual biological research, the regulatory networks under different physiological conditions, developmental stages or disease conditions differ significantly. For example, the transcriptional regulation pattern in cancer cells may have undergone large-scale remodeling compared to normal cells. Therefore, the inference of differential gene regulatory networks, i.e., comparing the differences in regulatory networks under two or more conditions, has become one of the key directions in recent years of systems biology research.
[0004] Traditional network inference methods (such as ARACNE, CLR, GENIE3, etc.) mainly focus on network reconstruction under a single condition, and are difficult to directly capture the regulatory changes between different conditions. Although some methods attempt to construct networks under multiple conditions independently and then compare their differences, due to noise accumulation and low statistical efficiency, this "post-comparison" approach is difficult to accurately identify true differential regulatory relationships.
[0005] In addition, most existing methods rely on assumptions of linear correlation, regression models or tree structure regression, making it difficult to capture non-linear regulatory relationships, especially in small sample high-dimensional data.
[0006] In recent years, network inference methods based on mutual information have received widespread attention due to their ability to capture non-linear relationships, no distribution assumptions, etc. However, how to effectively compare the mutual information network structure under different conditions, eliminate redundant regulatory edges, and improve the credibility of regulatory edges remains a problem to be solved.
[0007] Therefore, there is an urgent need for a technical solution that can: 1) simultaneously consider gene expression data under multiple conditions; 2) identify regulatory differences using non-parametric information theory methods; 3) directly infer the structure of the regulatory network difference and perform significance analysis, in order to improve the accuracy and interpretability of the reconstructed regulatory network of differential genes.
[0008] The present application is based on the above problems, and proposes a mutual information-based differential gene regulatory network inference method and system, which combines conditional mutual information and redundant edge removal mechanism to effectively identify significant changes in regulatory relationships under different conditions, and is suitable for various biological information scenarios. III. SUMMARY
[0009] The present application aims to provide a mutual information-based and regulatory redundancy removal mechanism-based differential gene regulatory network reconstruction method and system to improve the identification accuracy of regulatory difference edges and the biological interpretability of differential networks.
[0010] To achieve the above-mentioned purpose, the technical scheme of the present application is as follows:
[0011] A mutual information-based differential gene regulatory network inference method, comprising the following steps:
[0012] S01: data preprocessing
[0013] S11: input the original gene expression matrix where X (c) is the counts format gene expression data, Z (c) (n,j) represents the expression of the jth gene in the nth sample, n∈{1,2,…,N c} represents the number of samples or single cells in the cth cell type, c∈{1,2,…,C} represents the cth cell type, N c represents the number of samples or single cells in the cth cell type, and p represents the number of genes.
[0014] S12: normalize, handle missing values, and filter noise for the gene expression data Z (c) , filter out genes and samples with excessively high or low expression, standardize gene length and sequencing depth, and convert it to a TPM format gene expression matrix
[0015] S02 mutual information calculation:
[0016] Mutual information (MI) between gene expression data is a measure of the non-linear dependence relationship between two genes in the expression pattern. It is more powerful than the traditional Pearson correlation coefficient (which only detects linear relationships) and is suitable for discovering potential regulatory effects, co-expression modules, or constructing gene regulatory networks.
[0017] S21 In the cth cell type, the mutual information between each pair of genes is estimated by the Kraskov estimator.
[0018] S22 In the cth cell type, the mutual information between all pairs of genes forms a matrix where
[0019] M (c) (i,j) = I(X i ,X j )
[0020] where X i represents the expression vector of the ith gene, X j represents the expression vector of the jth gene, I (c) (X i ,X j ) represents the mutual information between the expression of the two genes, and its theoretical definition is
[0021]
[0022] where p(x i ) and p(x j ) represent the marginal probability density of the ith and jth gene expression, respectively, and p(x i ,x j ) represents the joint probability density of the ith and jth gene expression.
[0023] S03 Redundant indirect regulation filtering:
[0024] S31: The core idea is to analyze whether the regulation between genes is a direct regulation relationship through conditional mutual information (CMI) analysis, and then remove the indirect redundant regulation edges. In the cth cell type of the cell lineage, under the condition of controlling the third variable k, according to the redundancy removal strategy, such as Data processing inequality, remove indirect regulation for any ternary gene group i, j, k ∈ {1, 2, …, p}. The specific process is as follows:
[0025] S321: Set the mutual information threshold τ c , for each potential regulation edge (i, j), if I(i, j) < τ c , then it is considered that the regulation between the gene pair (i, j) in the cell type is noise and needs to be filtered out. Only the edges with I(X i ,X j )>τ are retained to form the candidate regulation edge set E.
[0026] S322: Conditioned mutual information calculation and redundancy determination
[0027] For each candidate edge X i →X j , find possible mediator X k , calculate
[0028] If one of the following conditions is satisfied, the edge is determined as redundant:
[0029] • I(X i ,X j |X k ) << I(Xi; Xj), i.e. the regulatory relationship is significantly reduced given X k ;
[0030] • Or satisfy the data processing inequality:
[0031] I(X i ,X j ) < min{I(X i ,X k ), I(X j ,X k )}
[0032] S333 Redundant regulatory edge removal Remove edges that satisfy the redundancy determination condition from the candidate edge set E, and the remaining edges constitute the final network structure G'.
[0033] S04: Regulatory direction determination:
[0034] S41 If there is time series data, determine the regulatory direction by Granger causality test or time delay mutual information. If there is time series data, determine the regulatory direction by Granger causality or time delay mutual information. If the past value of another variable can still significantly improve the prediction ability of the current value after controlling the influence of its own past value, we say that this variable Granger causes the current variable. Determine whether X Granger causes Y.
[0035] S42 If there is no time series, use data-driven regulatory direction scoring function. Determine the causal structure between variables by statistical conditional independence.
[0036] S05: Network construction and threshold optimization:
[0037] S51 Set global or local threshold according to mutual information strength to construct gene regulatory graph;
[0038] S52 Optimize threshold parameters combined with cross-validation or biological prior knowledge.
[0039] S06 Difference network construction module
[0040] This module is used to identify gene pairs with significant regulatory changes under different physiological or treatment conditions, and specifically includes the following steps:
[0041] S61 Mutual information difference calculation:
[0042] For each pair of genes X i ,X j , its mutual information values I (c=0) (X i ,X j ) and I (c=1) (X i ,X j ) are calculated under multiple conditions (e.g., normal group and disease group), and the difference is calculated:
[0043] ΔI(X i ,X j )=|I (c=0) (X i ,X j )-I (c=1) (X i ,X j )|
[0044] S62 Significance test:
[0045] Through non-parametric permutation test (Permutation test) or bootstrap method, the statistical significance analysis of mutual information difference of each pair of genes is carried out. The specific method is:
[0046] (1) Randomly shuffle sample labels, recalculate mutual information difference, repeat N times;
[0047] (2) Generate difference distribution under null hypothesis;
[0048] (3) Compare whether the original difference is in the upper 1-α quantile (e.g., 95%) of the distribution, to obtain the p-value;
[0049] (4) Multiple hypothesis testing adjustment, such as using Benjamini-Hochberg method to control FDR.
[0050] S63 Candidate edge screening:
[0051] Gene pairs with significant mutual information difference (p-value below threshold, and ΔI exceeding absolute value threshold) are retained as candidate difference regulation edges for further processing by the redundancy regulation filtering module.
[0052] S64 Robustness evaluation:
[0053] If multiple mutual information estimation methods are used (such as Kraskov estimation and kernel method), the screening robustness can be further improved through intersection or weighted fusion.
[0054] S7 output and visualization:
[0055] Output the inferred differential gene regulatory network structure (adjacency matrix or edge set form); support visualization output and standard format export.
[0056] This module is used to present the final differential gene regulatory network structure in a graphical way, and supports network export in multiple standard formats, which is convenient for subsequent analysis and integrated use. Its specific functions and implementation steps include:
[0057] (1) Network graphical display
[0058] • Use graph structure drawing library (such as Cytoscape, Graphviz or D3.js) to build network topology visualization;
[0059] • Gene nodes are represented as circles or ellipses, and regulatory edges are represented as directed arrows;
[0060] ● The thickness / color of the edge can be adjusted to reflect the difference in mutual information;
[0061] • Support automatic clustering coloring of nodes by functional category, expression intensity or enriched pathway;
[0062] ● Provide interactive functions such as node click to display annotation information, local zoom, edge filtering, etc.
[0063] (2) Differential regulation marking
[0064] ● Mark the regulatory edges that change significantly under different conditions (such as using red to represent enhanced regulation and using blue to represent weakened regulation);
[0065] • Support user to set threshold (such as mutual information change absolute value > 0.1) to filter the regulatory modules of interest;
[0066] ● Can highlight or export specific pathways or regulatory subnetworks separately.
[0067] (3) Network export function
[0068] • Provide one-key export function, support the following formats:
[0069] o GraphML: XML structure compatible with graph database and visualization tools;
[0070] o GML (Graph Modeling Language): simple text structure, supports multiple attributes;
[0071] o JSON or CSV: suitable for subsequent loading and processing in Python, R, JavaScript, etc. environments;
[0072] o PNG / SVG format images: for report or paper illustrations;
[0073] • Support exporting node and edge attributes (such as mutual information values, conditional mutual information, difference significance p values, etc.) together.
[0074] (4) Module integration features
[0075] • Linkage with difference edge filtering and redundancy identification module, can instantly view the regulatory change network;
[0076] • Provide graph structure comparison tools to support edge analysis between two regulatory networks (such as union, intersection, difference);
[0077] • Support saving user operation views and network states to achieve project-level reuse.
[0078] Advantages
[0079] Compared with the prior art, the present application has the following advantages:
[0080] 1) High accuracy: Introducing conditional mutual information and redundancy control mechanism, effectively identifying real regulatory relationships;
[0081] 2) Strong applicability: Support nonlinear relationship modeling, suitable for various gene expression platforms (such as bulk RNA-seq, scRNA-seq);
[0082] 3) Good interpretability: The output results can be directly used for downstream enrichment analysis and pathway modeling;
[0083] 4) Modular system design: Easy to transplant and integrate into existing bioinformatics analysis platforms. IV. DESCRIPTION OF DRAWINGS
[0084] In order to more clearly illustrate the technical solutions of the present application, the present application will be further described in conjunction with the following drawings. The accompanying drawings are only schematic descriptions and do not constitute a limitation on the present application.
[0085] Figure 1 : System flowchart of the method of the present application. This figure shows the complete process from the input of the original expression data, mutual information estimation, difference edge filtering, conditional mutual information analysis, redundancy edge removal to network output.
[0086] Figure 2 : Redundant regulatory identification schematic diagram. Take the gene X→Y←Z structure as an example, to illustrate how to determine whether X→Y is a redundant regulatory edge through conditional mutual information I(X,Y|Z).
[0087] Figure 3 Differential regulatory network result visualization. V. DETAILED DESCRIPTION
[0088] In order to make the technical solutions of the present application clearer and more explicit, the present application is further described below in combination with the drawings.
[0089] Please refer to Figure 1 The differential gene regulatory network inference method based on mutual information of the present application comprises the following modules:
[0090] S1: data input:
[0091] The input gene expression data is received, and multiple formats (such as CSV, TSV, ExpressionSet object, etc.) are supported, and standardization preprocessing is performed, including normalization (such as log2 or TPM conversion), missing value filling (such as k-nearest neighborhood method) and noise filtering.
[0092] S2: mutual information calculation module:
[0093] Mutual information (MI) between gene expression data is a measure of the non-linear dependence relationship between two genes in the expression pattern. It is more powerful than the traditional Pearson correlation coefficient (only detects linear relationship), and is suitable for discovering potential regulatory effects, co-expression modules or constructing gene regulatory networks. In the c-th cell type, the mutual information between all gene pairs constitutes a matrix Wherein
[0094] M (c) (i,j)=I(X i ,X j )
[0095] Wherein X i represents the expression vector of the i-th gene, X j represents the expression vector of the j-th gene, I (c) (X i ,X j ) represents the mutual information between the expression amounts of two genes, and its theoretical definition is
[0096]
[0097] For mutual information calculation between two variables, a proper method is needed to estimate the probability density, thus to calculate the mutual information between two variables. Kraskov is a commonly used method. In the cth cell type of the cell lineage, the mutual information value between each pair of genes is estimated by Kraskov operator (KSG, Kraskov– Grassberger), and a mutual information matrix M (k) (i,j) is constructed, which represents the mutual information between gene i and gene j in the cth cell type.
[0098] The Kraskov estimator estimates the mutual information (Mutual Information, MI) between each pair of genes, which is a non-parametric mutual information estimation method based on k-neighbor method for continuous variables. It can be used for continuous data such as gene expression data, avoiding the curse of dimensionality problem caused by histogram or kernel density estimation.
[0099] The Kraskov estimator is implemented by the NPEET library of Python to estimate the mutual information between any pair of genes.
[0100] The Kraskov estimator is sensitive to the number of samples, and the number of samples ≥ 50 is relatively stable.
[0101] It is recommended to normalize or standardize (such as z-score) the expression data first.
[0102] k is usually selected as 3-5, and the larger the k is, the smoother the estimation of local density is.
[0103] The mutual information matrix is symmetric, and the mutual information between a variable and itself is not considered, so the main diagonal of M is usually 0 (or ignored).
[0104] S03: Redundant regulatory filtering module:
[0105] S31: The core idea is to analyze the conditional mutual information (Conditional Mutual Information, CMI) to determine whether the regulation between genes is a direct regulation relationship, and then remove the indirect redundant regulation edges. In the cth cell type of the cell lineage, under the condition of controlling the third variable k, according to the redundancy removal strategy, such as Data processing inequality, remove the indirect regulation for any three gene groups i, j, k ∈ {1, 2, …, p}. The specific process is as follows:
[0106] S321: Set the mutual information threshold τ c , for each potential regulatory edge (i,j), if I(i,j)<τ cSo we consider the regulatory effect between genes (i, j) in this cell type as noise and need to filter it out by setting it to zero only keeping I(X i , j ) > τ.
[0107] S322: Conditional mutual information computation and redundancy determination
[0108] For each candidate edge X i → X j , find possible intermediate genes X k and compute
[0109] If one of the following conditions is satisfied, the edge is determined as redundant:
[0110] • I(X i , X j | X k ) « I(X i ; X j ), i.e. the regulatory relationship drops significantly given X k ;
[0111] • Or satisfy the data processing inequality:
[0112] I(X i , X j ) < min{I(X i , X k ), I(X j , X k )}
[0113] S333 Redundant regulatory edge removal Remove edges that satisfy the redundancy determination condition from the candidate edge set E, and the remaining edges constitute the final network structure G'.
[0114] S04: Regulatory direction determination module:
[0115] If there is time series data, determine the regulatory direction by Granger causality or time delay mutual information. If the past values of another variable can significantly improve the prediction ability of the current value after controlling the influence of its own past values, we say that this variable Granger causes the current variable. To determine whether X Granger causes Y, compare two models:
[0116] Model 1 (contains only the lagged terms of Y itself):
[0117] Y t = a0+ a1Y t-1 +... + a p Y t-p + εt
[0118] Model 2 (adding lagged terms of X):
[0119]
[0120] If model 2 is significantly better than model 1 (usually by F-test), then X Granger causes Y.
[0121] That is, assume
[0122] Null hypothesis H0: X does not Granger cause Y (i.e. γ1= γ2=... = γ p = 0)
[0123] Alternative hypothesis H1: X Granger causes Y (at least one γ j is not 0)
[0124] Use python's statsmodels library to perform Granger causality test.
[0125] If there is no time series, use data-driven regulatory direction scoring function. Determine the causal structure between variables through statistical conditional independence.
[0126]
[0127]
[0128] Use python's CausalDiscoveryToolbox library to determine the causal relationship without time series.
[0129] S05: Network construction and threshold optimization module
[0130] S51: Set global or local threshold according to mutual information strength, construct gene regulatory graph;
[0131] S52: Optimize threshold parameters combined with cross-validation or biological prior knowledge.
[0132] S06: Differential network construction module: by comparing the differences between different gene regulatory networks on cell lineages, infer the differential gene regulatory network
[0133] For gene pairs (i, j), if their regulatory relationship changes on different cell types, specifically, if one of the following two conditions is met:
[0134] (1) A c=0 (i,j) = 0 on cell type c = 0, and A c=1(i,j)≠0;
[0135] (2) On cell type c=0 A c=0 (i,j)≠0, and on cell type c=1 A c=1 (i,j)=0
[0136] We say that the regulatory relationship between gene pair (i,j) has changed significantly,
[0137] S61 Mutual information difference calculation:
[0138] For each gene pair X i ,X j , calculate its mutual information value I (c=0) (X i ,X j ) and I (c=1) (X i ,X j ) under multiple conditions (e.g. normal group and disease group), and then calculate the difference:
[0139] ΔI(X i ,X j ) = |I (c=0) (X i ,X j )-I (c=1) (X i ,X j )|
[0140] S62 Significance test:
[0141] Through non-parametric permutation test (Permutation test) or bootstrap method, statistical significance analysis is carried out on the mutual information difference value of each gene pair. The specific method is as follows:
[0142] (1) Randomly shuffle the sample labels and recalculate the mutual information difference value, repeat N times;
[0143] (2) Generate the difference value distribution under the null hypothesis;
[0144] (3) Compare whether the original difference value is in the upper 1-α quantile (e.g. 95%) of the distribution, so as to obtain the p value;
[0145] (4) Multiple hypothesis testing adjustment, such as using Benjamini-Hochberg method to control FDR.
[0146] S63 Candidate edge screening:
[0147] The gene pairs with significant mutual information difference (p-value below threshold and AI exceeding absolute value threshold) are retained as candidate differential regulation edges for further processing by the redundant regulation filtering module.
[0148] S64 Robustness evaluation:
[0149] If multiple mutual information estimation methods (such as Kraskov estimation and kernel method) are used, the screening robustness can be further improved through intersection or weighted fusion.
[0150] S07: Output and visualization module: output network structure, which can be exported as GraphML, GML or TXT format, support network graph interactive visualization based on Cytoscape.js or Plotly.
[0151] This module is used to present the final obtained differential gene regulation network structure in a graphical way, and supports network export in multiple standard formats, which is convenient for subsequent analysis and integrated use. Its specific functions and implementation steps include:
[0152] (1) Network graphical display
[0153] • Use graph structure drawing library (such as Cytoscape, Graphviz or D3.js) to build network topology visualization;
[0154] • Gene nodes are represented as circles or ellipses, and regulation edges are represented as directed arrows;
[0155] • The thickness / color of the edge can be adjusted to reflect the degree of mutual information difference;
[0156] • Support automatic clustering coloring of nodes by functional category, expression intensity or enriched pathway;
[0157] • Provide interactive functions such as node click to display annotation information, local zoom, edge filtering, etc.
[0158] (2) Differential regulation marking
[0159] • Mark the regulation edges that change significantly under different conditions (such as use red to represent enhanced regulation and use blue to represent weakened regulation);
[0160] • Support user to set threshold (such as mutual information change absolute value > 0.1) to filter the regulation modules of interest;
[0161] • Can highlight or export specific pathways or regulation sub-networks separately.
[0162] (3) Network export function
[0163] • Provide one-key export function, support the following formats:
[0164] o GraphML: XML structure compatible with graph databases and visualization tools;
[0165] o GML (Graph Modeling Language): simple text structure, supports multiple attributes;
[0166] o JSON or CSV: suitable for subsequent loading and processing in Python, R, JavaScript, etc. environments;
[0167] o PNG / SVG format images: for report or thesis illustrations;
[0168] • Support exporting node and edge attributes (such as mutual information values, conditional mutual information, difference significance p values, etc.) together.
[0169] (4) Module integration features
[0170] • With the difference edge filtering and redundancy identification module, you can instantly view the regulatory change network;
[0171] • Provide graph structure comparison tools to support edge analysis between two regulatory networks (such as union, intersection, difference); support saving user operation views and network states to achieve project-level reuse.
Claims
1. A method for reconstructing differential gene regulatory networks based on mutual information, characterized in that, Includes the following steps: (1) Data preprocessing: Collect gene expression data under two or more different biological conditions and perform data normalization and other processing respectively; (2) Mutual information calculation: The mutual information between all gene pairs was calculated under the expression data of all biological conditions to obtain multiple mutual information matrices; (3) Filtering out redundant indirect control effects: Based on the degree of decrease in conditional mutual information or information transmission inequality, determine whether there is redundant control and remove redundant control edges from the candidate edges. (4) Determining the direction of regulation: Based on whether the gene expression data contains time series information, the direction of regulation is determined through causal relationships; (5) Network construction and threshold optimization: Comparative analysis of multiple mutual information matrices was performed to identify gene pairs with significant changes in mutual information and screen them as candidate differential regulatory edges; (6) Differential network building module: gene pairs that exhibit significant regulatory changes under different physiological or treatment conditions; (7) Output the inferred differential gene regulatory network structure (adjacency matrix or edge set form); support visualization output and standard format export.
2. The method according to claim 1, characterized in that, The mutual information calculation in step (2) includes any one of the nonparametric methods, such as Kraskov nearest neighbor estimation, kernel density estimation, or copula transformation.
3. The method according to claim 1, characterized in that, Redundant control effects in step (3) are filtered out. The redundancy control judgment criterion is characterized by satisfying any of the following conditions: (a)I(X i ,X j |X k ) << I(X i ,X j That is, the regulatory relationship under a given X k It then declined significantly; (b) Or it satisfies the Data Processing Inequality: I(X i ,X j )<min{I(X i ,X k ),I(X j ,X k )}。 4. The method according to claim 1, characterized in that, The differential gene regulatory network can output as an adjacency matrix, edge list, or graph structure file, including GraphML, GML, or JSON formats.
5. A differential gene regulatory network reconstruction system based on mutual information, characterized in that, include: • Data preprocessing module, used to receive gene expression data under multiple conditions; • Mutual information computing module, used to construct mutual information networks under various conditions; • Redundant indirect control filtering module, used to identify candidate edges with significant changes in mutual information between conditions; • The control direction determination module is used to calculate conditional mutual information and eliminate redundant control edges; • The network construction and threshold optimization module is used to output the control network structure for different categories of samples; • The differential network building module is used to construct differential regulation networks between samples of different categories; • Output and visualization module, used to display and export differential network results.
Citation Information
Cited By
LCD display screen production quality collaborative management and control method based on detection data chain
CN121146296A