A method, device and medium for analyzing a MICP functional gene regulation network

By using binary processing of gene expression data and quantum mutual information methods, a gene regulatory network was constructed, which solved the problems of non-robustness in gene regulatory network analysis and difficulty in capturing high-order relationships in existing technologies. This enabled the capture of high-order regulatory patterns among multiple genes and the localization of core genes, thus optimizing strain modification and engineering applications.

CN121709015BActive Publication Date: 2026-05-08LONGYAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LONGYAN UNIV
Filing Date
2026-02-10
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing gene regulatory network analysis methods suffer from problems such as data noise, sparsity, unstable results, difficulty in capturing high-order multi-gene synergistic relationships, difficulty in identifying causal directions, and misjudgment of environmental driving effects when processing single-cell and environmental microbial data, especially in non-model microorganisms.

Method used

After dimensionality reduction through binarization and principal component analysis, the data are mapped to one-dimensional gene sequences by spatial filling curves. Quantum state basis vectors are constructed, and quantum mutual information is calculated by decomposing and compressing matrix product state tensors. A gene regulatory backbone network is constructed and multi-gene regulatory modules are identified. A comprehensive regulatory network is then constructed to identify core genes or functional modules.

Benefits of technology

While ensuring computational feasibility, this method captures high-order regulatory patterns among multiple genes, overcomes the limitations of traditional methods, locates core functional genes, optimizes strain modification and engineering applications, and provides theoretical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709015B_ABST
    Figure CN121709015B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of bioinformatics, and discloses a MICP functional gene regulation network analysis method, device, equipment and medium. The expression data of the MICP functional gene is binarized and mapped into a one-dimensional ordered sequence. Secondly, the gene state sequence is encoded into a quantum state vector, and a matrix product state tensor network is used for efficient compression representation. On this basis, not only the quantum mutual information between the genes is calculated to construct a binary regulation skeleton network, but also the multi-element quantum mutual information is further calculated to identify significant multi-gene synergistic or redundant regulation modules. Finally, reliable relationships are screened through statistical tests, and a comprehensive regulation network is integrated and constructed for function analysis. Through quantumization and tensor network, the bottleneck of traditional methods in high-order relationship capture is broken through, and the robustness and analysis depth of network inference are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a method, apparatus, device, and medium for analyzing MICP functional gene regulatory networks. Background Technology

[0002] Existing methods for inferring gene regulatory networks (such as correlation analysis, mutual information, regression models, and Bayesian networks) each have their theoretical advantages, but they generally have shortcomings when processing single-cell and environmental microbial data. The main problems include: data noise and sparsity leading to unstable results; difficulty in capturing high-order polygenic synergistic relationships; difficulty in identifying causal directions; and the frequent misjudgment of environmental driving effects and spurious correlations as direct regulation. Furthermore, while mutual information and permutation tests can detect nonlinear relationships, they suffer from large estimation biases and excessive computational costs when sample sizes are insufficient.

[0003] In the context of MICP research, gene regulation is influenced by the combined effects of multiple species, environmental factors, and metabolic pathways. Existing methods often struggle to reveal complex high-order regulatory patterns and lack reproducibility in real-world applications. While attempts such as multi-omics integration and deep learning methods can alleviate these problems to some extent, they rely on prior annotation and are difficult to validate experimentally, with limitations being particularly pronounced in non-model microorganisms.

[0004] Therefore, overall, existing technologies have significant shortcomings in capturing high-order dependencies, statistical robustness, causal orientation determination, and environmental adaptability. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method, apparatus, device, and medium for analyzing the MICP functional gene regulatory network.

[0006] According to one aspect of this application, a method for resolving the MICP functional gene regulatory network is disclosed, the method comprising:

[0007] The raw expression data of MICP functional genes from multiple samples are binarized to obtain a binary expression matrix, wherein the rows of the binary expression matrix represent genes, the columns represent samples, and the matrix elements are used to characterize the expression status of the corresponding gene in the corresponding sample.

[0008] For each gene in the binary expression matrix, its two-dimensional feature vector is mapped to a one-dimensional gene sequence through a space-filling curve, wherein genes that are close in position in the gene feature distribution map remain adjacent in the one-dimensional gene sequence.

[0009] The one-dimensional gene sequence determines the binary state sequence of genes for each sample, and encodes each binary state sequence of genes into a quantum state basis vector;

[0010] Construct a system quantum state vector representing the entire gene system based on all quantum state basis vectors of all samples;

[0011] The quantum state vector of the system is decomposed and compressed using a matrix product state tensor network to obtain the matrix product state representation of the quantum state vector;

[0012] Based on the matrix product state representation, calculate the binary quantum mutual information between any pair of genes and the multi-dimensional quantum mutual information between multiple genes;

[0013] A gene regulatory backbone network is constructed based on the binary quantum mutual information, and a multi-gene regulatory module is identified based on the multi-quantum mutual information.

[0014] A comprehensive regulatory network is constructed based on the gene regulatory backbone network and the multi-gene regulatory module;

[0015] The MICP functional gene expression data of the sample to be tested are input into the integrated regulatory network for analysis to identify the core genes or functional modules in the network that play a role in MICP efficiency, so as to determine the target sites for strain modification based on the identified core genes or functional modules.

[0016] In some embodiments, the binarization of the raw expression data of MICP functional genes from multiple samples to obtain a binary expression matrix includes:

[0017] The original expression data is subjected to dimensionality reduction processing based on principal component analysis, including:

[0018] The original expression data is zero-mean normalized, and the covariance matrix of the zero-mean normalized matrix is ​​calculated.

[0019] The covariance matrix is ​​subjected to eigenvalue decomposition, and the first target principal components are selected as low-dimensional representations.

[0020] The data points of each gene in the covariance matrix are projected onto a two-dimensional plane composed of the first target principal components to obtain the two-dimensional coordinates of each gene.

[0021] In some embodiments, mapping each gene in the binary expression matrix to a one-dimensional gene sequence based on its two-dimensional feature vector using a space-filling curve includes:

[0022] The two-dimensional coordinates of each gene are normalized and mapped to the target square region;

[0023] Calculate the one-dimensional index value of each gene on the Hilbert space curve based on its two-dimensional coordinates.

[0024] Sort all genes in ascending order of their index values ​​to obtain a one-dimensional permutation sequence.

[0025] In some embodiments, determining the binary state sequence of each sample based on the one-dimensional gene sequence and encoding each binary state sequence of the gene into a quantum state basis vector includes:

[0026] For each sample, the binary expression state of each gene is obtained sequentially according to the gene order defined by the one-dimensional gene sequence, forming the binary state sequence of the sample.

[0027] Each binary bit in the binary state sequence is mapped to the corresponding computational ground state of a qubit, where 0 is mapped to... 1 is mapped to 0 indicates that the gene is in an inactive state, and 1 indicates that the gene is active.

[0028] The binary state sequence of each sample is mapped to multiple qubits, and the direct product of the ground states is calculated to obtain the quantum state basis vector of each sample.

[0029] In some embodiments, constructing a system quantum state vector representing the entire gene system based on all quantum state basis vectors of all samples includes:

[0030] Statistical analysis of each quantum state basis vector in all samples The frequency of occurrence is used to obtain its probability distribution. ;

[0031] Based on the probability distribution Construct a system quantum state vector representing the entire gene system;

[0032] The system's quantum state vector is represented by the following formula:

[0033] ;

[0034] In the formula, These are quantum state basis vectors. It is the probability amplitude of the quantum state basis vector.

[0035] In some embodiments, the step of decomposing and compressing the system quantum state vector using a matrix product state tensor network to obtain the matrix product state representation of the quantum state vector includes:

[0036] The quantum state vector of the system is decomposed into a product sequence of tensors using the matrix product state method, where each tensor corresponds to a gene, and the bond dimension between adjacent tensors is used to characterize the correlation strength between genes.

[0037] Establish a chain-like matrix product state structure consisting of N tensors connected in series, as the matrix product state representation of the quantum state vector.

[0038] In some embodiments, the construction of a gene regulatory backbone network based on the binary quantum mutual information and the identification of multi-gene regulation modules based on the multi-dimensional quantum mutual information include:

[0039] A symmetric gene dependency matrix is ​​constructed based on the binary quantum mutual information between all gene pairs.

[0040] The permutation test is performed on each binary quantum mutual information in the symmetric gene dependence matrix to screen out regulatory gene pairs;

[0041] Based on the regulatory gene pairs, the gene regulatory backbone network is constructed, wherein each gene is a network node and gene pairs are connected by edges.

[0042] Permutation tests are performed on all the aforementioned multivariate quantum mutual information values ​​to screen out multivariate control relationships;

[0043] By integrating the aforementioned multi-regulatory relationships, a multi-gene regulatory module is identified and formed.

[0044] According to another aspect of this application, a MICP functional gene regulatory network analysis device is also disclosed, the device comprising:

[0045] The binary expression matrix determination module is used to binarize the raw expression data of MICP functional genes from multiple samples to obtain a binary expression matrix. In the binary expression matrix, the rows represent genes, the columns represent samples, and the matrix elements are used to characterize the expression status of the corresponding gene in the corresponding sample.

[0046] A one-dimensional gene sequence determination module is used to map each gene in the binary expression matrix into a one-dimensional gene sequence based on its two-dimensional feature vector through a space-filling curve, wherein genes that are close in position in the gene feature distribution map remain adjacent in the one-dimensional gene sequence.

[0047] The quantum state basis vector determination module is used to determine the binary state sequence of each sample based on the one-dimensional gene sequence, and to encode each binary state sequence of the gene into a quantum state basis vector;

[0048] The system quantum state vector construction module is used to construct a system quantum state vector representing the entire gene system based on all quantum state basis vectors of all samples.

[0049] The matrix product state representation determination module is used to decompose and compress the quantum state vector of the system using a matrix product state tensor network to obtain the matrix product state representation of the quantum state vector;

[0050] The quantum mutual information determination module is used to calculate the binary quantum mutual information between any pair of genes and the multivariate quantum mutual information between multiple genes based on the matrix product state representation.

[0051] A gene regulation network construction module is used to construct a gene regulation backbone network based on the binary quantum mutual information, and to identify a multi-gene regulation module based on the multi-quantum mutual information.

[0052] A comprehensive regulatory network construction module is used to construct a comprehensive regulatory network based on the gene regulatory backbone network and the multi-gene regulatory module.

[0053] The analysis and identification module is used to input the MICP functional gene expression data of the sample to be tested into the integrated regulatory network for analysis, so as to identify the core genes or functional modules in the network that play a role in MICP efficiency, so as to determine the target sites for strain modification based on the identified core genes or functional modules.

[0054] According to another aspect of this application, an electronic device is also disclosed, the electronic device including a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to cause the electronic device to perform the various steps of the MICP functional gene regulatory network analysis method as described in any of the preceding claims.

[0055] According to another aspect of this application, a computer-readable storage medium is also disclosed, on which instructions are stored, which, when executed by a processor, implement the various steps of the MICP functional gene regulatory network analysis method as described in any of the preceding claims.

[0056] The present invention includes, but is not limited to, the following beneficial effects: (1) By introducing tensor networks and quantum mutual information methods, the present invention can capture high-order regulatory patterns among multiple genes under the premise of ensuring computational feasibility, overcome the limitations of traditional methods, and provide theoretical support for in-depth understanding of the molecular mechanism of MIP, optimization of strain modification and engineering application; (2) The present invention relies on the tensor network framework to break through the core bottleneck of the regulation analysis of key functional genes in microbial induced calcium carbonate precipitation (MICP), and can capture high-order gene regulatory modules that traditional methods cannot identify, and locate core functional genes; (3) The two-dimensional coordinates obtained after PCA dimensionality reduction are the ideal input for subsequent mapping of genes into one-dimensional sequences using space-filling curves, because PCA retains the proximity relationship between data points to the greatest extent during dimensionality reduction. After one-dimensional serialization, functionally related genes can still be as close as possible, thereby encoding biological prior knowledge into the topological structure of the subsequent model. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0058] Figure 1 This is a flowchart of the MICP functional gene regulatory network analysis method according to an embodiment of this application;

[0059] Figure 2 This is a flowchart illustrating the determination of the binary representation matrix according to an embodiment of this application;

[0060] Figure 3 This is a flowchart illustrating the determination of a one-dimensional gene sequence according to an embodiment of this application;

[0061] Figure 4 This is a flowchart illustrating the construction of quantum state basis vectors according to an embodiment of this application;

[0062] Figure 5 This is a flowchart illustrating the construction of a gene regulatory network according to an embodiment of this application;

[0063] Figure 6 This is a structural block diagram of the MICP functional gene regulatory network analysis device according to an embodiment of this application;

[0064] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0065] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0066] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Figure 1 This is a flowchart of the MICP functional gene regulatory network analysis method according to an embodiment of this application. (See attached document.) Figure 1 It includes the following steps:

[0067] S100. The original expression data of MICP functional genes from multiple samples are binarized to obtain a binary expression matrix.

[0068] Understandably, MICP is a microbial mineralization process with broad application prospects, involving fields such as concrete crack repair, soil solidification, environmental remediation, and carbon dioxide sequestration.

[0069] Specifically, the purpose of this step is to acquire and organize the necessary gene expression data for key functional genes involved in microbial induced calcium carbonate precipitation (MICP), and to transform the raw experimental measurement data into a form suitable for subsequent tensor network modeling and regulatory network inference through data processing methods. Since the MICP process typically involves multiple functional genes (such as the urease gene ureC, carbonic anhydrase gene, and calcium ion transport-related genes), their expression data may originate from single-cell transcriptome sequencing, population transcriptome sequencing, or multi-omics joint analysis.

[0070] For example, this step first cleans and standardizes the collected data, and then performs binarization to ensure that downstream calculations can be carried out based on accurate and structured input data. The following is a description of data collection, cleaning, and standardization:

[0071] 1) Data Collection

[0072] First, it is necessary to collect expression data of key functional genes from MICP-related microorganisms (such as strains of Bacillus, Pseudomonas, and Alcaligenes known to have calcium carbonate precipitation capabilities). Data sources may include, but are not limited to:

[0073] Single-cell transcriptome sequencing data (scRNA-seq): used to capture the heterogeneity of gene expression in different cells during MICP, which helps to identify key regulatory genes that play a central role in the population.

[0074] Population transcriptome (RNA-seq) data: used to reflect gene expression patterns at the overall level, facilitating the analysis of gene regulatory trends of the MICP process under different environmental conditions.

[0075] Public databases and experimental data: such as expression profiles of relevant strains in the NCBI GEO database, or MICP-related samples collected by the research team based on specific environments (such as soil, concrete cracks, and marine sediments).

[0076] During data collection, the focus was on genes directly related to carbonate precipitation, such as the ureABC gene cluster encoding urease, auxiliary genes involved in ammonia production, carbonic anhydrase genes promoting bicarbonate ion production, and membrane protein genes related to calcium ion uptake and transport. These genes collectively determine the mineralization capacity and precipitation efficiency of microorganisms in the MIP process.

[0077] 2) Data cleaning and standardization

[0078] Raw gene expression data often contains noise and missing values, therefore it must undergo rigorous cleaning and standardization. The specific process is as follows:

[0079] (a) Removal of low-quality data: This includes removing samples with insufficient sequencing coverage, too few short reads, or gene expression close to the background level, in order to avoid spurious correlations in subsequent modeling.

[0080] (b) Normalization: Methods such as TPM (Transcripts Per Million), FPKM (Fragments Per Kilobase per Million), or RPKM (Reads Per Kilobase of exon model per Million) are used to unify the expression data of different samples to a comparable scale, thereby eliminating the bias caused by different sequencing depths.

[0081] (c) Missing value imputation: If there is a missing expression of key genes in some samples, the missing values ​​can be imputed by interpolation, similar gene expression patterns, or based on statistical models to ensure the integrity of the matrix.

[0082] Furthermore, after the above processing, some noise in the original data is removed. To further incorporate gene expression information into the tensor network model, continuous expression levels need to be binarized to distinguish between active and inactive gene states. The specific method is as follows:

[0083] (A) Gaussian Mixture Model (GMM) Fitting: For the expression distribution of each gene, it is assumed that its expression pattern can be composed of two Gaussian distributions: one distribution corresponds to low expression (inactive) and the other corresponds to high expression (active). The model parameters, including the mean, variance, and mixture ratio, are estimated using the Expectation-Maximization (EM) algorithm.

[0084] (B) State partitioning: After the model is fitted, the gene expression value of each sample is assigned to the state with the highest posterior probability, thus obtaining the binarized result, where 1 represents gene activity and 0 represents gene inactivity.

[0085] (C) Robustness test: To avoid bias caused by underfitting, the Bayesian Information Criterion (BIC) can be used to evaluate the rationality of the model and, if necessary, compare the merits of the binary model and the ternary model.

[0086] The above processing yields a binary gene state matrix, where each element is either 0 or 1, representing the gene's activity in the corresponding sample. This result not only simplifies subsequent computational complexity but also helps focus on regulatory relationships rather than absolute expression values ​​in modeling. In the binary expression matrix, rows represent genes, columns represent samples, and matrix elements characterize the expression state of the corresponding gene in the corresponding sample.

[0087] Furthermore, after binarization, the results can be transformed into an input format suitable for tensor network modeling. For example, the matrix can first be structured to generate a binary matrix of genes × samples. Then, the gene state vector for each sample can be encoded into a sequence of 0s and 1s, which will be mapped to quantum state vectors in subsequent steps. Further, the processed data can be stored in a standard format (such as CSV or HDF5) and backed up on different storage media for later use in tensor decomposition and network inference.

[0088] Based on the above steps, the original MICP-related gene expression data were collected, cleaned, standardized, and binarized, transforming them into an input matrix suitable for quantum state modeling. This step lays the foundation for subsequent dimensionality compression, tensor network modeling, and the inference of gene regulatory relationships.

[0089] Furthermore, after completing the data preprocessing and binarization described above, a high-dimensional gene expression matrix is ​​obtained. Since the MICP process involves a large number of genes, directly performing tensor network modeling in high-dimensional space would lead to excessively high computational complexity and make it difficult to maintain the biological proximity relationships between genes. Therefore, the high-dimensional gene expression matrix obtained above needs to be dimensionality-reduced to obtain the final binary expression matrix.

[0090] For example, Figure 2 A flowchart for determining the binary representation matrix is ​​provided; please refer to [link / reference]. Figure 2 Principal component analysis was used to reduce the dimensionality of the original expression data to obtain a binary expression matrix, including the following steps:

[0091] S200. Perform zero-mean normalization on the original expression data and calculate the covariance matrix of the matrix after zero-mean normalization.

[0092] S202. Perform eigenvalue decomposition on the covariance matrix and select the first target principal components as low-dimensional representations.

[0093] S204. Project the data points of each gene in the covariance matrix onto a two-dimensional plane composed of the first target principal components to obtain the two-dimensional coordinates of each gene.

[0094] Specifically, we can define the matrix that needs to be reduced in dimensionality as follows: ,in For the number of genes, This represents the number of samples. Further analysis of the matrix... After zero-mean normalization, the covariance matrix is ​​calculated to capture the correlation of gene expression patterns. Further, eigenvalue decomposition is performed on the covariance matrix, and the first target principal components (e.g., the first two, three, etc.) are selected as low-dimensional representations. Then, the data points of each gene are projected onto a two-dimensional plane composed of the first two principal components, thus obtaining the two-dimensional coordinates of each gene. Zero-mean subtraction refers to subtracting the average value of the data in each row (i.e., each gene) of matrix Z from the average value of the data in that row (the expression value of that gene in all samples). After performing this operation on all rows (all genes) of matrix Z, the zero-mean matrix is ​​obtained.

[0095] The result of this process is that the originally high-dimensional and complex gene expression patterns are compressed into point clouds on a two-dimensional plane, which reduces the computational burden while preserving most of the major expression variation features. For key MICP genes (such as urease, carbonic anhydrase, and calcium transport-related genes), this process helps to identify their relative positions in the overall expression space.

[0096] S102. For each gene in the binary expression matrix, map it to a one-dimensional gene sequence through a space-filling curve based on its two-dimensional feature vector.

[0097] Specifically, Figure 3 This is a flowchart illustrating the determination of a one-dimensional gene sequence according to an embodiment of this application. The flowchart is used to explain step S102. (See attached document.) Figure 3 It includes the following steps:

[0098] S300. Normalize the two-dimensional coordinates of each gene and map them to the target square region.

[0099] For example, the target square region can be [0,1]×[0,1].

[0100] S302. Calculate the one-dimensional index value of each gene on the Hilbert space curve based on its two-dimensional coordinates.

[0101] Hilbert space curves are a type of space-filling curve that can continuously and recursively traverse a two-dimensional plane like a line. During this process, points that are close in position on the two-dimensional plane will also be close in position (index value) on this curve.

[0102] S304. Sort all genes in ascending order of index value to obtain a one-dimensional permutation sequence.

[0103] S104. Based on the one-dimensional gene sequence, determine the binary state sequence of each sample and encode each binary state sequence of a quantum state as a quantum state basis vector.

[0104] Specifically, the method for determining the quantum state basis vectors in step S104 is as follows: Figure 4 As shown, it includes the following steps:

[0105] S400. For each sample, according to the gene order defined by the one-dimensional gene sequence, obtain the binary expression state of each gene in sequence to form the binary state sequence of the sample.

[0106] Specifically, it forms a binary state sequence consisting of 0s and 1s, for example: [1,0,1,1,0,…].

[0107] S402. Map each binary bit in the binary state sequence to the corresponding computational ground state of a qubit.

[0108] Specifically, the 0 in the binary state sequence is mapped to 1 is mapped to 0 indicates that the gene is in an inactive state, and 1 indicates that the gene is in an active state.

[0109] S404. Map the binary state sequence of each sample to multiple qubits, calculate the direct product of the ground states, and obtain the quantum state basis vector of each sample.

[0110] Specifically, each sample generates a string of quantum letters, for example... These are combined through direct product operations. Direct product operations are mathematical operations in quantum mechanics that describe the joint state of multiple independent systems, and are symbolized as follows: For the example above, the combined state is: In quantum information, this state, formed by the direct product of multiple qubit ground states, is called a product ground state. It is usually abbreviated as... .and then This is the final quantum state basis vector that represents the specific gene expression pattern of the sample.

[0111] S106. Construct a system quantum state vector representing the entire gene system based on all quantum state basis vectors of all samples.

[0112] Specifically, statistical analysis is performed on each quantum state basis vector in all samples. The frequency of occurrence is used to obtain its probability distribution. Based on probability distribution Construct a system quantum state vector representing the entire gene system; whereby the system quantum state vector is represented by the following formula:

[0113] ;

[0114] In the formula, These are quantum state basis vectors. It is the probability amplitude of the quantum state basis vector.

[0115] S108. The system quantum state vector is decomposed and compressed using a matrix product state tensor network to obtain the matrix product state representation of the quantum state vector.

[0116] Specifically, the quantum state vector of the system is decomposed into a product sequence of tensors using the matrix product state method. Each tensor corresponds to a gene, and the bond dimension between adjacent tensors is used to characterize the correlation strength between genes. Then, a chain matrix product state structure composed of N tensors is established as the matrix product state representation of the quantum state vector.

[0117] S110. Based on matrix product state representation, calculate the binary quantum mutual information between arbitrary gene pairs and the multivariate quantum mutual information between multiple genes.

[0118] Specifically, after completing the above steps, the matrix product state (MPS) representation of key functional genes in MICP has been obtained. This representation not only compresses high-dimensional data but also preserves multi-level correlation information between genes. However, tensor network structures alone are insufficient to reveal the specific gene regulatory networks. To further infer the regulatory relationships between genes, it is necessary to utilize quantum mutual information (QMI) from quantum information theory to quantify the dependence strength between gene pairs and verify its significance through statistical methods.

[0119] Quantum mutual information is a metric used to measure the overall correlation between two subsystems, encompassing the combined effects of classical and quantum correlations. In gene regulatory network inference, it is used to measure the degree of dependence between the expression states of two genes.

[0120] The quantum mutual information of two gene-corresponding subsystems A and B is defined as follows:

[0121] ;

[0122] in, For von Neumann entropy, These are the reduced density matrices for single genes. This is the joint density matrix of the two genes.

[0123] In MPS representation, the reduced density matrix of a single gene or gene pair can be efficiently calculated through tensor shrinkage, thereby obtaining the desired entropy value and QMI. A larger QMI value indicates a strong dependency between the two genes in their expression state, possibly indicating direct or indirect regulatory effects; a smaller QMI value indicates that the genes are independent of each other or have a weak dependency.

[0124] To measure higher-order relationships among the three genes, ternary quantum mutual information (TQMI) is introduced.

[0125] For three genes Its ternary QMI can be expressed as:

[0126] ;

[0127] in, The von Neumann entropy represents the corresponding gene or gene combination.

[0128] like This indicates that there is a synergistic relationship among the three, meaning that the combined effect of the three genes is stronger than any two of them combined.

[0129] like This indicates a redundant relationship, meaning that the function of a certain gene is partially replaced by two other genes.

[0130] S112, Constructing a gene regulatory backbone network based on binary quantum mutual information, and identifying multi-gene regulatory modules based on multi-dimensional quantum mutual information.

[0131] Specifically, Figure 5 This is a flowchart illustrating the construction of a gene regulatory network according to an embodiment of this application. The flowchart is used to exemplarily illustrate step S112. (See attached document.) Figure 5 It includes the following steps:

[0132] S500: Based on the binary quantum mutual information between all gene pairs, a symmetric gene dependency matrix is ​​constructed.

[0133] Specifically, for binary quantum mutual information, after calculating the quantum mutual information of all gene pairs, the results can be organized into a symmetric gene dependency matrix. , where matrix elements Indicates gene and The quantum mutual information between genes; the diagonal elements are empty or set to zero because the mutual information between genes and themselves has no practical significance; the matrix reflects all possible pairwise dependencies in the entire gene set.

[0134] S502. Perform a permutation test on each binary quantum mutual information in the symmetric gene dependence matrix to screen out regulatory gene pairs.

[0135] It is understandable that the symmetric gene dependency matrix obtained by the above steps is an unscreened candidate regulatory network, which may contain both real regulatory relationships and random noise or accidental co-occurrences.

[0136] Step 502 involves performing a permutation test on each binary quantum mutual information in the symmetric gene dependency matrix to screen for regulatory gene pairs. Specifically, by randomly shuffling the labels of genes in the sample, their true correlations are disrupted, thereby generating a quantum mutual information distribution under the dependency-free assumption. This permutation operation is repeated several times (e.g., 1000 times) for each gene pair, resulting in a set of virtual quantum mutual information values ​​as the null distribution. The observed quantum mutual information is compared with the null distribution, and the p-values ​​of the right and left tails are calculated. If the right tail is significant, it indicates that the dependency between the two genes is stronger than in the random case, and a positive regulatory relationship may exist. If the left tail is significant, it indicates that the dependency between the two genes is weaker than in the random case, and an active separation or mutual exclusion mechanism may exist. Specifically, a statistical significance level (e.g., p < 0.05) can be set to screen significant gene pairs as regulatory gene pairs. The right tail can be understood as being far from the right extreme of the random distribution, and the left tail as being far from the left extreme of the random distribution.

[0137] S504. Construct a gene regulatory backbone network based on regulatory gene pairs.

[0138] Specifically, each gene is treated as a node in the network. If the quantum mutual information value between two genes is statistically significant, an edge is added between them, indicating the existence of a potential regulatory relationship. Optionally, the magnitude of the quantum mutual information is used as the weight of the edge to measure the regulatory strength. Furthermore, graph theory visualization tools (such as Cytoscape or Gephi) can be used to draw the regulatory network, visually displaying the interaction structure between genes.

[0139] For example, if a significant and high QMI is found between the urease gene (ureC) and the carbonic anhydrase gene during the MICP process, it can be inferred that they have a close synergistic effect in the regulation process, jointly promoting the formation of calcium carbonate precipitation.

[0140] This step quantifies the dependencies between genes using quantum mutual information methods and eliminates random noise through permutation tests, ultimately constructing a regulatory network of key functional genes in MICP. This network accurately reflects the potential regulatory relationships between genes.

[0141] S506. Perform permutation tests on all multi-element quantum mutual information values ​​to screen out multi-element control relationships.

[0142] S508 integrates multiple regulatory relationships, identifies and forms multi-gene regulatory modules.

[0143] Specifically, similar to binary quantum mutual information, the calculation results of higher-order quantum mutual information also need to be verified for significance using statistical methods. A zero distribution of ternary quantum mutual information can be generated by randomly shuffling gene tags in the sample. This zero distribution can be repeated several times (e.g., 1000 times) to obtain a virtual distribution of ternary quantum mutual information. The difference between the observed ternary quantum mutual information and the zero distribution can be calculated to obtain the p-value, and significant three-gene regulatory modules can be screened out.

[0144] For example, if a ternary QMI shows a significant positive synergistic relationship among the urease gene (ureC), carbonic anhydrase gene, and calcium ion transporter gene, it can be inferred that these three may together constitute the core regulatory unit in the MCP.

[0145] After obtaining significant multi-gene regulatory relationships, it is necessary to further incorporate them into a network structure for modular analysis, including:

[0146] 1. Module Identification: Significant ternary and multivariate regulatory relationships are integrated into modules through community detection or clustering algorithms. For example, the ureC–carbonic anhydrase–calcium ion transport gene may form a mineralization functional module.

[0147] 2. Biological Function Explanation: Based on existing MIP research results, analyze the role of the modules in metabolic pathways. For example, urease hydrolyzes urea to produce carbonate ions, carbonic anhydrase catalyzes the reaction of carbon dioxide and water to produce bicarbonate ions, and calcium transport genes are responsible for introducing calcium ions into the surrounding cellular environment. These three work synergistically to form calcium carbonate precipitate.

[0148] 3. Exploration of hierarchical structure: In addition to focusing on three-gene modules, we can recursively extend to four-gene or larger regulatory units to gradually reveal the hierarchical organization of the MIP regulatory network.

[0149] Through the steps described above, the analysis of regulatory relationships is extended from binary to multivariate, and ternary and multivariate quantum mutual information is used to reveal potential cooperative, redundant, and mutually exclusive mechanisms in the MICP process. This not only enriches our understanding of gene regulatory networks but also provides new evidence for the identification and application of functional modules.

[0150] S114. Construct a comprehensive regulatory network based on the gene regulatory backbone network and multi-gene regulatory modules.

[0151] Specifically, based on the steps described above, pairwise regulatory relationships between key functional genes in MICP and higher-order regulatory modules of multiple genes were obtained. However, these relationships still exist in the form of matrices or discrete modules, and have not yet formed an overall framework that can be intuitively understood and applied. In order to enable the results to be directly used for theoretical analysis and practical engineering applications, these relationships need to be systematically integrated into a complete gene regulatory network. Therefore, the purpose of step 114 includes: integrating binary and multi-regulatory relationships into a unified network model, visualizing the network, performing structural analysis, and interpreting its functions.

[0152] Specifically, the construction of the integrated control network in this step is based on the following steps:

[0153] 1. Node definition: All MICP-related functional genes are used as nodes in the network, including the urease gene cluster (ureABC), carbonic anhydrase gene, calcium transport gene, and other regulatory genes identified in previous analyses.

[0154] 2. Introduction of edges:

[0155] For binary relations: if the QMI between two genes is significant, then add an edge between the two nodes.

[0156] For multivariate relationships: For significant ternary or higher-order modules, hyperedges can be introduced or virtual nodes can be transformed to reflect the joint regulatory relationships of multiple genes.

[0157] 3. Weight assignment: The binary quantum mutual information value or the ternary quantum mutual information value is used as the weight of the edge to quantitatively represent the strength of the control relationship.

[0158] S116. Input the MICP functional gene expression data of the sample to be tested into the integrated regulatory network for analysis, so as to identify the core genes or functional modules that play a role in MICP efficiency in the network, so as to determine the target of strain modification based on the identified core genes or functional modules.

[0159] Specifically, after the MICP functional gene expression data of the test samples are input into the integrated regulatory network, the degree centrality, betweenness centrality, and proximity centrality of the nodes are calculated based on the constructed MICP key functional gene regulatory network to identify genes that play a core role in the MICP regulatory network. Core genes and functional modules are identified by jointly analyzing network topological features and higher-order quantum information metrics. Specifically, degree centrality and betweenness centrality algorithms are used to locate key hub nodes and bridge nodes that maintain connectivity and control metabolic flow in the network. Simultaneously, a community detection algorithm combined with the positive and negative values ​​of ternary quantum mutual information (TQMI) is applied to cluster gene clusters with significant synergistic effects (TQMI>0) into independent functional modules, thereby accurately identifying the synergistic regulatory units that dominate the calcium carbonate precipitation process. For example, if ureC is central in multiple modules, it indicates that it is a key driver gene in the mineralization process. Furthermore, network partitioning algorithms (such as the Louvain method) are used to identify functional modules, revealing the division of labor and cooperation among different gene groups. Furthermore, analyzing the shortest paths or regulatory links between genes can explore potential regulatory cascade effects in the calcium carbonate precipitation process. Identifying core regulatory genes within the network can provide targets for genetic engineering, such as enhancing the synergistic module between urease and calcium transport genes to improve calcium carbonate precipitation efficiency. In practical MICP applications (such as concrete remediation, soil solidification, and carbon dioxide sequestration), culture conditions or exogenous additives can be adjusted based on the regulatory network to promote the expression and regulation of key genes. In contaminated site remediation, the regulatory network can be used to predict environmental factors (such as...). The effect of concentration on gene regulation provides theoretical support for optimizing microbial community function.

[0160] Furthermore, based on the above network analysis results, core genes with high centrality in the synergistic functional modules were selected as overexpression targets to enhance mineralization function, or upstream repressor factors with significant negative dependence on key pathway genes were identified as knockout targets, thereby optimizing the mineralization efficiency of the strain in a targeted manner.

[0161] Through the above steps, the binary and multivariate gene regulatory relationships obtained earlier are integrated into a complete MICP key functional gene regulatory network, which is then visualized, structurally analyzed, and biologically interpreted. This network not only reveals the core regulatory mechanisms in the calcium carbonate precipitation process but also provides practical guidance for strain modification, process optimization, and environmental applications.

[0162] Furthermore, Figure 6 This is a structural block diagram of the MICP functional gene regulatory network analysis device according to an embodiment of the application, such as... Figure 6 As shown, the device includes:

[0163] The binary expression matrix determination module is used to binarize the raw expression data of MICP functional genes from multiple samples to obtain a binary expression matrix. In the binary expression matrix, the rows represent genes, the columns represent samples, and the matrix elements are used to characterize the expression status of the corresponding gene in the corresponding sample.

[0164] The one-dimensional gene sequence determination module is used to map each gene in the binary expression matrix into a one-dimensional gene sequence based on its two-dimensional feature vector through a space-filling curve. In this module, genes that are close in position in the gene feature distribution map are kept adjacent in the one-dimensional gene sequence.

[0165] The quantum state basis vector determination module is used to determine the binary state sequence of each sample based on the one-dimensional gene sequence, and encode each binary state sequence of the gene into a quantum state basis vector;

[0166] The system quantum state vector construction module is used to construct a system quantum state vector representing the entire gene system based on all quantum state basis vectors of all samples.

[0167] The matrix product state representation determination module is used to decompose and compress the quantum state vector of the system using a matrix product state tensor network to obtain the matrix product state representation of the quantum state vector;

[0168] The quantum mutual information determination module is used to calculate the binary quantum mutual information between any pair of genes and the multivariate quantum mutual information between multiple genes based on the matrix product state representation.

[0169] A gene regulation network construction module is used to construct a gene regulation backbone network based on binary quantum mutual information, and a multi-gene regulation module is used to identify multiple genes based on multi-dimensional quantum mutual information.

[0170] A comprehensive regulatory network construction module is used to construct a comprehensive regulatory network based on the gene regulatory backbone network and multi-gene regulatory modules.

[0171] The analysis and identification module is used to input the MICP functional gene expression data of the sample to be tested into the integrated regulatory network for analysis, so as to identify the core genes or functional modules in the network that play a role in MICP efficiency, so as to determine the target sites for strain modification based on the identified core genes or functional modules.

[0172] The application of the relevant modules of the device in this example can be referred to the relevant introduction of the method principle above, and will not be repeated here.

[0173] above Figure 6 The MICP functional gene regulatory network analysis device in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The electronic device in this embodiment of the invention will be described in detail from the perspective of hardware processing.

[0174] Figure 7 This is a schematic diagram of the structure of an electronic device 700 provided in an embodiment of the present invention. The electronic device 700 can vary significantly due to differences in configuration or performance, and may include one or more central processing units (CPUs) 710 (e.g., one or more processors) and a memory 720, and one or more storage media 730 (e.g., one or more mass storage devices) for storing application programs 733 or data 732. The memory 720 and storage media 730 can be temporary or persistent storage. The program stored in the storage media 730 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the electronic device 700. Furthermore, the processor 710 may be configured to communicate with the storage media 730 and execute the series of instruction operations in the storage media 730 on the electronic device 700.

[0175] Electronic device 700 may also include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input / output interfaces 760, and / or one or more operating systems 731, such as Windows Server, MacOSX, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 7 The illustrated electronic device structure does not constitute a limitation on electronic devices and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0176] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of any of the above-described MICP functional gene regulatory network parsing methods.

[0177] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0178] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0179] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for analyzing the MICP functional gene regulatory network, characterized in that, The method includes: The raw expression data of MICP functional genes from multiple samples are binarized to obtain a binary expression matrix, wherein the rows of the binary expression matrix represent genes, the columns represent samples, and the matrix elements are used to characterize the expression status of the corresponding gene in the corresponding sample. The two-dimensional coordinates of each gene are normalized and mapped to the target square region; Calculate the one-dimensional index value of each gene on the Hilbert space curve based on its two-dimensional coordinates. All genes are sorted in ascending order of their index values ​​to obtain a one-dimensional gene sequence. Genes that are close in position in the gene feature distribution map are kept adjacent in the one-dimensional gene sequence. Based on the one-dimensional gene sequence, determine the binary state sequence of the genes for each sample, and encode each binary state sequence of the genes into a quantum state basis vector; Statistical analysis of each quantum state basis vector in all samples The frequency of occurrence is used to obtain its probability distribution. ; Based on the probability distribution Construct a system quantum state vector representing the entire gene system, where the system quantum state vector is based on The formula represents that, in the formula, These are quantum state basis vectors. It is the probability amplitude of the quantum state basis vector; The quantum state vector of the system is decomposed into a product sequence of tensors using the matrix product state method, where each tensor corresponds to a gene, and the bond dimension between adjacent tensors is used to characterize the correlation strength between genes. Establish a chain-like matrix product state structure consisting of N tensors connected in series, as the matrix product state representation of the quantum state vector; Based on the matrix product state representation, calculate the binary quantum mutual information between any pair of genes and the multi-dimensional quantum mutual information between multiple genes; A symmetric gene dependency matrix is ​​constructed based on the binary quantum mutual information between all gene pairs. The permutation test is performed on each binary quantum mutual information in the symmetric gene dependence matrix to screen out regulatory gene pairs; Based on the aforementioned regulatory gene pairs, a gene regulatory backbone network is constructed, wherein each gene serves as a network node, and gene pairs are connected by edges. Permutation tests are performed on all multi-dimensional quantum mutual information to screen out multi-dimensional control relationships; By integrating the aforementioned multi-regulatory relationships, a multi-gene regulatory module is identified and formed. A comprehensive regulatory network is constructed based on the gene regulatory backbone network and the multi-gene regulatory module; The MICP functional gene expression data of the sample to be tested are input into the integrated regulatory network for analysis to identify the core genes or functional modules in the network that play a role in MICP efficiency, so as to determine the target sites for strain modification based on the identified core genes or functional modules.

2. The method for analyzing the MICP functional gene regulatory network according to claim 1, characterized in that, The binarization of the raw expression data of MICP functional genes from multiple samples to obtain a binary expression matrix includes: The original expression data is subjected to dimensionality reduction processing based on principal component analysis, including: The original expression data is zero-mean normalized, and the covariance matrix of the zero-mean normalized matrix is ​​calculated. The covariance matrix is ​​subjected to eigenvalue decomposition, and the first target principal components are selected as low-dimensional representations. The data points of each gene in the covariance matrix are projected onto a two-dimensional plane composed of the first target principal components to obtain the two-dimensional coordinates of each gene.

3. The method for analyzing the MICP functional gene regulatory network according to claim 1, characterized in that, The process involves determining the binary state sequence of genes for each sample based on the one-dimensional gene sequence, and encoding each binary state sequence of genes into a quantum state basis vector. For each sample, the binary expression state of each gene is obtained sequentially according to the gene order defined by the one-dimensional gene sequence, forming the binary state sequence of the sample. Each binary bit in the binary state sequence is mapped to the corresponding computational ground state of a qubit, where 0 is mapped to... 1 is mapped to 0 indicates that the gene is in an inactive state, and 1 indicates that the gene is active. The binary state sequence of each sample is mapped to multiple qubits, and the direct product of the ground states is calculated to obtain the quantum state basis vector of each sample.

4. A device for analyzing the MICP functional gene regulatory network, characterized in that, The device includes: The binary expression matrix determination module is used to binarize the raw expression data of MICP functional genes from multiple samples to obtain a binary expression matrix. In the binary expression matrix, the rows represent genes, the columns represent samples, and the matrix elements are used to characterize the expression status of the corresponding gene in the corresponding sample. The one-dimensional gene sequence determination module is used to normalize the two-dimensional coordinates of each gene and map them to the target square region; calculate the one-dimensional index value of each gene on the Hilbert space curve based on the two-dimensional coordinates of each gene; sort all genes in ascending order of index value to obtain a one-dimensional gene sequence, wherein genes that are close in position in the gene feature distribution map are kept adjacent in the one-dimensional gene sequence. The quantum state basis vector determination module is used to determine the binary state sequence of each sample based on the one-dimensional gene sequence, and to encode each binary state sequence of the gene into a quantum state basis vector; The system quantum state vector construction module is used to statistically analyze the basis vectors of each quantum state in all samples. The frequency of occurrence is used to obtain its probability distribution. Based on the probability distribution Construct a system quantum state vector representing the entire gene system, where the system quantum state vector is based on The formula represents that, in the formula, These are quantum state basis vectors. It is the probability amplitude of the quantum state basis vector; The matrix product state representation determination module is used to decompose the quantum state vector of the system into a product sequence of tensors through the matrix product state method, and establish a chain matrix product state structure composed of N tensors connected in series as the matrix product state representation of the quantum state vector. Each tensor corresponds to a gene, and the bond dimension between adjacent tensors is used to characterize the correlation strength between genes. The quantum mutual information determination module is used to calculate the binary quantum mutual information between any pair of genes and the multivariate quantum mutual information between multiple genes based on the matrix product state representation. A gene regulation network construction module is used to construct a symmetric gene dependency matrix based on the binary quantum mutual information between all gene pairs, perform permutation tests on each binary quantum mutual information in the symmetric gene dependency matrix to screen out regulatory gene pairs, construct a gene regulation backbone network based on the regulatory gene pairs, perform permutation tests on all multivariate quantum mutual information to screen out multivariate regulatory relationships, integrate the multivariate regulatory relationships, identify and form a multigene regulation module, wherein each gene is a network node, and gene pairs are connected by edges; A comprehensive regulatory network construction module is used to construct a comprehensive regulatory network based on the gene regulatory backbone network and the multi-gene regulatory module. The analysis and identification module is used to input the MICP functional gene expression data of the sample to be tested into the integrated regulatory network for analysis, so as to identify the core genes or functional modules in the network that play a role in MICP efficiency, so as to determine the target sites for strain modification based on the identified core genes or functional modules.

5. An electronic device, characterized in that, The electronic device includes a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to cause the electronic device to perform the steps of the MICP functional gene regulatory network analysis method as described in any one of claims 1-3.

6. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement each step of the MICP functional gene regulatory network parsing method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Digital modeling method, medium and equipment for enhancing connectivity of low-porosity rock core

    CN120894516A

  • Full-length gene sequence modeling method and system based on neural network

    CN121306244A