A method for gene expression network analysis based on the minimum entropy algorithm

The method uses the minimum entropy algorithm to evaluate gene importance and construct gene expression networks, improving clustering accuracy by selecting strongly correlated genes and enabling network visualization.

CN116884487BActive Publication Date: 2025-07-15HAINAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310912377.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-07-15
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

The prior art cannot directly select highly correlated genes from the calculation process to map gene expression networks.

Method used

The minimum entropy algorithm is used to evaluate the importance of genes, and combined with Pearson correlation coefficient and Infomap multi-level network clustering model, a gene expression network is constructed, and the correlation matrix between genes is obtained and clustered analysis is performed.

Benefits of technology

It improves the accuracy of gene clustering results, can effectively select highly correlated genes to build gene expression networks, and realizes visualization of gene expression networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116884487B_ABST
    Figure CN116884487B_ABST
Patent Text Reader

Abstract

The present invention provides a method for analyzing gene expression networks based on the minimum entropy algorithm, comprising the following steps: data collection and preprocessing; obtaining gene correlations; converting data formats; establishing a gene expression network based on a gene correlation matrix; clustering the gene expression network using an Infomap multilevel network clustering model, identifying the relationships between all genes by means of random walks, dividing genes encoded with the same information into the same group, and visualizing the network for each obtained group. This application constructs a gene expression network by selecting genes with strong correlations based on the minimum entropy algorithm, solves the problem that the prior art cannot directly select genes with strong correlations from the calculation process to draw an expression network, and selects the most representative genes as clustering features based on evaluating the importance of each gene for the clustering task in the minimum entropy algorithm, thereby improving the accuracy of the clustering results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene expression analysis, and particularly to a method for analyzing gene expression networks based on the minimum entropy algorithm. Background Art

[0002] Gene expression analysis has a wide range of applications in biological research and can be used to study problems in different species, different biological processes, different diseases, etc. The methods for analyzing gene expression data can be divided into the following types: differential analysis, clustering analysis, biological network analysis, functional annotation analysis, etc. These methods are often combined and used comprehensively to achieve more accurate and comprehensive results. The minimum entropy algorithm is an algorithm used for decision tree construction, aiming to select the best features for splitting by minimizing the value of entropy. Compared with other clustering algorithms, the advantage of the minimum entropy algorithm is that it does not require prior knowledge, has stronger robustness to noise and missing data, and is easy to interpret the results. In addition, the minimum entropy algorithm is efficient in processing large-scale data, which makes it a powerful tool for analyzing large data sets.

[0003] However, in the process of using existing software, it is impossible to directly select genes with strong correlation from the calculation process to draw an expression network. By designing and improving relevant parameters in the present invention, users can better adjust their clustering results. The final data and pictures that can be obtained include the total gene correlation file and the correlation files of each module, the expression networks of all genes and the expression network diagrams of individual modules, sample clustering diagrams, K-value evaluation diagrams, gene and inter-gene correlation heatmaps. Based on the obtained inter-gene relationships, a secondary network of gene libraries such as pathways can be constructed. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method for analyzing gene expression networks based on the minimum entropy algorithm. By obtaining the correlation between genes, a gene correlation matrix is obtained, and genes with strong correlation are selected based on the minimum entropy algorithm to construct a gene expression network, so as to solve the problem that in the prior art, it is impossible to directly select genes with strong correlation from the calculation process to draw an expression network.

[0005] To achieve the above invention purpose, the present invention provides a method for analyzing gene expression networks based on the minimum entropy algorithm, and the method includes the following steps:

[0006] S1. Data collection and preprocessing: Collect various types of expression data or chip data, perform preprocessing, and draw a sample correlation matrix to remove outlier samples;

[0007] S2. Gene - to - gene correlation acquisition: Use the Pearson correlation coefficient to calculate the correlation between all genes and construct a gene correlation matrix. Among them, in statistics, the Pearson correlation coefficient is used to measure the correlation (linear correlation) between two variables X and Y. The value range of the Pearson correlation coefficient is between - 1 and 1. The closer the value is to 1 or - 1, the stronger the correlation. The closer the value is to 0, the weaker or non - existent the correlation.

[0008] S3. Convert data format: Based on the fields fromNode, toNode, and weight, convert the gene correlation matrix into a DataFrame format, and perform label encoding on the fields fromNode and toNode.

[0009] S4. Establish a gene expression network based on the gene correlation matrix: Adopt the minimum entropy algorithm for clustering analysis and construct a gene expression network.

[0010] S5. Use the Infomap multi - level network clustering model to cluster the gene expression network, identify the relationships between all genes by means of random walks, divide the genes encoded with the same information into the same community, and perform network visualization on each obtained community.

[0011] It should be noted that the specific steps of S1 are as follows:

[0012] 1) Collect the dataset to be clustered, including RNA - sequencing data or microarray data. The dataset contains N samples, and each sample consists of d attributes, which can be represented as an N x d matrix. Represent all samples as a matrix, where each row of the matrix represents a sample vector.

[0013] 2) Calculate the distance between each sample: By calculating the Euclidean distance between each sample and other samples, obtain an N x N distance matrix. For two d - dimensional sample vectors X and Y, the Euclidean distance d(X,Y) is:

[0014] d(X,Y)=sqrt((x1 - y1)^2+(x2 - y2)^2+...+(xd - yd)^2)

[0015] where x1, x2,..., xd and y1, y2,..., yd are the values of vector X and vector Y in each dimension respectively.

[0016] 3) Select the number of clusters k, and then use the k - means algorithm to cluster the samples into k clusters.

[0017] 4) Observe whether there are outliers or abnormal values, including removing invalid data, standardizing, selecting feature genes, and plotting the sample correlation matrix.

[0018] It should be noted that the specific steps for converting the gene correlation matrix into a DataFrame format based on the fields fromNode, toNode, and weight in S3 and performing label encoding on the fields fromNode and toNode are as follows:

[0019] 1) Use the pandas library in Python to store the gene correlation matrix in a variable named corr_matrix, where corr_matrix is a gene correlation matrix;

[0020] 2) Import the pandas library: Use the DataFrame function of pandas to create an empty DataFrame array: Set the gene correlation matrix as the value of the DataFrame array, and the DataFrame array uses gene names as row and column indexes;

[0021] 3) Set the row and column indexes of the DataFrame to gene names or other corresponding identifiers.

[0022] It should be noted that the specific steps for performing clustering analysis using the minimum entropy algorithm and constructing a gene expression network in S4 are as follows:

[0023] 1) Select the corresponding parameters of the entropy function and distance metric;

[0024] 2) Minimize the information entropy between samples;

[0025] 3) Divide the data points into different clusters;

[0026] 4) Construct a gene expression network. In the gene expression network, each gene can be represented as a node, and the correlation between genes can be represented as an edge.

[0027] It should be noted that the community classification evaluation formula of the Infomap multi-level network clustering model in S5 is expressed as:

[0028] $Q=\sum_{i}(e_{ii}-a_{i}^2)+\frac{1}{2}\sum_{i\neq j}(e_{ij}-a_{i}a_{j})\delta(\sigma_i,\sigma_j)$

[0029] Among them, $Q$ represents the modularity for community division, which is used to measure the quality of the community structure. $e_{ij}$ represents the edge weight between node $i$ and node $j$. $a_{i}$ represents the degree of node $i$. $\delta(\sigma_i,\sigma_j)$ is a Kronecker Delta function, which has a value of 1 if node $i$ and node $j$ belong to the same module, and 0 otherwise. $\sigma_i$ represents the module to which node $i$ belongs. The goal of the Infomap algorithm is to find an optimal $\sigma_i$ allocation scheme to minimize $Q$.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] (1) Evaluate the importance of each gene for the clustering task based on the minimum entropy algorithm, and select the most representative genes as clustering features to improve the accuracy of the clustering results.

[0032] (2) Select genes with strong correlation based on the minimum entropy algorithm to construct a gene expression network, and solve the problem that the prior art cannot directly select genes with strong correlation from the calculation process to draw an expression network. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only the preferred embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 is a flowchart of a gene expression network analysis method based on the minimum entropy algorithm provided by an embodiment of the present invention.

[0035] Figure 2 is a sample clustering diagram of a gene expression network analysis method based on the minimum entropy algorithm provided by an embodiment of the present invention.

[0036] Figure 3 is a gene clustering diagram of a gene expression network analysis method based on the minimum entropy algorithm provided by an embodiment of the present invention.

[0037] Figure 4-a 、 Figure 4-b 、 Figure 4-c is a visualization schematic diagram of a gene expression network of a gene expression network analysis method based on the minimum entropy algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The principles and features of the present invention will be described below in conjunction with the accompanying drawings. The listed embodiments are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0039] Referring to Figure 1 , this embodiment provides a gene expression network analysis method based on the minimum entropy algorithm. The specific steps are as follows:

[0040] S1. Referring to Figure 2 , data collection and preprocessing: Collect the data set to be clustered, including RNA sequencing data or microarray data. The data set contains N samples, and each sample is composed of d attributes, which can be represented as an N x d matrix. Represent all samples as a matrix, where each row of the matrix represents a sample vector; Calculate the distance between each sample: By calculating the Euclidean distance between each sample and other samples, an N x N distance matrix is obtained; For two d-dimensional sample vectors X and Y, the Euclidean distance d(X,Y) is:

[0041] d(X,Y) = sqrt((x1 - y1)^2 + (x2 - y2)^2 +... + (xd - yd)^2)

[0042] where x1, x2,..., xd and y1, y2,..., yd are the values of vector X and vector Y in each dimension respectively;

[0043] Select the number of clusters k, and then use the k-means algorithm to cluster the samples into k clusters; Observe whether there are outliers or anomalies, including removing invalid data, standardizing, selecting feature genes, and plotting the sample correlation matrix.

[0044] S2. Obtaining gene-gene correlations: Use the Pearson correlation coefficient to calculate the correlations between all genes and construct a gene correlation matrix; Among them, the Pearson correlation coefficient is used in statistics to measure the correlation (linear correlation) between two variables X and Y. The value range of the Pearson correlation coefficient is between -1 and 1. The closer the value is to 1 or -1, the stronger the correlation. The closer the value is to 0, the weaker or non-existent the correlation;

[0045] S3. Convert the data format: Based on the fields fromNode, toNode, and weight, use the pandas library in Python to store the gene correlation matrix in a variable named corr_matrix, where corr_matrix is a gene correlation matrix; Import the pandas library: Use the DataFrame function of pandas to create an empty DataFrame array: Set the gene correlation matrix as the value of the DataFrame array, enabling the conversion of the gene correlation matrix to the DataFrame format, where the DataFrame array uses gene names as row and column indices.

[0046] S4. Establish a gene expression network based on the gene correlation matrix: Select the corresponding parameters of the entropy function and distance metric; Minimize the information entropy between samples; Divide the data points into different clusters; Construct a gene expression network, where in the gene expression network, each gene can be represented as a node and the correlation between genes can be represented as an edge.

[0047] S5. Refer to Figure 3 -4, use the Infomap multi-level network clustering model to cluster the gene expression network, identify the relationships between all genes using the random walk method, divide the genes encoded with the same information into the same community, and perform network visualization for each obtained community; The community classification evaluation formula of the Infomap multi-level network clustering model in S5 is expressed as:

[0048] $Q = \sum_{i}(e_{ii} - a_{i}^2) + \frac{1}{2}\sum_{i\neq j}(e_{ij} - a_{i}a_{j})\delta(\sigma_i, \sigma_j)$

[0049] where $Q$ represents the modularity of the community division, used to measure the quality of the community structure, $e_{ij}$ represents the edge weight between node $i$ and node $j$, $a_{i}$ represents the degree of node $i$, $\delta(\sigma_i, \sigma_j)$ is a Kronecker Delta function, if node $i$ and node $j$ belong to the same module, the value of this function is 1, otherwise it is 0, $\sigma_i$ represents the module to which node $i$ belongs, and the goal of the Infomap algorithm is to find an optimal $\sigma_i$ allocation scheme to minimize $Q$.

[0050] It should be noted that evaluating the importance of each gene for the clustering task based on the minimum entropy algorithm and selecting the most representative genes as clustering features improves the accuracy of the clustering results.

[0051] According to the visualization results of the gene expression community, users can analyze and explore the gene expression community structure, solving the problem that in the prior art, it is impossible to directly select genes with strong relevance from the calculation process to draw an expression network and realize visualization operations.

[0052] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for analyzing gene expression networks based on the minimum entropy algorithm, characterized in that The method includes the following steps: S1. Data collection and preprocessing: Collect various RNA sequencing data, perform preprocessing, and draw a sample correlation matrix to remove outlier samples. Specifically: 1) Collect the RNA sequencing data set to be clustered. The data set contains N samples, and each sample consists of d attributes, which can be represented as an N x d matrix. Represent all samples as a matrix, where each row of the matrix represents a sample vector; 2) Calculate the distance between each sample: By calculating the Euclidean distance between each sample and other samples, an N x N distance matrix is obtained; for two d-dimensional sample vectors X and Y, the Euclidean distance d(X,Y) is: d(X,Y) = sqrt((x1 - y1)^2 + (x2 - y2)^2 +... + (xd - yd)^2) where x1, x2,..., xd and y1, y2,..., yd are the values of vector X and vector Y in each dimension respectively; 3) Select the number of clusters k, and then use the k-means algorithm to cluster the samples into k clusters; 4) Observe whether there are outliers or anomalies, including removing invalid data, standardizing, selecting feature genes, and drawing a sample correlation matrix; S2. Obtain gene-gene correlations: Use the Pearson correlation coefficient to calculate the correlations between all genes and construct a gene correlation matrix; where the Pearson correlation coefficient is used in statistics to measure the correlation (linear correlation) between two variables X and Y. The value range of the Pearson correlation coefficient is between -1 and 1. The closer the value is to 1 or -1, the stronger the correlation; the closer the value is to 0, the weaker or non-existent the correlation; S3. Convert the data format: Based on the fields fromNode, toNode, and weight, convert the gene correlation matrix into a DataFrame format and perform label encoding on the fields fromNode and toNode; S4. Establish a gene expression network based on the gene correlation matrix: Adopt the minimum entropy algorithm for clustering analysis and construct a gene expression network. The specific steps are as follows: 1) Select the corresponding parameters of the entropy function and distance metric; 2) Minimize the information entropy between samples; 3) Divide the data points into different clusters; 4) Construct a gene expression network. In the gene expression network, each gene can be represented as a node, and the correlation between genes can be represented as an edge; S5. Use the Infomap multi-level network clustering model to cluster the gene expression network, identify the relationships between all genes by means of random walks, divide the genes encoded with the same information into the same community, and perform network visualization on each obtained community.

2. The gene expression network analysis method based on the minimum entropy algorithm according to claim 1, wherein The specific steps of converting the gene correlation matrix into a DataFrame format based on the fields fromNode, toNode, and weight and performing label encoding on the fields fromNode and toNode in S3 are as follows: 1) Use the pandas library in Python to store the gene correlation matrix in a variable named corr_matrix, where corr_matrix is a gene correlation matrix; 2) Import the pandas library: Use the DataFrame function of pandas to create an empty DataFrame array: Set the gene correlation matrix as the values of the DataFrame array, where the DataFrame array uses gene names as row and column indices; 3) Set the row and column indices of the DataFrame to gene names or other corresponding identifiers.

3. A method for analyzing gene expression networks based on the minimum entropy algorithm according to claim 1, characterized in that, The community classification evaluation formula of the Infomap multi-level network clustering model in S5 is expressed as: $Q = \sum_{i}(e_{ii}-a_{i}^2)+\frac{1}{2}\sum_{i\neq j}(e_{ij}-a_{i}a_{j})\delta(\sigma_i,\sigma_j)$ where $Q$ represents the modularity of the community partition, which is used to measure the quality of the community structure, $e_{ij}$ represents the edge weight between node $i$ and node $j$, $a_{i}$ represents the degree of node $i$, $\delta(\sigma_i,\sigma_j)$ is a Kronecker Delta function, which has a value of 1 if node $i$ and node $j$ belong to the same module, otherwise 0, $\sigma_i$ represents the module to which node $i$ belongs, and the goal of the Infomap algorithm is to find an optimal $\sigma_i$ allocation scheme to minimize $Q$.

Citation Information

Patent Citations

  • Method for constructing gene regulation network

    CN109215735A

  • Method and apparatus for neural network quantization

    US20180107925A1