Method, system and computer medium for identifying core regulatory factors in a regulatory network
By calculating the perturbation scores encoded by the auto-perturbation and external perturbation in the molecular regulatory network, the core regulatory factors in the regulatory network are accurately identified, which solves the problem of inaccurate identification in the existing technology, improves the accuracy of identification, and provides a new method for tumor research.
Patent Information
- Application Number
- CN202111505947.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-10
AI Technical Summary
It is difficult for the prior art to accurately identify core regulatory factors in regulatory networks, especially in specific cellular states, pathological states and genomic states.
By acquiring the molecular regulatory network, the auto-perturbation and external perturbation codes of the regulatory network nodes are obtained based on genome variation or nonspecific modification and regulation, the perturbation score is calculated, and whether the node is a core regulatory factor is judged based on the threshold.
It realizes the accurate identification of core regulatory factors in the regulatory network under specific conditions, improves the accuracy of identification, and provides a theoretical basis for the research on tumor microenvironment mechanisms and the development of targeted drugs.
Smart Images

Figure CN114333993B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of core regulatory factor recognition, and relates to a method, a system and a computer medium for recognizing core regulatory factors in a regulatory network. Background Art
[0002] Cancer is one of the three major diseases that seriously endanger human health in the world and is also a disease that endangers the health of Chinese residents. The annual incidence and death numbers of cancer respectively account for about 23.7% and 30% of the global total. The GLOBOCAN 2020 database shows that there were 19.29 million newly diagnosed cancer cases and 9.95 million cancer deaths globally in 2020. In the past more than 10 years, the survival rate of malignant tumors has shown a gradually increasing trend. At present, the 5-year relative survival rate of malignant tumors in China is about 40.5%. Therefore, the research on malignant tumors has become a hot issue in medical research, and the fundamental task of basic research on malignant tumors is to clarify the molecular regulatory mechanism of tumor occurrence and development.
[0003] In the past decade, with the rapid development of high-throughput detection technology (previously, only one or a few sequences could be detected simultaneously by Sanger sequencing. High-throughput detection technology uses a microarray, utilizes the precise pairing principle of molecular hybridization and a professional fluorescence detection software to detect millions of preset sequences to be detected on the array, achieving the purpose of simultaneously detecting a large number of genomic sequences of interest), combined with the generation and application of high-quality cancer genome data of tens of thousands of patients, strict statistical tools, and a large number of relatively complete patient clinical and follow-up information records, it is possible to study complex diseases such as cancer from multiple perspectives.
[0004] In recent years, exploring the mechanisms of tumor occurrence and development from a multi-omics perspective has become a hot area, and significant results have been achieved in the identification of early tumor mutation markers, the identification and identification of pan-tumor molecular subtypes, the study of tumor chemotherapy drug tolerance mechanisms, the identification of tumor epigenome markers, and the identification of tumor prognostic markers at various omics levels. With the increase in the ability to identify multi-omics level changes related to cancer, the academic community has reached some understandings, thus forming methods to clarify the causes of the disease. First, it has been determined that so far, gene-level variations usually only explain a small part of the tumor progression mechanism. Secondly, changes in the expression levels of tumor-related genes are also affected by the level of DNA methylation (a chemical modification on DNA double helix molecules, where the ends of these molecules are converted to methyl CH3 by other chemical groups through corresponding enzymes, which will reduce the efficiency of genes regulating transcription and the high methylation level of DNA in this section, and play a role in regulating gene transcription), and the methylation level of the same locus itself is different in different sexes and ages. Third, differential expression at the transcriptome level is often the result of multi-level regulation of different molecules, which depends on the stage of tumor progression. Taken together, these insights provide a theoretical basis for personalized cancer treatment, which involves integrating different omics data types to identify stage-specific tumor-associated molecular regulatory networks and their mechanisms.
[0005] Due to the complex multi-level and multi-regulatory characteristics of tumors, multi-omics data is the key to exploring the molecular regulatory mechanisms of tumors. Life systems tend to be in dynamic equilibrium, and so do molecular regulatory networks. Tumor occurrence and tumor progression are the result of the destruction of the stability (i.e., excessive perturbations) of normal molecular regulatory networks. Core regulatory factors (referring to key genes that regulate a phenotype, like tyrosinase in albinism, where tyrosinase deficiency will directly lead to albinism) are key nodes in specific phenotypic regulatory networks. Identifying core regulatory factors (complex diseases are not like single-gene diseases, where the phenotype is controlled by a complex regulatory network, and there are several more critical genes in the regulatory network) will pave the way for further research into the mechanisms of the tumor microenvironment and the development of new therapeutic targets and therapeutic combinations. However, the ability to accurately identify these factors is currently insufficient. Although a series of algorithms have been developed to identify core regulatory factors in regulatory networks, such as cytohubba and MCODE, these methods are mainly based on degree size and other factors. These algorithms are acceptable for the task of determining the approximate range of core regulatory factors. However, they are ineffective in accurately determining the core regulatory factors of molecular regulatory networks under specific cell states, pathological states, and genomic states (genomic mutations are a hallmark of cancer). Therefore, it is urgent to develop new algorithms for accurately identifying core regulatory factors in molecular regulatory networks. Summary of the invention
[0006] The object of the present invention is to provide a method, system and computer medium for identifying core regulatory factors in a regulatory network, so as to accurately identify core regulatory factors in the regulatory network.
[0007] To achieve the above object, the basic solution of the present invention is: A method for accurately identifying core regulatory factors in a regulatory network based on multi-omics data, comprising the following steps:
[0008] Obtain a molecular regulatory network;
[0009] Based on genomic variations or non-specific modification regulations, obtain the self-disturbance suffered by the nodes of the regulatory network;
[0010] Based on specific modification regulation relationships, obtain the external disturbance encoding suffered by the nodes of the regulatory network;
[0011] Calculate the disturbance score according to the self-disturbance and the external disturbance encoding;
[0012] Set a threshold, compare the disturbance score with the threshold, and judge whether the node of the regulatory network is a core regulatory factor according to the comparison result.
[0013] The working principle and beneficial effects of this basic solution are as follows: Utilize the correlation of the expression levels of molecules with each other under specific pathological or cellular states to measure the influence weights of different molecules on the same downstream molecule, comprehensively detect genomic mutations, copy number variations and epigenomic changes of nodes to measure the presence and magnitude of self-disturbance of nodes, decompose the regulatory effect received by a node into self-disturbance and external disturbance (upstream signal), and the regulatory effects of each upstream signal on the node are weighted, optimizing the accuracy of identifying core regulatory factors.
[0014] This method can be applied to different disease studies to identify the core regulatory factors of a specific phenotype in the disease, which can greatly improve the research efficiency. For drug development, the core factors identified by this method can be used to develop targeted drugs after being verified by in vitro and in vivo experiments. Compared with other potential targets that are not this method, the core factors provided by this method are identified on the premise of globally considering a large number of compensatory molecular mechanisms of the entire regulatory network, and theoretically have less drug resistance. It can greatly reduce the economic and human costs of secondary research and development.
[0015] Further, based on node criteria and edge criteria, the method for obtaining a molecular regulatory network is as follows:
[0016] Classify the cell fraction estimates and assign them to terciles, namely low, medium, and high. If more than 66% of the samples within the cluster are classified as "medium" or "high", the node will be retained for a node set; if the consistency score of two connected nodes is greater than the P25 of the consistency score distribution, the edge will be retained, obtaining an edge set. The regulatory network consists of its node set and edge set together;
[0017] The consistency score for downstream promotion is calculated as follows:
[0018]
[0019] The consistency score for downstream inhibition is calculated as follows:
[0020]
[0021] where n low,low is the number of samples with node 1 low and node 2 low; n low,high is the number of samples with node 1 low and node 2 high; n high,high is the number of samples with node 1 high and node 2 high; n high,low is the number of samples with node 1 high and node 2 low.
[0022] It is simple to operate and convenient to use.
[0023] Furthermore, the method for obtaining the encoding of the external perturbation received by the regulatory network node is as follows:
[0024] Let the encoding of the k-th perturbation event of node i be:
[0025]
[0026] The calculation structure is simple and convenient to use.
[0027] Furthermore, the method for determining whether a regulatory network node is a core regulatory factor is as follows:
[0028] For the weighted graph G(V,E) with node V(G)>1 and edge E(G)>0, the sample space Ω is obtained by randomly sampling Kp times. Let the weighted perturbation score of node i be PS ran (i), i∈Ω. Given a node j∈V(G), if the weighted score PS obs (i) satisfies:
[0029]
[0030] When P<0.05, then node j is a core regulatory factor in the weighted graph G(V,E).
[0031] Weighting the effects of different regulatory factors better fits the real regulatory network and improves the recognition accuracy.
[0032] Furthermore, compared with normal samples, the expression change of gene i in a specific state is quantified as:
[0033]
[0034] where pi is the p-value for differential expression hypothesis testing, and φ^(-1) is the inverse Gaussian cumulative distribution function; thus, the perturbation score is defined as:
[0035]
[0036] where CS j is the consistency score of gene i and upstream regulatory gene j in an immune cluster, i.e., the external perturbation weight; OR i,a is the odds ratio of the a-th self-perturbation event of gene i in the same group as CS j , i.e., the self-perturbation weight; m represents the number of genes.
[0037] Quantifying external perturbation and self-perturbation factors and calculating the weighted perturbation score of network nodes facilitate subsequent identification of regulatory factors.
[0038] The present invention also provides a system for accurately identifying core regulatory factors in a regulatory network based on the method described in the present invention, including a molecular regulatory network construction unit, a self-perturbation acquisition unit, an external perturbation acquisition unit, and a processing unit;
[0039] The molecular regulatory network construction unit is used to obtain a molecular regulatory network, and the molecular regulatory network construction unit is respectively connected to the input ends of the self-perturbation acquisition unit and the external perturbation acquisition unit;
[0040] The self-perturbation acquisition unit is used to obtain the self-perturbation received by the nodes of the regulatory network;
[0041] The external perturbation acquisition unit is used to obtain the external perturbation encoding received by the nodes of the regulatory network;
[0042] The processing unit is respectively connected to the output ends of the self-perturbation acquisition unit and the external perturbation acquisition unit. The processing unit is used to calculate the perturbation score and compare the perturbation score with a threshold to determine whether the nodes of the regulatory network are core regulatory factors.
[0043] Using this system, the regulatory effect received by a node of a regulatory network is decomposed into self-perturbation and external perturbation, and they are respectively mathematically quantified to achieve accurate identification of core regulatory factors.
[0044] The present invention also provides a computer medium, and a program for executing the method described in the present invention is stored in the computer medium.
[0045] The computer medium is convenient to use. By using the computer medium, the identification operation of core regulatory factors can be performed on various devices, expanding the scope of use. Brief Description of the Drawings
[0046] Figure 1It is a schematic flowchart of the method for accurately identifying core regulatory factors in a regulatory network based on multi-omics data in the present invention. Detailed implementation manners
[0047] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention.
[0048] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0049] In the description of the present invention, unless otherwise specified and defined, it should be noted that the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.
[0050] The present invention provides a method for constructing a molecular regulatory network, a method for identifying core regulatory factors in the regulatory network, and a method for establishing an interactive analysis and visualization platform for the molecular regulatory network based on the web. The specific technical solutions are as follows:
[0051] In view of the defects and deficiencies of the prior art, the present invention discloses a method for constructing a high-confidence molecular regulatory network. (1) Using text mining technology to mine the direct regulatory effects existing between each pair of nodes, and then using the expression levels under specific pathological or cell states and the correlation between this pair of nodes to measure whether this regulatory relationship exists in the current network. (2) On the premise of knowing the regulatory modes of all nodes, performing (1) can predict its regulatory mechanism. The construction method of this solution includes the following steps:
[0052] Determine the environment of the molecular regulatory network to be constructed. The environment can be selected according to the interests of the implementer and there are no conditional restrictions, because the regulatory network only exists under a certain specific state and it exists for regulating a certain phenotype. Preferably, the environment includes a determined tumor type or subtype.
[0053] Determine the molecular types of all nodes to be included in the regulatory network to be constructed. The molecular types are determined based on biological knowledge. For unclear molecular types of a gene, they can also be determined by consulting materials, such as biological literature, genecards database, etc. Preferably, the molecular types include transcription factors, receptors, ligands, miRNA (microRNA) and lncRNA (long non-coding RNA). Determining the molecular types will help determine whether there is a corresponding direct regulatory relationship between two molecules during text mining or database mining. For example, transcription factors regulate the transcription efficiency of target genes by binding to the target gene promoter before transcription.
[0054] Based on text mining literature and database mining technology, the regulatory relationships containing target molecule types that have been experimentally verified are integrated to obtain a regulatory relationship list, and the relationship list that matches the target molecule type in the list is extracted. The regulatory relationship list includes the pairwise regulatory relationships experimentally verified by literature reports, as well as the pairwise regulatory relationships that have been clearly recorded in the database and have been confirmed to exist.
[0055] Specifically, it can be but not limited to using a generalized linear model, a Boolean model or a Bayesian model to determine whether the mutual regulatory effect of all nodes (two-to-two relationships) of the molecular type to be analyzed exists. According to the extracted relationship list and the inferred results, a high-confidence molecular regulatory network for the corresponding environment is constructed. It can be implemented using a variety of software, various programming languages or platforms, such as R language, cytoscape software, etc., for visual network construction. For example, using R language, with seven tumor immune subtypes as the environment, a subtype-enhanced immune regulatory network is constructed. The first column represents the upstream node (gene name) of the regulatory relationship, the second column represents the downstream node (gene name) of the regulatory relationship, the third column represents the two nodes in which immune cells are co-expressed, the fourth column represents the regulatory relationship in which tumor immune subtype, the fifth column represents the quantification of the intensity of the regulatory relationship (based on a generalized linear model), and the sixth column represents the identification of the regulation and direction of the upstream to downstream nodes.
[0056] In a preferred embodiment of the present invention, the text mining literature and database mining technology include the following steps:
[0057] By mining the key areas of literature and databases, we can extract the regulatory relationships that exist therein;
[0058] Based on artificial intelligence, existing literature is used for model training to obtain a literature integration mathematical model. The model can adopt a convolutional neural network model CNN, a recurrent neural network model RNN, etc.
[0059] Apply the literature integration mathematical model to the new literature set to obtain the desired verified gene regulatory relationship. For example, gene A plays a role in promoting tumors by regulating gene B. The regulatory relationship can be extracted by mining the abstract of the paper (understanding the semantics), and the relationship "A regulates B" can be obtained and recorded. However, when the number of documents is very large, manual reading and sorting of documents will consume a lot of manpower and material resources, or even be impossible to complete. Therefore, it is very convenient to perform semantic understanding of biomedical texts based on artificial intelligence. By applying the pre-trained model to the new literature set, the desired verified gene regulatory relationship can be obtained.
[0060] In a preferred embodiment of the present invention, the method for determining whether the mutual regulation of all nodes of the molecular type to be analyzed exists is as follows:
[0061] Calculate the Spearman correlation of the two nodes. If the correlation coefficient |rho|≥0.3 and P<0.05, proceed to the next step. Use a generalized linear regression model (such as a logistic regression model) to calculate the relationship between the two nodes and the phenotype (such as tumor vs. normal). If P<0.05, it is considered that the two nodes have a regulatory relationship and thus affect the phenotype.
[0062] In a preferred embodiment of the present invention, a method for constructing a high-confidence molecular regulatory network corresponding to an environment based on the extracted relationship list and the inferred results is:
[0063] Based on the node criteria and edge criteria, the cell score estimates are classified and divided into tertiles, namely low, medium, and high. If more than 66% of the samples in the cluster are classified as "medium" or "high", the node will be retained to a node set; if the consistency score of two connected nodes is greater than the P25 of the consistency score distribution, the edge is retained to obtain an edge set. The regulatory network is composed of its node set and edge set.
[0064] The solution also provides a computer medium, in which a program for executing the method of the present invention is stored. The computer medium is easy to use, and can be used to construct a molecular regulatory network on a variety of devices, thereby expanding the scope of use.
[0065] The core regulatory factor recognition algorithms in the prior art mainly characterize node centrality in network analysis by quantifying degree centrality (Degree Centrality, which is used to mathematically quantify the importance of a node in a network for the entire network). Degree centrality is the most direct metric. The larger the node degree of a node, the higher its degree centrality, and the more important this node is in the network. In graph theory and networks, degree refers to the number of connections of a point in the network (graph) to other points. For a directed network, degree is divided into in-degree and out-degree. The in-degree refers to the number of edges pointing to this node, and the out-degree refers to the number of edges starting from this node and pointing to other nodes. For an undirected network, degree is the number of edges connected to the node, regardless of the direction of connection.
[0066] Existing algorithms assume that each gene in the molecular regulatory network has the same contribution to the signal network. In fact, even if different genes co-regulate the same function, their weights are often different. Genes also have the ability to regulate themselves (self-disturbance, defined as mutations in their own gene DNA, resulting in the inability of the gene to function properly or a significant promotion). For example, genomic variations occur. Considering the role of upstream signals in a general way cannot effectively decompose the effects of these signals (a gene is subject to multiple regulations, which may come from transcription factors, miRNAs, various enzymes, and DNA mutations in the genome respectively. The effects of these regulatory signals on the target gene may be superimposed, that is, they promote together, so it is considered to decompose these regulatory effects). As a result, the recognition results of core regulatory factors are not accurate enough. As Figure 1 shown, the present invention discloses a method for accurately identifying core regulatory factors in a regulatory network based on multi-omics data, using the correlation of expression levels between molecules under specific pathological or cellular states to measure the influence weights of different molecules on the same downstream molecule. And by comprehensively detecting genomic mutations, copy number variations, and changes in the epigenome of nodes, the presence and magnitude of self-disturbance of nodes are measured. The regulatory effects received by a node are decomposed into self-disturbance and external disturbance (upstream signals), and the regulatory effects of each upstream signal on the node are weighted. The method of this embodiment includes the following steps:
[0067] Obtain a molecular regulatory network, construct a molecular regulatory network or obtain a pre-constructed regulatory network from the literature;
[0068] Based on genomic variations or non-specific modification regulations, obtain the self-disturbance received by the nodes of the regulatory network;
[0069] Based on specific modification regulatory relationships, obtain the external disturbance codes received by the nodes of the regulatory network;
[0070] Calculate the disturbance score according to the self-disturbance and external disturbance codes;
[0071] Set a threshold, compare the perturbation score with the threshold, and based on the comparison result, determine whether the regulatory network node is a core regulatory factor.
[0072] In a preferred embodiment of the present invention, based on the node criterion and the edge criterion, the method for obtaining the molecular regulatory network is as follows:
[0073] Classify the estimated cell fraction values and assign them to terciles, namely low, medium, and high. If more than 66% of the samples within the cluster are classified as "medium" or "high", the node will be retained for a node set. The same applies to the expression of ligands and receptors. Ligands and receptors are both proteins and belong to gene-type nodes, and this score is their respective expression levels. The estimated cell fraction value is based on an algorithm CIBERSORT, which is used to estimate the proportions of various immune cell subsets and is a numerical estimate of the nodes in the network to quantify their levels in the body. If the node is a gene, it is the expression level; if the node is a cell, it is the number or proportion of the cells.
[0074] If the consistency score between two connected nodes is greater than the P25 of the consistency score distribution, retain the edge to obtain an edge set. The regulatory network is jointly composed of its node set and edge set. P25 is the 25th percentile, a type of percentile. n consistency scores form a statistical distribution. Sort this distribution from smallest to largest, and the value at the 25th percentile is P25.
[0075] The consistency score is the proportion of the co-variation of the levels between two nodes. The higher the proportion of co-variation, the greater the intensity of their mutual regulation, that is, the magnitude of the influence of the upstream regulation (external perturbation) on it, which is the weight mentioned above. The consistency score is calculated as follows:
[0076]
[0077] Similarly, a tumor-enhancing interaction regulatory network involving TF (transcription factor), miRNA, and lncRNA was identified based on node criteria (>66% of clustering samples classified as "medium" or "high") and edge criteria (distribution with consistency score > P25). Considering the ceRNA theory, miRNA-target and lncRNA-miRNA are expected to have a negative correlation. The ceRNA theory means endogenous competitive inhibitory RNA. Since lncRNA can inhibit miRNA, and miRNA can inhibit its own target genes (including transcription factors), then lncRNA can indirectly promote the target genes of miRNA by inhibiting miRNA. Since the lncRNA, miRNA, and target genes here are all endogenous, that is, none of them are input in vitro and are produced by their own cells, so it is called endogenous competitive inhibition. In addition, the inhibitory effects are all manifested as negative correlations. The consistency score is calculated as follows:
[0078]
[0079] where n low,low is the number of samples with node 1 low and node 2 low; n low,high is the number of samples with node 1 low and node 2 high; n high,high is the number of samples with node 1 high and node 2 high; n high,low is the number of samples with node 1 high and node 2 low. Each node (such as a gene) has a value (representing the expression level) in each sample. m nodes × n samples form an m×n matrix. Here, low or high means whether the value of the node in the distribution of all samples is in the low (less than P33, i.e., the 33rd percentile) or high (greater than P67, i.e., the 67th percentile) of the three quantiles (low, medium, high).
[0080] In a preferred embodiment of the present invention, the method for encoding the external perturbations received by the regulatory network nodes is as follows:
[0081] Let the kth perturbation event encoding of node i be:
[0082]
[0083] In a preferred embodiment of the present invention, the method for determining whether a regulatory network node is a core regulatory factor is as follows:
[0084] For a weighted graph G(V,E) with node V(G)>1 and edge E(G)>0, the sample space Ω is obtained by randomly sampling Kp times. Let the weighted perturbation score of node i be PS ran (i), i ∈ Ω. Given a node j ∈ V(G), if the weighted score PS obs (i) satisfies:
[0085]
[0086] When P < 0.05, node j is a core regulatory factor in the weighted graph G(V, E) and has a very high orderliness. A P-value needs to be calculated for each node, indicating whether the magnitude of the perturbation score value of this node is a statistically random event (if it is not a random event, it means it is unusually large, and unusually large means significant in statistics). When calculating the P-value of any node, the perturbation score of this node is denoted as PS obs Among all other nodes, any one node is selected to calculate its perturbation score, denoted as PS ran With replacement sampling 100,000 times, a distribution of PS ran is formed. The P-value is the percentile at which PS obs is located in the distribution of PS ran (sorted from large to small). Generally, P < 0.05 is used as the threshold. When P < 0.05, it is considered that the perturbation score of this node is large enough.
[0087] Compared with normal samples, the expression change of gene i in a specific state is quantified as:
[0088]
[0089] where p i is the p-value of the differential expression hypothesis test (Welch t-test), and φ^(-1) is the inverse Gaussian cumulative distribution function; therefore, the perturbation score is defined as:
[0090]
[0091] where CS is the abbreviation of Concordance Score, and CS j is the concordance score between gene i and the upstream regulatory gene j in an immune cluster, that is, the external perturbation weight; OR is the odds ratio, which is a value often calculated in statistics to represent the risk coefficient (the magnitude of the influence), and OR i,a is the odds ratio (Fisher's test) of the a-th self-perturbation event of gene i in the same group as CS j , that is, the self-perturbation weight; m represents the number of genes.
[0092] The present invention also provides a system for accurately identifying core regulatory factors in a regulatory network based on the method of the present invention, including a molecular regulatory network construction unit, a self-perturbation acquisition unit, an external-perturbation acquisition unit, and a processing unit.
[0093] A molecular regulatory network construction unit is used to obtain a molecular regulatory network. The molecular regulatory network construction unit is respectively connected to the input ends of a self-perturbation acquisition unit and an external-perturbation acquisition unit. The self-perturbation acquisition unit is used to obtain the self-perturbation received by the nodes of the regulatory network. The external-perturbation acquisition unit is used to obtain the external-perturbation encoding received by the nodes of the regulatory network. A processing unit is respectively connected to the output ends of the self-perturbation acquisition unit and the external-perturbation acquisition unit. The processing unit is used to calculate a perturbation score and compare the perturbation score with a threshold value to determine whether the nodes of the regulatory network are core regulatory factors.
[0094] This solution utilizes the regulatory network existing in a specific cell or pathological state. The self-perturbation is based on genomic variations or non-specific (genome-wide, such as DNA methylation) modification regulations, and the external perturbation is based on specific (targeting only one or a class of targets) modification regulation relationships. The Concordance score1 (agreement score) is used to define the agreement for the interaction nodes with downstream promoting effects, and conversely, the Concordance score 2 is used to define the agreement for the interaction nodes with downstream inhibitory effects. Only high-risk self-perturbation factors are included in the algorithm for consideration, that is, log10(odds ratio)≥0.5. This algorithm can be developed using the R language, has platform compatibility, and is relatively convenient to implement.
[0095] In the prior art, the regulatory network analysis based on the Cytoscape software is widely used. The regulatory network refers to the network formed by the interaction relationships between genes within a cell or within a genome, specifically referring to the effects caused by gene regulation. Regulation means that the function of one gene controls whether another gene plays a role or changes the magnitude of this role through a certain way, such as transcriptional regulation.
[0096] Currently, the commonly used software or method for analyzing regulatory networks is to use the Cytoscpae software, which is time-consuming and laborious for offline input, configuration, layout, calculation, etc. of regulatory network files. For networks with extremely high complexity, the requirements for personal computers are also correspondingly increased, which is not conducive to lightweight office work. The offline regulatory network analysis not only has a relatively high threshold but also is not conducive to customizing the search for sub-networks of genes of interest and their related biological information.
[0097] The present invention discloses a method for constructing an interactive analysis and visualization platform for a web-based molecular regulatory network, which uses a web-based online database to provide online interactive analysis and visualization of the molecular regulatory network. For each node (such as a gene), very detailed biological information (sequence, subcellular localization, tissue expression level, protein quantification, etc.) is collected and provided for easy reference. The method of this solution includes the following steps:
[0098] The regulatory relationships of the network are dependent on specific cells or pathological states. For example, there is a mutual regulatory effect between genes A and B in cardiomyocytes, but this regulatory effect does not exist in hepatocytes; or this regulatory effect exists in liver cancer, but does not exist in normal liver cells. At the same time, it also depends on the entire network being searched.
[0099] Integrate and standardize the access to the regulatory network data offline to obtain the node set file and the edge set file, and import them into the mysql (relational database management system) database on the server side. A regulatory network is jointly described by its node set and edge set (the regulatory relationships between two nodes), corresponding to the nodes.txt and edges.txt files respectively, and import them into the mysql database respectively.
[0100] Preferably, the regulatory network has multiple layout options and can arbitrarily extract the sub-networks of interest, and can automatically calculate the positions of its nodes according to multiple layout algorithms. The layout algorithms include random, grid, concentric, breadthfirst or cose. Each layout has its established layout method and established program to determine the styles presented by each node in each layout mode. For example, the grid layout presents the network as a matrix. When calculating the positions of each node on the screen, only parameters such as the number of nodes, the diameter of the nodes, the degree of the nodes, the node spacing, and the screen size need to be obtained to determine.
[0101] Develop software to extract the parameters of the gene sub-network (including the sub-network of the gene of interest) on the server side, as well as generate the data required for the web-side response and visualization, and can communicate uniformly in json format. According to the extracted data, use computer languages to develop the web-side interface and event response program. Preferably, the computer language uses html language, css language or javascript language.
[0102] Develop communication interfaces between the web side and the server side using programming languages for transmitting parameters, etc. The programming language course can use PHP (PHP: Hypertext Preprocessor, a scripting language executed on the server side, especially suitable for web development and can be embedded in HTML), or ASP, NET language, etc. Transmit the generated JSON parameter file to the server side and call the Perl program. The execution program is completed by PHP code. The server-side Apache (pronounced as Apache, the world's number one web server software. It can run on almost all widely used computer platforms and is widely used due to its cross-platform and security features. It is one of the most popular web server software) service interprets the returned PHP code, completes the events defined by the PHP code (i.e., extracts the subnetworks and generates the result file), and transmits it back to the browser for parsing and visualization. Complete the construction of the visualization platform. This platform uses the cytoscape.js extension to make the network visualization module compatible with multiple platforms. As long as it can be visualized using the offline cytoscape software, the web application will be more convenient.
[0103] In a preferred embodiment of the present invention, the method for extracting data on the server side based on software using the Perl programming language is as follows:
[0104] Read the parameters of the subnetwork to be extracted. The parameters of the subnetwork include the id of the gene of interest and the name of the subnetwork. According to the parameters of the subnetwork, read the data of the subnetwork from the MySQL database. Store the data of the subnetwork in JSON file format and TXT format files for constructing the network diagram on the web side.
[0105] In a preferred embodiment of the present invention, the method for developing the web side interface and event response program is as follows:
[0106] Use HTML (Hypertext Markup Language) to define the structural framework of the web page, and then use CSS (Cascading Style Sheets) language to render and beautify the style appearance of HTML components. The structural framework of the web page includes the web page logo and title, navigation bar, regulation network diagram container, network diagram toolbox container, network diagram annotation container, network diagram detailed description container, and detailed parameter containers for nodes, edges, and regulation axes in the network, documents, and contact information.
[0107] Implement page positioning and redirection, network layout switching, network-related data downloading, auto-suggestion for search boxes, and subnet search using the JavaScript language. Taking subnet search as an example, its internal process sequentially includes collecting parameters and storing them in variables, receiving the parameter variables by the PHP language, sequentially checking whether each parameter meets the requirements, filling in the missing parameters with default values, and generating a cleaned parameter list, which is transmitted to a program using the Perl language on the server in JSON format;
[0108] The requirements include: ① The gene names use the unified name of HGNC gene symbol, referring to https: / / www.genenames.org / ; ② The subnet types belong to C1-C7 or 33 tumor types in The Cancer Genome Atlas (TCGA) of human cancers, referring to https: / / www.cancer.gov / about-nci / organization / ccg / research / structural-genomics / tcga / studied-cancers; ③ The missing parameters are filled with the default value Null.
[0109] The present invention also provides an interactive analysis and visualization platform for constructing a web-based molecular regulatory network, including a server side, a web side, and a processor. The processor executes the method of the present invention, controls the connection between the server side and the web side, and completes the construction of the visualization platform. The core of this platform is developed using HTML, CSS, JavaScript, PHP, and Perl languages. Server-side software is developed for integration, texturizing the molecular regulatory network, extracting the subnets required for web visualization, and generating standardized format files, and communicating with the web side. On the web side, based on user operations, interactive analysis and layout of the regulatory network are performed, making the analysis more real-time and convenient, and making the network more concise and beautiful. It can greatly facilitate researchers to analyze and visualize the molecular regulatory network, with low threshold, fast, convenient, and customized retrieval of the regulatory network of genes of interest and their detailed information.
[0110] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0111] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for accurately identifying core regulatory factors in a regulatory network based on multi-omics data, characterized in that, Including the following steps: Obtain the molecular regulatory network; Based on genomic variations or non-specific modification regulations, obtain the self-perturbations received by the regulatory network nodes; Based on specific modification regulatory relationships, obtain the external perturbation codes received by the regulatory network nodes; Calculate the perturbation score according to the self-perturbations and external perturbation codes; Set a threshold, compare the perturbation score with the threshold, and according to the comparison result, determine whether the regulatory network node is a core regulatory factor; Based on the node criteria and edge criteria, the method for obtaining the molecular regulatory network is as follows: Classify the estimated cell fractions and assign them to terciles, namely low, medium, and high. If more than 66% of the samples within the cluster are classified as "medium" or "high", the node will be retained in a node set; if the consistency score of two connected nodes is greater than P25 of the consistency score distribution, the edge will be retained, obtaining an edge set. The regulatory network is jointly composed of its node set and edge set; The calculation method of the consistency score for downstream promotion is as follows: The calculation method of the consistency score for downstream inhibition is as follows: where n low,low is the number of samples with node 1 low and node 2 low; n low,high is the number of samples with node 1 low and node 2 high; n high,high is the number of samples with node 1 high and node 2 high; n high,low is the number of samples with node 1 high and node 2 low; The method for obtaining the external perturbation codes received by the regulatory network nodes is as follows: Let the k-th perturbation event code of node i be: The method for determining whether the regulatory network node is a core regulatory factor is as follows: For a weighted graph \(G(V, E)\) with \(|V(G)|>1\) nodes and \(|E(G)| > 0\) edges, the sample space \(\Omega\) is obtained by randomly sampling with replacement from \(V(G)\) \(K_p\) times. Let the weighted perturbation score of node \(i\) be \(PS ran (i)\), where \(i\in\Omega\). Given a node \(j\in V(G)\), if the weighted score \(PS obs (i)\) satisfies: When P < 0.05, node j is a core regulatory factor in the weighted graph G(V,E); Compared with normal samples, the expression change of gene i in a specific state is quantified as: where p i is the p-value of the differential expression hypothesis test, and φ^(-1) is the inverse Gaussian cumulative distribution function; thus, the perturbation score is defined as: Among them, CS j is the consistency score of gene i and upstream regulatory gene j in an immune cluster, that is, the external perturbation weight; OR i,a is the odds ratio of the a-th self-perturbation event of gene i in the same group as CS j , that is, the self-perturbation weight; m represents the number of genes.
2. A system for accurately identifying core regulatory factors in a regulatory network based on the method described in claim 1, characterized in that, Including a molecular regulatory network construction unit, a self-perturbation acquisition unit, an external perturbation acquisition unit, and a processing unit; The molecular regulatory network construction unit is used to obtain the molecular regulatory network, and the molecular regulatory network construction unit is respectively connected to the input ends of the self-perturbation acquisition unit and the external perturbation acquisition unit; The self-perturbation acquisition unit is used to obtain the self-perturbations received by the regulatory network nodes; The external perturbation acquisition unit is used to obtain the external perturbation codes received by the regulatory network nodes; The processing unit is respectively connected to the output ends of the self-perturbation acquisition unit and the external perturbation acquisition unit. The processing unit is used to calculate the perturbation score and compare the perturbation score with the threshold to determine whether the regulatory network node is a core regulatory factor.
3. A computer medium, characterized in that, The computer medium stores a program for executing the method described in claim 1.
Citation Information
Patent Citations
Construction method of high-confidence molecular regulation and control network and computer medium
CN114283879A
Method for constructing visual platform of molecular regulation and control network and visual platform
CN114283891A