Method and computer medium for constructing high-confidence molecular regulatory networks
Through text mining and database mining technology, combined with generalized linear models and other methods, a high confidence molecular regulatory network is built, which solves the problems of low confidence and difficult to determine the regulation method in the existing technology, and achieves more accurate prediction of gene regulation mechanisms and support for tumor research.
Patent Information
- Application Number
- CN202111505949.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The prior art has low confidence when constructing analytical gene networks, which are difficult to determine the regulation mode and direction, and it is difficult to deal with the complex multi-level multiple regulatory characteristics of tumors.
Through text mining and database mining technologies, the regulatory relationships that have been verified experimentally are integrated, and a high confidence molecular regulatory network is constructed using methods such as generalized linear models, and reverse verification and iteration are carried out to determine the regulation method and direction.
It improves the confidence of the molecular regulatory network, can predict and understand gene regulatory mechanisms more accurately, and is helpful for tumor biology research and drug development.
Smart Images

Figure CN114283879B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of molecular regulatory networks and relates to a method for constructing a high-confidence molecular regulatory network and a computer medium. Background Art
[0002] Cancer is one of the three major diseases that seriously endanger human health in the world, and it is also one of the diseases that endanger the health of Chinese residents. Its annual incidence and mortality account for about 23.7% and 30% of the world respectively. The GLOBOCAN2020 database shows that in 2020, there were 19.29 million new cases of cancer and 9.95 million cancer deaths worldwide. In the past 10 years, the survival rate of malignant tumors has shown a gradual upward trend. At present, the 5-year relative survival rate of malignant tumors in my country is about 40.5%. Therefore, the study of malignant tumors has become a hot topic in medical research, and the fundamental task of basic research on malignant tumors is to clarify the molecular regulatory mechanism of tumor occurrence and development.
[0003] With the innovation of high-throughput technology, the generation and application of high-quality cancer genome data from tens of thousands of patients, rigorous statistical tools, and a large number of relatively complete patient clinical and follow-up information records, it is possible to study cancer, a complex disease, from multiple perspectives. Cancer function and development are controlled by a network that regulates gene expression, not the action of a single gene. The importance of precise regulation of gene expression through development and cell differentiation is well known, and perturbations in gene regulation are associated with many complex diseases, including cancer.
[0004] Gene Regulatory Network (GRN), referred to as regulatory network, refers to the network formed by the interaction between genes in a cell or a genome, specifically the interaction between genes caused by gene regulation. GRN is the mechanism for controlling gene expression in organisms, and the main process of gene expression is transcription + translation.
[0005] There are many ways to construct and analyze gene networks, including but not limited to Boolean networks, linear models, Markov models, differential equation models, Bayesian network models, mutual information association models, random equation models, etc. Boolean networks are the simplest models. The states of each gene in Boolean networks are only "on" and "off". "On" means that the gene is expressed, and "off" means that the gene is not expressed. However, this network is too simplified and has limitations. Linear models are continuous GRN models. The expression level of a gene is represented by the weighted sum of the expression levels of several other genes. The weight is the quantification of the relationship between genes: positive weights represent gene excitation, negative weights represent gene inhibition, and 0 weights represent that the two genes have no relationship. Linear model networks can only process gene expression data with linear relationships, and their application scope is small. Markov chain is a random process suitable for analyzing gene expression data of time series. Markov chain assumes that the gene expression level at a certain moment determines the gene expression level at the next moment. In the process of constructing GRN, the feature extraction and clustering of gene expression profile based on Markov model show good adaptability, but if the accuracy of the model is to be improved, the order of Markov model needs to be greatly increased, and the operation is complicated. The differential equation model assumes that a gene is a variable, and a network composed of n genes can be represented by the following n-dimensional differential equation table. It is powerful and flexible, and is conducive to describing the complex relationships in the gene network. Based on Bayes' theorem and assumptions, the probabilistic relationship between random variables is represented in the form of a directed acyclic graph (DAG). Each gene in the network is a node, and each regulatory relationship is an edge. The model can handle random events, control noise, and obtain the causal relationship between variables.
[0006] The methods used in the prior art to construct and analyze gene networks are purely prediction-based methods with low confidence, and do not combine existing knowledge for reverse verification to optimize the network model. It is also difficult to determine the mode of regulation (transcriptional regulation, translational regulation, etc.) and direction (inhibition or promotion). Due to the complex multi-level and multi-regulatory characteristics of tumors, it is urgent to develop new methods to construct high-confidence molecular regulatory networks on the premise of determining the mode of regulation. Summary of the invention
[0007] The purpose of the present invention is to provide a method for constructing a high-confidence molecular regulatory network and a computer medium, so as to obtain a high-confidence molecular regulatory network for easy use.
[0008] In order to achieve the above object, the basic scheme of the present invention is: a method for constructing a high-confidence molecular regulatory network, comprising the following steps:
[0009] Determine the environment of the molecular regulatory network to be constructed;
[0010] Determine the molecular types of all nodes to be included in the regulatory network to be constructed;
[0011] Based on text mining literature and database mining technology, the regulatory relationships containing target molecule types that have been experimentally verified are integrated to obtain a regulatory relationship list, and the relationship list that matches the target molecule type is extracted from the list;
[0012] Determine whether the mutual regulatory effects of all nodes of the molecular type to be analyzed exist;
[0013] Based on the extracted relationship list and inferred results, a high-confidence molecular regulatory network corresponding to the environment is constructed.
[0014] The working principle and beneficial effects of this basic solution are: using text mining technology to explore the direct regulatory effects of each pair of nodes, and then using the expression level and the correlation of the pair of nodes under specific pathological or cellular conditions to measure whether the regulatory relationship exists in the current network. Under the premise that the regulatory methods of all nodes are known, their regulatory mechanisms can be predicted. Innovatively, the regulatory relationship based on experimental verification, the regulatory effect between network nodes is inferred based on the generalized linear model, and reverse verification and iteration are performed based on the established knowledge system to determine its regulatory method and obtain a molecular regulatory network with high confidence.
[0015] Further, the environment includes tumor type or subtype.
[0016] High-confidence tumor molecular regulatory networks will greatly enhance tumor biologists' understanding of tumor mechanisms, help clarify the molecular mechanisms of tumor occurrence and development, and will also greatly facilitate experimental verification and drug application.
[0017] Furthermore, the molecular types include transcription factors, receptors, ligands, miRNA and lncRNA.
[0018] Determine the molecule type for subsequent network construction operations.
[0019] Further, the text mining literature and database mining technology includes the following steps:
[0020] By mining the key areas of literature and databases, we can extract the regulatory relationships that exist therein;
[0021] Based on artificial intelligence, the existing literature is used for model training to obtain a literature integration mathematical model;
[0022] Apply the literature integration mathematical model to the new literature set to obtain the desired verified gene regulatory relationships.
[0023] Simple operation and easy to use.
[0024] Furthermore, the regulatory relationship list includes pairwise regulatory relationships verified by experiments reported in literature, as well as pairwise regulatory relationships that are clearly recorded in the database and have been confirmed to exist.
[0025] The control relationship list contains the required data for subsequent use.
[0026] Furthermore, the method for determining whether the mutual regulation of all nodes of the molecular type to be analyzed exists is as follows:
[0027] Calculate the Spearman correlation between two nodes. If the correlation coefficient |rho|≥0.3 and P<0.05, proceed to the next step;
[0028] The generalized linear regression model was used to calculate the relationship between the two nodes and the phenotype. If P < 0.05, it was considered that there was a regulatory relationship between the two nodes and thus affected the phenotype.
[0029] The judgment process is simple and easy to use.
[0030] Furthermore, based on the extracted relationship list and inference results, the method for constructing a high-confidence molecular regulatory network corresponding to the environment is:
[0031] Based on the node criteria and edge criteria, the cell score estimates are classified and divided into tertiles, namely low, medium, and high. If more than 66% of the samples in the cluster are classified as "medium" or "high", the node will be retained to a node set; if the consistency score of two connected nodes is greater than the P25 of the consistency score distribution, the edge is retained to obtain an edge set. The regulatory network is composed of its node set and edge set.
[0032] Simple operation and easy to use.
[0033] The solution also provides a computer medium, in which a program for executing the method of the present invention is stored.
[0034] Computer media are easy to use and can be used to construct molecular regulatory networks on a variety of devices, thereby expanding the scope of use. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the process of constructing a high-confidence molecular regulatory network of the present invention;
[0036] Figure 2 It is a schematic diagram of the structure of a molecular regulatory network of a preferred embodiment of the present invention;
[0037] Figure 3 It is a schematic diagram of a regulatory relationship list of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0038] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0039] In the description of the present invention, it is necessary to understand that the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0040] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal connection between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.
[0041] The present invention provides a method for constructing a molecular regulatory network, a method for identifying core regulatory factors in a regulatory network, and a method for establishing a web-based interactive analysis and visualization platform for molecular regulatory networks. The specific technical solutions are as follows:
[0042] The current method of constructing molecular regulatory networks is mainly based on putative molecular interactions, and does not use rapidly updated literature knowledge to reversely verify and optimize the network. It is only based on a specific model to predict whether there is an interaction between two nodes, which is the "edge" reflected in the network. The confidence of this predicted interaction is low, and cell and molecular biology experiments are needed to verify this interaction. Researchers have studied the regulatory relationship (or interaction) between certain genes and published papers. This program collects the gene-gene interactions verified in these papers through text mining, which can be used to reversely verify which ones have been verified and which ones have not been verified in the regulatory network predicted by the model, thereby greatly improving the confidence of the regulatory network.
[0043] The relationship between nodes in the molecular regulatory network of the prior art can only tell whether there is an effect or the weight of the effect, but cannot tell how they interact. There are many ways of gene regulation. For example, gene A has a regulatory effect on gene B. If it is chemically modified, it may be acetylation, methylation, phosphorylation, ubiquitination, etc. The regulatory relationship in the network predicted by the existing method can only infer whether A has a regulatory effect on B (semi-qualitative) and the size of the effect of A on B (quantitative). However, it is not clear what kind of chemical modification it is, which is not convenient to use.
[0044] In view of the defects and shortcomings of the prior art, the present invention discloses a method for constructing a high-confidence molecular regulatory network. (1) Text mining technology is used to mine the direct regulatory effects of each pair of nodes, and then the expression level and the correlation of the pair of nodes under specific pathological or cellular conditions are used to measure whether the regulatory relationship exists in the current network. (2) Under the premise that the regulatory methods of all nodes are known, (1) can be performed to predict their regulatory mechanisms. Figure 1 As shown, the scheme construction method includes the following steps:
[0045] Determine the environment of the molecular regulatory network to be constructed. The environment can be selected according to the interests of the implementer and is not conditionally limited, because the regulatory network only exists in a certain state and exists to regulate a certain phenotype. Preferably, the environment includes a determined tumor type or subtype.
[0046] Determine the molecular types of all nodes to be included in the regulatory network to be constructed. The molecular types are determined based on biological knowledge. For unclear molecular types of a gene, they can also be determined by consulting materials, such as biological literature, genecards database, etc. Preferably, the molecular types include transcription factors, receptors, ligands, miRNA (microRNA) and lncRNA (long non-coding RNA). Determining the molecular types will help determine whether there is a corresponding direct regulatory relationship between two molecules during text mining or database mining. For example, transcription factors regulate the transcription efficiency of target genes by binding to the target gene promoter before transcription.
[0047] Based on text mining literature and database mining technology, the regulatory relationships containing target molecule types that have been experimentally verified are integrated to obtain a regulatory relationship list, and the relationship list that matches the target molecule type in the list is extracted. The regulatory relationship list includes the pairwise regulatory relationships experimentally verified by literature reports, as well as the pairwise regulatory relationships that have been clearly recorded in the database and have been confirmed to exist.
[0048] Specifically, it can be but not limited to using a generalized linear model, a Boolean model or a Bayesian model to determine whether there is a mutual regulatory effect between all nodes (pairwise relationships) of the molecular type to be analyzed. Based on the extracted relationship list and the inferred results, a high-confidence molecular regulatory network for the corresponding environment is constructed. This can be implemented using a variety of software, various programming languages or platforms, such as R language, cytoscape software, etc., for visual network construction. For example, using R language, with seven tumor immune subtypes as the environment, a subtype-enhanced immune regulatory network is constructed. Figure 2 As shown in the table below, its regulatory relationships are listed as follows Figure 3 As shown, the first column represents the upstream node (gene name) of the regulatory relationship, the second column represents the downstream node (gene name) of the regulatory relationship, the third column represents in which immune cells the two nodes are co-expressed, the fourth column represents in which tumor immune subtypes the regulatory relationship exists, the fifth column represents the quantification of the strength of the regulatory relationship (based on the generalized linear model), and the sixth column represents the identification of the regulation from the upstream to the downstream node and its direction.
[0049] In a preferred embodiment of the present invention, the text mining literature and database mining technology include the following steps:
[0050] By mining the key areas of literature and databases, we can extract the regulatory relationships that exist therein;
[0051] Based on artificial intelligence, existing literature is used for model training to obtain a literature integration mathematical model. The model can adopt a convolutional neural network model CNN, a recurrent neural network model RNN, etc.
[0052] Apply the literature integration mathematical model to the new literature set to obtain the desired verified gene regulatory relationship. For example, gene A plays a role in promoting tumors by regulating gene B. The regulatory relationship can be extracted by mining the abstract of the paper (understanding the semantics), and the relationship "A regulates B" can be obtained and recorded. However, when the number of documents is very large, manual reading and sorting of documents will consume a lot of manpower and material resources, or even be impossible to complete. Therefore, it is very convenient to perform semantic understanding of biomedical texts based on artificial intelligence. By applying the pre-trained model to the new literature set, the desired verified gene regulatory relationship can be obtained.
[0053] In a preferred embodiment of the present invention, the method for determining whether the mutual regulation of all nodes of the molecular type to be analyzed exists is as follows:
[0054] Calculate the Spearman correlation of the two nodes. If the correlation coefficient |rho|≥0.3 and P<0.05, proceed to the next step. Use a generalized linear regression model (such as a logistic regression model) to calculate the relationship between the two nodes and the phenotype (such as tumor vs. normal). If P<0.05, it is considered that the two nodes have a regulatory relationship and thus affect the phenotype.
[0055] In a preferred embodiment of the present invention, a method for constructing a high-confidence molecular regulatory network corresponding to an environment based on the extracted relationship list and the inferred results is:
[0056] Based on the node criteria and edge criteria, the cell score estimates are classified and divided into tertiles, namely low, medium, and high. If more than 66% of the samples in the cluster are classified as "medium" or "high", the node will be retained to a node set; if the consistency score of two connected nodes is greater than the P25 of the consistency score distribution, the edge is retained to obtain an edge set. The regulatory network is composed of its node set and edge set.
[0057] The solution also provides a computer medium, in which a program for executing the method of the present invention is stored. The computer medium is easy to use, and can be used to construct a molecular regulatory network on a variety of devices, thereby expanding the scope of use.
[0058] The present invention also provides a method for accurately identifying core regulatory factors in molecular regulatory networks based on multi-omics data, using the correlation between the expression levels of molecules and molecules under specific pathological or cellular states to measure the influence weights of different molecules on the same downstream molecules. And by comprehensively detecting the genomic mutations, copy number variations and epigenomic changes of the nodes, the presence and size of node self-disturbance are measured. The regulatory effects on a node are decomposed into self-disturbance and external disturbances (upstream signals), and the regulatory effects of each upstream signal on the node are weighted. The method comprises the following steps:
[0059] Obtain molecular regulatory networks, construct molecular regulatory networks, or obtain established regulatory networks from literature. Based on genomic variation or non-specific modification regulation, obtain the self-disturbance of regulatory network nodes. Based on specific modification regulation relationships, obtain the external disturbance encoding of regulatory network nodes. Calculate the disturbance score based on the self-disturbance and external disturbance encoding. Set a threshold, compare the disturbance score with the threshold, and determine whether the regulatory network node is a core regulatory factor based on the comparison result.
[0060] Based on the node standard and edge standard, the method of obtaining the molecular regulatory network is as follows:
[0061] The cell score estimates are classified and divided into tertiles, low, medium, and high. If more than 66% of the samples in the cluster are classified as "medium" or "high", the node will be reserved for a node set. The same is true for the expression of ligands and receptors. Both ligands and receptors are proteins and belong to gene-type nodes. This score is their respective expression levels. The cell score estimate is based on an algorithm CIBERSORT, which is used to estimate the proportion of various immune cell subsets. It is a numerical estimate of the nodes in the network to quantify its level in the body. If the node is a gene, it is the expression level; if the node is a cell, it is the number or proportion of cells.
[0062] If the consistency score of two connected nodes is greater than P25 of the consistency score distribution, the edge is retained to obtain an edge set. The regulatory network consists of its node set and edge set. P25 is the 25th percentile, a type of percentile. n consistency scores form a statistical distribution. The distribution is sorted from small to large, and the value at the 25th percentile is P25.
[0063] The consistency score is the ratio of the horizontal common changes between two nodes. The higher the ratio of common changes, the greater the intensity of their mutual regulation, that is, the magnitude of the impact of upstream regulation (external disturbance) on it, which is called the weight above. The consistency score is calculated as follows:
[0064]
[0065] Similarly, tumor-enhancing interaction regulatory networks involving TF (transcription factors), miRNA, and lncRNA were identified based on node criteria (>66% of clustered samples were classified as "medium" or "high") and edge criteria (distribution of consistency scores >P25). Considering that miRNA-target and lncRNA-miRNA are expected to have a negative correlation based on the ceRNA theory, the ceRNA theory means endogenous competitive inhibitory RNA, because lncRNA can inhibit miRNA, and miRNA can inhibit its own target genes (including transcription factors), then lncRNA can indirectly promote the target genes of miRNA by inhibiting miRNA. Since the lncRNA, miRNA, and target genes here are all endogenous, that is, they are not input in vitro, but produced by their own cells, it is called endogenous competitive inhibition. In addition, the inhibitory effects are all reflected as negative correlations. The consistency score is calculated as follows:
[0066]
[0067] Among them, n low,low is the number of samples with node 1 low and node 2 low; n low,highis the number of samples where node 1 is low and node 2 is high; n high,high is the number of samples with node 1 high and node 2 high; n high,low is the number of samples with high values in node 1 and low values in node 2. Each node (such as a gene) has a value (indicating the expression level) in each sample, and m nodes × n samples form an m × n matrix. The low or high here indicates whether the node is low (less than P33, i.e., the 33rd percentile) or high (greater than P67, i.e., the 67th percentile) in the distribution of the values of all samples.
[0068] In a preferred embodiment of the present invention, the method for obtaining the external disturbance code of the control network node is as follows:
[0069] Suppose the kth disturbance event of node i is encoded as:
[0070]
[0071] In a preferred embodiment of the present invention, the method for determining whether a regulatory network node is a core regulatory factor is as follows:
[0072] For a weighted graph G(V,E) with nodes V(G)>1 and edges E(G)>0, the sample space Ω is obtained by random sampling Kp times. Let the weighted perturbation score of node i be PS ran (i), i∈Ω, given a node j∈V(G), if the weighted score PS obs (i) Satisfy:
[0073]
[0074] Then, node j is the core regulatory factor in the weighted graph G(V,E) and has a very high order. Each node must calculate a P value, indicating whether the size of the perturbation score of the node is a statistically random event (if it is not a random event, it means that it is unusually large, and unusually large means statistically significant). When calculating the P value of any node, the perturbation score of the node is calculated and recorded as PS obs Among all other nodes, any node is selected to calculate its disturbance score and recorded as PS ran , with replacement sampling 100,000 times, forming a PS ran The distribution of PS obs In PS ran The percentile of the distribution of is the value of P. Generally, P<0.05 is used as the threshold. If P<0.05, it is considered that the perturbation score of this node is large enough.
[0075] Compared with the normal sample, the expression change of gene i in a specific state is quantified as:
[0076]
[0077] Among them, p i is the p-value of the differential expression hypothesis test (Welch t-test), and φ^(-1) is the inverse Gaussian cumulative distribution function; therefore, the perturbation score is defined as:
[0078]
[0079] Among them, CS is the abbreviation of Concordance Score. j is the consistency score between gene i and upstream regulatory gene j in an immune cluster, i.e., the external perturbation weight; OR is odds ratio, which is a value often calculated in statistics to represent the risk coefficient (the size of the impact). i,a Is with CS j The odds ratio (Fisher's test) of the a-th self-disturbance event of gene i in the same group, that is, the self-disturbance weight; m represents the number of genes.
[0080] The present invention also provides a system for accurately identifying core regulatory factors in a regulatory network based on the method of the present invention, comprising a molecular regulatory network construction unit, a self-disturbance acquisition unit, an external disturbance acquisition unit and a processing unit.
[0081] The molecular regulatory network construction unit is used to obtain the molecular regulatory network. The molecular regulatory network construction unit is connected to the input ends of the self-disturbance acquisition unit and the external disturbance acquisition unit respectively. The self-disturbance acquisition unit is used to obtain the self-disturbance of the regulatory network node. The external disturbance acquisition unit is used to obtain the external disturbance code of the regulatory network node. The processing unit is connected to the output ends of the self-disturbance acquisition unit and the external disturbance acquisition unit respectively. The processing unit is used to calculate the disturbance score, and compare the disturbance score with the threshold value to determine whether the regulatory network node is a core regulatory factor.
[0082] This scheme uses the regulatory network that exists in specific cells or pathological states. Self-disturbance is based on genomic variation or non-specific (whole genome, such as DNA methylation) modification regulation, and external perturbation is based on specific (only for one or a class of targets) modification regulation relationship. Concordance score 1 is used to define consistency for downstream promotion interaction nodes, and conversely, Concordance score 2 is used to define consistency for downstream inhibition interaction nodes. Only high-risk self-disturbance factors are taken into account in the algorithm, that is, log10 (odds ratio) ≥ 0.5. This algorithm can be developed using R language, has platform compatibility, and is relatively easy to implement.
[0083] The present invention also provides a method for constructing a web-based interactive analysis and visualization platform for molecular regulatory networks, using a web-based online database to provide online interactive analysis and visualization of molecular regulatory networks, and for each node (such as a gene), collect and provide very detailed biological information (sequence, cell localization, tissue expression level, protein quantification, etc.) for easy reference. This method includes the following steps:
[0084] The regulatory relationship of the network depends on specific cells or pathological conditions. For example, there is mutual regulation between genes AB in myocardial cells, but not in liver cells; or there is such regulation in liver cancer, but not in normal liver cells. At the same time, it also depends on the entire network being searched.
[0085] Offline, the regulatory network data is integrated and accessed in a standardized manner to obtain node set files and edge set files and import them into the MySQL (relational database management system) database on the server side. A regulatory network is described by its node set and edge set (regulatory relationships between two nodes), corresponding to the nodes.txt and edges.txt files, which are imported into the MySQL database respectively.
[0086] Preferably, the regulatory network has multiple layout options and can arbitrarily extract the subnetwork of interest (the regulatory network composed of nodes that directly interact with the gene of interest), and the node position can be automatically calculated according to multiple layout algorithms, including random, grid, concentric, breadthfirst or cose. Each layout has its established layout mode, and the established program determines the style of each node in each layout mode. For example, the grid layout presents the network as a matrix. When calculating the position of each node on the screen, it can be determined by obtaining parameters such as the number of nodes, the diameter of the node, the degree of the node, the node spacing, the screen size, etc.
[0087] Software is developed to extract the parameters of the gene subnetwork (including the subnetwork of the gene of interest) on the server side, and to generate the data required for the web-side response and visualization, which can be uniformly communicated in json format. Based on the extracted data, a web-side interface and event response program are developed using a computer language. Preferably, the computer language uses html, css or javascript.
[0088] Use programming languages to develop web-side and server-side communication interfaces for transmitting parameters, etc. Programming language courses use PHP (PHP: Hypertext Preprocessor, a scripting language executed on the server side, especially suitable for Web development and can be embedded in HTML) language, or asp, net language, etc. The generated JSON parameter file is transmitted to the server side and the perl program is called. The execution of the program is completed by PHP code. The server-side Apache (transliterated as Apache, is the world's number one Web server software. It can run on almost all widely used computer platforms. Due to its cross-platform and security, it is widely used and is one of the most popular Web server software) service will interpret the returned PHP code, complete the events defined by the PHP code (i.e. extract the sub-network and generate the result file), and transmit it back to the browser for parsing and visualization. Complete the construction of the visualization platform, which uses the cytoscape.js extension to make the network visualization module compatible with multiple platforms. As long as the offline cytoscape software can be used for visualization, the web-side application will be more convenient.
[0089] In a preferred embodiment of the present invention, based on software using the perl programming language, the method for extracting data on the server side is as follows:
[0090] Read the parameters of the subnetwork to be extracted, including the ID of the gene of interest and the name of the subnetwork. According to the parameters of the subnetwork, read the subnetwork data from the MySQL database. Store the subnetwork data in JSON file format and TXT format file for building the network diagram on the web.
[0091] In a preferred embodiment of the present invention, the method for developing a web interface and an event response program is as follows:
[0092] The structure of the web page is defined using the HTML hypertext markup language, and then the CSS style sheet language is used to render and beautify the style and appearance of the HTML components. The structure of the web page includes the web page logo and title, navigation bar, control network diagram container, network diagram toolbox container, network diagram annotation container, network diagram detailed description container, and detailed parameter containers of nodes, edges, and control axes in the network, documents, and contact information.
[0093] Use javascript to achieve page positioning and jumping, network layout switching and network-related data downloading, search box automatic prompts and sub-network search. Taking sub-network search as an example, its internal process includes parameter collection and storage in variables, receiving parameter variables by PHP language, checking whether each parameter meets the requirements, filling missing parameters with default values and generating a cleaned parameter list, which is transmitted to the server program using Perl language in JSON format;
[0094] Requirements include: ① Gene names use the unified name of HGNC gene symbol, refer to https: / / www.genenames.org / ; ② The subnetwork type belongs to C1-C7 or the 33 tumor types in the Human Cancer Genome Atlas TCGA, refer to https: / / www.cancer.gov / about-nci / organization / ccg / research / structural-genomics / tcga / studied-cancers; ③ Missing parameters are filled with the default value Null.
[0095] The present invention also provides an interactive analysis and visualization platform for building a web-based molecular regulatory network, including a server side, a web side and a processor. The processor executes the method described in the present invention, controls the server side to connect with the web side, and completes the construction of a visualization platform. The core of the platform is developed using HTML, CSS, JavaScript, PHP, and Perl languages. The server-side software is developed for integration, texting the molecular regulatory network, extracting the subnetwork required for web visualization, generating standardized format files, and communicating with the web side. On the web side, based on user operations, interactive analysis and layout of the regulatory network are performed, making the analysis more real-time and convenient, and making the network more concise and beautiful. It can greatly facilitate researchers to analyze and visualize the molecular regulatory network, and retrieve the regulatory network of interest genes and their detailed information in a low-threshold, fast, convenient, and customized manner.
[0096] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0097] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for constructing a high-confidence molecular regulatory network. It is characterized in that The steps include: Determine the environment of the molecular regulatory network to be constructed; Determine the molecular types of all nodes to be included in the regulatory network to be constructed; Based on text mining literature and database mining technology, the regulatory relationships containing target molecule types that have been experimentally verified are integrated to obtain a regulatory relationship list, and the relationship list that matches the target molecule type is extracted from the list; Determine whether the mutual regulatory effects of all nodes of the molecular type to be analyzed exist; Based on the extracted relationship list and inferred results, a high-confidence molecular regulatory network corresponding to the environment is constructed; The environment includes tumor type or subtype; The molecular types include transcription factors, receptors, ligands, miRNAs and lncRNAs; Text mining literature and database mining technology include the following steps: By mining the key areas of literature and databases, we can extract the regulatory relationships that exist therein; Based on artificial intelligence, the existing literature is used for model training to obtain a literature integration mathematical model; Apply the literature integration mathematical model to the new literature set to obtain the desired verified gene regulatory relationships; The regulatory relationship list includes the pairwise regulatory relationships verified by experiments reported in literature, as well as the pairwise regulatory relationships that have been clearly recorded in the database and confirmed to exist; The method for determining whether mutual regulation exists among all nodes of the molecular type to be analyzed is as follows: Calculate the Spearman correlation between two nodes. If the correlation coefficient |rho|≥0.3 and P<0.05, proceed to the next step; The generalized linear regression model was used to calculate the relationship between the two nodes and the phenotype. If P < 0.05, it was considered that there was a regulatory relationship between the two nodes and thus affected the phenotype. Based on the extracted relationship list and inference results, the method for constructing a high-confidence molecular regulatory network corresponding to the environment is: Based on the node criteria and edge criteria, the cell score estimates are classified and divided into tertiles, namely low, medium and high. If more than 66% of the samples in the cluster are classified as "medium" or "high", the node will be retained to a node set; if the consistency score of two connected nodes is greater than the P25 of the consistency score distribution, the edge is retained to obtain an edge set. The regulatory network is composed of its node set and edge set.
2. A computer medium, It is characterized in that The computer medium stores a program for executing the method of claim 1.
Citation Information
Patent Citations
Method for constructing visual platform of molecular regulation and control network and visual platform
CN114283891A