Screening method and device of compound toxicity antagonist, terminal equipment and storage medium
By constructing a key target gene set and an analysis network of adverse results, identifying the core gene set of compounds, solving the problem of unclear toxicity mechanism of compounds and improving the screening efficiency of compound toxic antagonists.
Patent Information
- Application Number
- CN202510459092.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, the toxicity mechanism of the compound is unclear, especially the lack of systematic analysis of complex multi-target toxicity, resulting in inefficient screening of compound toxic antagonists.
By constructing a key target gene set, using a bad result analysis network for enrichment analysis, identifying key enrichment events, and constructing a target adverse result analysis subnet of target compounds, identifying core key events and core gene sets, and then determining the toxic antagonist of the compound.
The rapid localization of core gene sets of multi-target toxic compounds is achieved, and the screening efficiency of compound toxic antagonists is improved.
Smart Images

Figure CN120452597A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of toxicology, and in particular to a method, apparatus, terminal device and storage medium for screening compound toxicity antagonists. Background Art
[0002] As the impact of chemical pollutants on human health becomes increasingly serious, current techniques for analyzing toxicity mechanisms and screening for antagonists to these effects rely primarily on researchers' prior experience and knowledge, as well as traditional in vitro or in vivo experimental methods. In existing technologies, the adverse outcome pathway (AOP), as a multi-scale framework, can be used to describe the toxic effects of chemicals. It can systematically organize and evaluate the connections between molecular initiating events (MIEs) and key events (KEs) at different biological levels, all the way to adverse outcomes (AOs) at the macro level.
[0003] The AOP method can systematically evaluate the toxicity mechanisms of a class of compounds. By integrating multi-omics data such as transcriptomes and proteomes, the toxic effect targets of the compounds can be identified. Large-scale in vitro cell viability experiments can preliminarily screen for toxic effect antagonists. In the field of drug repositioning, the CMAP database based on the principle of transcriptome feature reversal can successfully achieve drug screening. This method assumes that if a drug can reverse the disease-related gene expression pattern to the gene expression pattern of a healthy state, then the drug may have a therapeutic effect on the disease.
[0004] However, due to the lack of clarity regarding the mechanisms of chemical toxicity, particularly the lack of systematic analysis of complex multi-target toxicity, it is difficult to precisely locate the toxic targets that compounds induce disease, specifically the core target genes that interact with the compounds and thus induce disease. Therefore, relying solely on the manual establishment of AOPs by experts in complex toxic effect mechanisms is time-consuming and labor-intensive, resulting in currently inefficient screening for compound toxicity antagonists. Summary of the Invention
[0005] The embodiments of the present invention provide a method, apparatus, terminal device and storage medium for screening compound toxicity antagonists, which can quickly locate the core gene set of current complex multi-target toxic compounds, thereby effectively improving the screening efficiency of compound toxicity antagonists.
[0006] One embodiment of the present invention provides a method for screening compound toxicity antagonists, comprising:
[0007] Acquire a plurality of key target genes, and construct a key target gene set based on the plurality of key target genes; wherein the key target genes include: a first key target gene that acts on the target compound, and a second key target gene related to the target disease caused by the target compound;
[0008] Taking the target disease as an adverse outcome, an enrichment analysis is performed on a pre-constructed adverse outcome analysis network based on the key target gene set, and a plurality of enriched key events enriched with the plurality of key target genes are determined from the adverse outcome analysis network; wherein the adverse outcome analysis network is a network formed by integrating a plurality of adverse outcome analysis pathways based on the same key events in a plurality of different adverse outcome analysis pathways;
[0009] Extracting several related key events associated with the enriched key events from the adverse outcome analysis network, and taking the enriched key events and the related key events as target key events to construct a target adverse outcome analysis subnetwork for the target compound;
[0010] Calculating the degree of each target key event in the target adverse outcome analysis subnetwork, taking the target key event corresponding to the largest degree as the core key event, and determining the core gene set in the core key event;
[0011] Based on the core gene set, antagonists of the toxicity of the target compound are determined.
[0012] Furthermore, the acquisition of several key target genes includes:
[0013] Obtain a preset chemical disease gene association database;
[0014] A first key target gene associated with the target compound is extracted from the chemical gene association file in the chemical disease gene association database, and a second key target gene associated with the target disease is extracted from the disease gene association file in the chemical disease gene association database.
[0015] Furthermore, the method for screening a compound toxicity antagonist described in the above embodiment further comprises:
[0016] When it is determined that the first key target gene and the second key target gene do not exist in the chemical disease gene association database, obtaining transcriptome data of the target compound;
[0017] identifying differentially expressed genes induced by the target compound based on the transcriptome data, and generating a plurality of gene expression profile matrices;
[0018] A plurality of the gene expression profile matrices are merged to generate a disordered gene expression module of the target compound, and the differentially expressed genes in the disordered gene expression module are used to generate a key gene set for the compound.
[0019] Furthermore, after generating the key gene set of the compound, the method further includes:
[0020] Obtain a preset protein interaction network;
[0021] Mapping the key gene set to the protein interaction network to construct a dysregulated protein regulatory network;
[0022] Clustering the disordered protein regulatory network using a greedy algorithm to identify key functional modules in the disordered protein regulatory network;
[0023] By analyzing the network topology, the node degrees of differentially expressed genes in key functional modules are calculated, and a related core gene set is generated based on a number of target differentially expressed genes whose node degrees are greater than a preset threshold;
[0024] Based on the relevant core gene set, toxicity antagonists of the target compound are determined.
[0025] Furthermore, the construction of the adverse outcome analysis network includes:
[0026] Obtaining a plurality of adverse outcome pathways; wherein the adverse outcome pathways include: molecular initiation events, key events, key event relationships, and adverse outcomes;
[0027] The same key event is used as the intersection node of each of the adverse result paths, and the key event relationship is used as the directed edge connecting the key events. Several of the adverse result paths are integrated to generate an adverse result analysis network.
[0028] Furthermore, the core gene set includes: a first up-regulated gene whose regulatory relationship with the target compound is up-regulated, and a first down-regulated gene whose regulatory relationship with the target compound is down-regulated;
[0029] Determining the toxicity antagonist of the target compound based on the core gene set includes:
[0030] Obtain gene regulation tables for chemical compounds and traditional Chinese medicine molecules;
[0031] Screening the first up-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a first compound and a first traditional Chinese medicine molecule that have a down-regulation relationship with the first up-regulated gene;
[0032] Screening the first down-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a second compound and a second traditional Chinese medicine molecule that have an up-regulated regulatory relationship with the first down-regulated gene;
[0033] The first compound, the first traditional Chinese medicine molecule, the second compound, and the second traditional Chinese medicine molecule are used as toxicity antagonists of the target compound.
[0034] Furthermore, the relevant core gene set includes: a second up-regulated gene whose regulatory relationship with the target compound is up-regulated, and a second down-regulated gene whose regulatory relationship with the target compound is down-regulated;
[0035] Determining the toxicity antagonist of the target compound based on the relevant core gene set includes:
[0036] Obtain gene regulation tables for chemical compounds and traditional Chinese medicine molecules;
[0037] Screening the second up-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a third compound and a third traditional Chinese medicine molecule that have a down-regulation regulatory relationship with the second up-regulated gene;
[0038] Screening the second down-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a fourth compound and a fourth traditional Chinese medicine molecule that have an up-regulated regulatory relationship with the second down-regulated gene;
[0039] The third compound, the third traditional Chinese medicine molecule, the fourth compound, and the fourth traditional Chinese medicine molecule are used as toxicity antagonists of the target compound.
[0040] Another embodiment of the present invention provides a device for screening compound toxicity antagonists, comprising:
[0041] A key target gene acquisition module is used to acquire a number of key target genes and construct a key target gene set based on the key target genes; wherein the key target genes include: a first key target gene that acts on the target compound, and a second key target gene related to the target disease caused by the target compound;
[0042] An enrichment analysis module is configured to perform enrichment analysis on a pre-constructed adverse outcome analysis network based on the key target gene set and taking the target disease as the adverse outcome, and to determine from the adverse outcome analysis network a number of enriched key events that are enriched with a number of key target genes; wherein the adverse outcome analysis network is a network formed by integrating a number of adverse outcome analysis pathways based on the same key events in a number of different adverse outcome analysis pathways;
[0043] a subnetwork construction module, configured to extract several related key events associated with the enriched key events from the adverse outcome analysis network, and use the enriched key events and the related key events as target key events to construct a target adverse outcome analysis subnetwork for the target compound;
[0044] a gene set determination module, configured to calculate the degree of each target key event in the target adverse outcome analysis subnetwork, take the target key event corresponding to the largest degree as the core key event, and determine the core gene set in the core key event;
[0045] The antagonist screening module is used to determine the toxicity antagonist of the target compound based on the core gene set.
[0046] Another embodiment of the present invention provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for screening a compound toxicity antagonist as described in any one of the embodiments is implemented.
[0047] Another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute a method for screening compound toxicity antagonists as described in any of the above embodiments.
[0048] The following beneficial effects are achieved by implementing the present invention:
[0049] The present invention discloses a screening method, device, terminal device and storage medium for a compound toxicity antagonist. The method constructs a key target gene set by integrating a first key target gene that can act on a target compound and a second key target gene related to a target disease, and then uses the target disease as an adverse result. An enrichment analysis is performed in a pre-constructed adverse result analysis network, and several enrichment key events enriched in several key target genes are determined from the adverse result analysis network, that is, enrichment key events that may be related to the target compound inducing the target disease are identified in the adverse result analysis network. In order to prevent the potential toxicity of the compound from being missed, the relevant key events associated with the enrichment key events are taken as target key events, a target adverse result analysis subnetwork of the target compound is constructed, and the core key event corresponding to the largest degree in the target adverse result analysis subnetwork is identified. It can be understood that the core key event is the gene change that must be carried out for the target compound to induce the target disease. Then, according to the core gene set that constitutes the core key event, the toxicity antagonist of the target compound can be accurately determined. Therefore, the present invention can quickly locate the core gene set of the current complex multi-target toxic compound, thereby effectively improving the screening efficiency of the compound toxicity antagonist. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 The figure is a flow chart of a method for screening compound toxicity antagonists provided in one embodiment of the present invention.
[0051] Figure 2 Schematic diagram of the structure of a device for screening compound toxicity antagonists provided in one embodiment of the present invention.
[0052] Figure 3 It is a structural diagram of an adverse result analysis network provided by an embodiment of the present invention.
[0053] Figure 4 Schematic diagram of enrichment analysis of a key target gene set provided in one embodiment of the present invention.
[0054] Figure 5 2 is a schematic diagram of the structure of a target adverse result analysis subnetwork provided by an embodiment of the present invention.
[0055] Figure 6 1 is a schematic diagram of a SATN online platform page provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0056] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned figure descriptions are intended to cover non-exclusive inclusions.
[0058] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.
[0059] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0060] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0061] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0062] In the description of the embodiments of the present application, unless otherwise expressly specified or limited, technical terms such as "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; internal connections between two components or interactions between two components. Those skilled in the art can understand the specific meanings of the above terms in the embodiments of the present application based on specific circumstances.
[0063] See also Figure 1 , is a schematic flow chart of a method for screening compound toxicity antagonists provided in one embodiment of the present invention, comprising:
[0064] S1. Obtaining a plurality of key target genes and constructing a key target gene set based on the plurality of key target genes; wherein the key target genes include: a first key target gene that acts on a target compound, and a second key target gene that is associated with a target disease caused by the target compound;
[0065] Preferably, the step of obtaining several key target genes includes:
[0066] S11. Obtaining a preset chemical disease gene association database;
[0067] S12. Extracting a first key target gene associated with the target compound from the chemical gene association file in the chemical disease gene association database, and extracting a second key target gene associated with the target disease from the disease gene association file in the chemical disease gene association database.
[0068] In a preferred embodiment of the present invention, when determining the target compound and target disease or specific toxicity to be studied, the associated genes (CGs) and their association patterns related to the chemical are extracted from the chemical-gene association file of the existing CTD database. If a certain chemical-gene association has multiple association patterns in opposite directions, the association is eliminated. At the same time, specific disease-related genes (DGs) are obtained from the disease-gene association file of CTD and the DisGeNET database, and the gene sets from these two sources are taken as a union. Finally, a complete set of key target genes (KTGs) is constructed by taking the intersection of CGs and DGs.
[0069] Specifically, taking the identification of male reproductive toxicity antagonists induced by the environmental pollutant di(2-ethylhexyl) phthalate (DEHP) as an example, in order to comprehensively explore the disturbed genes in male reproductive toxicity (MRT) induced by di(2-ethylhexyl) phthalate (DEHP) and its metabolite mono(2-ethylhexyl) phthalate (MEHP), the following genes were obtained from the CTD and DisGeNET databases based on the chemical-gene-disease association. Figure 4 KTGs shown in A. First, 22 disease terms related to DEHP or MEHP were screened from 94 disease terms related to the male reproductive system. A total of 552 KTGs were obtained, such as Figure 4 As shown in C, 340 of them are up-regulated genes and 212 are down-regulated genes. Figure 4 KEG pathway enrichment analysis was performed on the 552 KTGs shown in B, and 40 significantly enriched pathways were found (adjusted P < 0.05), including cell apoptosis, steroid hormone synthesis, p53 signaling pathway, MAPK signaling pathway, and cell cycle. Figure 4 D shows the top 20 most significant pathways. IPA analysis showed that the apoptosis signaling pathway (z value = 2.138) and p53 signaling pathway (z value = 2.121) were significantly activated, while the cell cycle (z value = -0.816) was significantly inhibited ( Figure 4 D). GO enrichment analysis was then performed, and a total of 196 significantly enriched GO terms were identified (p < 0.05, Table S7), among which KTGs were significantly associated with cell apoptosis and proliferation ( Figure 4 E).
[0070] Preferably, the method for screening a compound toxicity antagonist described in the above embodiment further comprises:
[0071] S13. When it is determined that the first key target gene and the second key target gene do not exist in the chemical disease gene association database, obtaining transcriptome data of the target compound;
[0072] S14. Identifying differentially expressed genes induced by the target compound based on the transcriptome data, and generating a plurality of gene expression profile matrices;
[0073] S15. Merging several of the gene expression profile matrices to generate a disordered gene expression module of the target compound, and generating a key gene set for the compound using the differentially expressed genes in the disordered gene expression module.
[0074] In a preferred embodiment of the present invention, if the first key target gene and the second key target gene do not exist in the CTD database, they can be obtained by compound-induced dysregulated transcriptional profile data. First, relevant transcriptome data are obtained from open source databases such as GEO, and limma or DESeq2 tools are used to identify gene expression profile matrices for gene chip sequencing data and second-generation sequencing data according to different sequencing type tissue sources. When the source contains multiple data sets, the RRA (Robust Rank Aggregation) algorithm is used to integrate the data sets. Weighted co-expression network analysis is used to obtain dysregulated gene expression modules that may have potential specific functions. The genes in the module are designated as key gene sets related to compound-specific toxicity.
[0075] S2. Using the target disease as an adverse outcome, and based on the key target gene set, performing enrichment analysis on a pre-constructed adverse outcome analysis network, determining from the adverse outcome analysis network several enriched key events that are enriched with several key target genes; wherein the adverse outcome analysis network is a network formed by integrating several different adverse outcome analysis pathways based on the same key events in the pathways;
[0076] In a preferred embodiment of the present invention, the clusterProfiler v3.12.0R package was used to perform enrichment analysis between the key target gene (KTG) set and the KE. The raw P values were adjusted for multiple comparisons using the Benjamini–Hochberg method. Only enriched KEs (enriched key events) with P < 0.05 were displayed. Each enriched KE was considered significant if at least three key target genes were enriched and the adjusted P value was less than 0.05. The activation (> 0) or inhibition (< 0) of the KE was predicted by the NES score.
[0077] Preferably, the construction of the adverse outcome analysis network includes:
[0078] S21. Obtaining a plurality of adverse outcome pathways; wherein the adverse outcome pathways include: molecular initiating events, key events, key event relationships, and adverse outcomes;
[0079] S22. Using the same key event as the intersection node of each of the adverse result paths, and using the key event relationship as the directed edge connecting the key events, a plurality of the adverse result paths are integrated to generate an adverse result analysis network.
[0080] In a preferred embodiment of the present invention, an XML file was downloaded from the AOP-Wiki database and parsed using the xml2 R package to extract information on all adverse outcome pathways (AOPs), molecular initiating events (MIEs), key events (KEs), key event relationships (KERs), and adverse outcomes (AOs). The specific steps are as follows:
[0081] The ID, title, OECD status, and SAAOP status of each AOP were obtained, and each KE was annotated, including its ID, title, abbreviation, and biological organization level. The type of each KE was determined and classified as MIE, KE, or AO. To ensure high confidence in the AOP network, AOPs with an SAAOP status of "archived" were deleted, and defective AOPs lacking any of the MIE, KE, or AO elements were removed. Each AOP was modeled as a directed graph, with KEs as nodes and KERs as directed edges, connecting upstream and downstream KEs through KERs. The AOP network was constructed using the igraph R package and visualized using Cytoscape software, ultimately generating an AOP network connected by KEs and KERs.
[0082] Furthermore, to establish associations between KEs and genes within the AOP, a KE gene set annotation system was constructed. First, the associations between known KEs and Gene Ontology Biological Processes (GOBPs) were determined based on the key event composition file "aop_ke_ec.tsv" in the AOPWiki database. Second, semantic analysis was used to calculate the similarity between KEs and pathway or GOBP entries to annotate the KEs. Specifically, the original text of the KE, pathway, and GOBP term names was converted to uppercase. The textreuse tool was used to segment the sentences and remove all punctuation. Furthermore, all abbreviations were converted to full spellings. Multiple words in the original text that expressed the same meaning were replaced with a single word, or frequently occurring unified terms describing action modes were removed to improve matching quality. Finally, the most common words in the original text that did not add additional meaning, such as "in," "of," "to," "the," and "from," were removed. Finally, single or multiple words expressing the same meaning were mapped to a unified standard vocabulary. KEs were created by calculating the Jaccard similarity coefficient between pathway and GO BP terms. KE annotations with a similarity coefficient greater than 0.5 were retained, while KEs with a similarity coefficient less than 0.5 were annotated as "missing." Entities were sorted in descending order by similarity coefficient for different entity types. The top three best matches were retained if the similarity coefficient was greater than 0.5, while KEs with a similarity coefficient less than 0.5 were annotated as "missing."
[0083] Furthermore, the KE gene set annotation system also includes a manual curation function, whereby the best matches are manually evaluated to ensure the accuracy of the automated annotation. For genes annotated as "missing," KE further searches the database to determine if there is a correct match, filling in the gaps and improving the annotation.
[0084] For similar KEs of GO BP terms, we grouped them according to GO clustering and guided the merging of KEs. We clustered GO BPs using the "Rel" algorithm (cutoff = 0.65). Finally, we obtained two KE annotation sets: "GO category-(GO BP term-KE)-gene" and "(pathway-KE)-gene."
[0085] Specifically, in the case of identifying antagonists of male reproductive toxicity induced by the environmental pollutant di(2-ethylhexyl) phthalate (DEHP), 424 AOPs were extracted from the AOP-Wiki database, and 299 high-confidence AOPs were retained after screening. These AOPs were associated with 1080 key events (KEs) and 1555 key event relationships (KERs) to construct the following: Figure 3 A global AOP network in A. Figure 3 As shown in Figure B, among the 1080 KEs, 206 were designated as molecular initiating events (MIEs) and 156 were designated as adverse outcomes (AOs). Figure 3 As shown in C, from the perspective of biological organization, the number of KEs at the cellular level is the largest (350), while the number of KEs at the population level is the smallest (29). The schematic diagram of obtaining the two types of KE annotation sets of "GO category-(GO BP term-KE)-gene" and "(pathway-KE)-gene" is shown in Figure 3 As shown in D. Figure 3 E is an example diagram of KE annotation of male gonadal development.
[0086] S3, extracting several related key events associated with the enriched key events in the adverse outcome analysis network, and taking the enriched key events and the related key events as target key events to construct a target adverse outcome analysis subnetwork for the target compound;
[0087] In a preferred embodiment of the present invention, the enriched KEs and pathway-associated KEs are mapped to the global AOP network, the association order between key events and the key event types (MIE, KE, AO) are determined, and subgraphs are extracted.
[0088] To make the connections between KEs more comprehensive, we calculated the distances between KE nodes in the directed graph using igraph within the subgraph. A distance of 1 indicates a direct connection between two key events in the directed graph. Nodes with distances of 2 and 3 were also selected as indirect KE connections and potential KERs. The direct and indirect connection graphs were combined, and similar KEs were manually merged to construct the target adverse outcome analysis subnetwork.
[0089] Furthermore, after the target adverse outcome analysis subnetwork is constructed, it is evaluated according to the AOP development guidelines, and the rationality of the KERs is assessed using evidence from the existing literature. This includes: 1) the biological domain to which the AOP applies; 2) the necessity of all KEs; and 3) the evidence supporting all KERs. The importance of KEs and KERs in the AOP is classified as high, medium, or low confidence based on direct, indirect, or uncontradictory evidence. Based on this confidence, the target adverse outcome analysis subnetwork is optimized, for example, by removing low-confidence target key events and their associated pathways.
[0090] Specifically, to construct an AOP network related to male reproductive toxicity (MRT) with di(2-ethylhexyl) phthalate (DEHP) as a stressor, gene-related key events (gKEs) and pathway-related key events (pKEs) were mapped to the global AOP network through semantic similarity analysis. Subnetworks were extracted from the global AOP network through similarity analysis, and 39 direct associations (distance = 1) and 117 indirect associations (distance = 2 and 3) were obtained ( Figure 5 A). KEs representing the same process but with different titles in different AOPs were merged. For example, MIE1278 “ROS formation” in AOP ID 423 was merged with MIE 1115 “Increase, reactive oxygen species” in AOP ID 476. Finally, AOPs related to MRT induced by DEHP as a stressor were obtained (see Figure 5 B).
[0091] This AOP contains 16 KEs and 24 KERs, where the initial event (MIE) is "increased ROS" and the ultimate outcome (AO) is "decreased male reproductive function." This AOP network consists of 16 KEs and 24 KERs, where the MIE is "increased ROS" and the AO is "decreased male reproduction." Four KEs are at the molecular level, six at the cellular level, and four at the organ or tissue level. The KE "increased apoptosis" has the highest degree (degree = 7), indicating that it is an important KE in this AOP.
[0092] The weight of evidence (WoE) for KEs and KERs was assessed according to the OECD guidelines for the development and evaluation of AOPs. Evidence from literature sources was collected for this purpose. The biological plausibility of most KERs in this AOP network was considered high. Twelve KERs were inferred from existing KERs in the AOP directed graph and were found to have high biological plausibility. Sufficient evidence was found to indicate that the eight newly defined KERs also had biological plausibility.
[0093] S4. Calculating the degree of each target key event in the target adverse outcome analysis subnetwork, taking the target key event corresponding to the largest degree as the core key event, and determining the core gene set in the core key event;
[0094] In a preferred embodiment of the present invention, the degree of each KE in the target adverse outcome analysis subnetwork is calculated through network topology analysis, and the one with the largest degree is selected as the core KE. The intersection of the gene set corresponding to the core KE and the KTGs (key target genes) is used as the KEGs. Alternatively, the intersection of the gene set corresponding to the KE of interest and the KTGs is directly designated as the gene set (KEGs) of the core key event of the target AOP network.
[0095] S5. Determine the toxicity antagonist of the target compound based on the core gene set.
[0096] Preferably, the core gene set includes: a first up-regulated gene whose regulatory relationship with the target compound is up-regulated, and a first down-regulated gene whose regulatory relationship with the target compound is down-regulated;
[0097] Determining the toxicity antagonist of the target compound based on the core gene set includes:
[0098] S51. Obtain gene regulation tables of chemical compounds and traditional Chinese medicine molecules;
[0099] S52, screening the first up-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a first compound and a first traditional Chinese medicine molecule that have a down-regulating regulatory relationship with the first up-regulated gene;
[0100] S53, screening the first down-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a second compound and a second traditional Chinese medicine molecule that have an up-regulated regulatory relationship with the first down-regulated gene;
[0101] S54. Use the first compound, the first traditional Chinese medicine molecule, the second compound, and the second traditional Chinese medicine molecule as toxicity antagonists of the target compound.
[0102] In a preferred embodiment of the present invention, compound-gene associations are obtained from drug transcriptomic data from DRKG, GNBR, Drug Targetor v1.21, and DrugBank databases, and the regulatory relationship between the compound and the gene (such as upregulation, downregulation, low expression, and high expression) is determined. The compound ID is uniformly converted to MESH ID, and the gene ID is converted to ENTREZ ID. Compound-gene associations from different sources are merged and duplicates are deleted. If a "compound-gene" association appears multiple times and the regulatory pattern is opposite, if the regulatory pattern with a higher number of occurrences accounts for more than 75% of the total number of occurrences of the association, the regulatory pattern with a higher number of occurrences is retained, otherwise the association is deleted.
[0103] Interactions from the STRING database and with a combined score > 400 were selected for the construction of the PPI network (protein interaction network). Then, the target gene (CTG) of each compound was mapped to the PPI network, and the degree of each node was calculated using the “degree” function in the R package igraphv1.2.2.
[0104] Nodes with upregulated CTGs have a positive degree, while those with downregulated CTGs have a negative degree. The larger the absolute value of a node's degree, the more important the gene. By sorting the nodes in descending degree order, a ranked gene importance list (Chemical Target Gene List, CTGRLS) is generated for each compound, also known as the compound gene regulation table. TCM molecule-gene associations are obtained from TCM molecular transcriptome data in the HERB database. Using the same method as CTGRLS, a ranked gene importance list (Chinese Medicine Target Gene List, CMTGRLS) is generated for each TCM molecule, also known as the TCM molecular gene regulation table.
[0105] The up-regulated KEGs and down-regulated KEGs were used as input to scan CTGRLS and CMTGRLS using the GSEA algorithm (screening criteria: number of overlapping genes > 2, FDR < 0.25). According to the GSEA results, the phenomenon in which up-regulated KEGs were enriched in down-regulated CTGRLS or CMGL (upNES < 0); the phenomenon in which down-regulated KEGs were enriched in up-regulated CTGRLS or CMTGRLS (dnNES > 0) was defined as an "opposite" pattern.
[0106] If the chemical or traditional Chinese medicine corresponding to CTGRLS or CMTGRLS exhibits an "opposite" pattern, it may reverse the KEG expression pattern and be identified as a potential toxicity antagonist. The target normalized enrichment score (opNES) of the opposite pattern compound is defined as: opNES = min(upNES + dnNES). The opNES is negative; a smaller score indicates a greater ability to reverse the expression pattern.
[0107] Specifically, a total of 7,116 CTGRLs were obtained from four databases, encompassing 7,116 chemicals, 19,444 genes, and 269,395 chemical-gene interactions. A further 1,017 CMTGRLs, covering 19,362 protein-coding genes, were obtained from a database of herbal medicines. These data constitute the background data for the Antscreen tool. Lists of upregulated and downregulated genes were submitted separately to Antscreen, and a gene importance ranking list was selected. Then, by clicking the "Submit" button, the predicted candidate antagonists were obtained.
[0108] We targeted the core KEG "Increase, Apoptosis" in the AOP network and screened for antagonists capable of reversing its expression pattern. We entered 33 downregulated KEGs and 98 upregulated KEGs into the AntScreen tool. We screened CTGRL and CMTGRL, ultimately identifying 52 candidate antagonists (FDR < 0.25), of which 9 antagonists were classified as validated, 24 as potential, and 19 as unreported. Figure 6 E shows the top 5 potential antagonists in each group. The top 2 antagonists in the verified group are quercetin (opNES=-3.16) and taurine (opNES=-2.88); the top 2 antagonists in the potential group are methionine (opNES=-3.22) and apigenin (opNES=-3.13); the top 2 antagonists in the unreported group are phlorizin (opNES=-2.83) and oridonin (opNES=-2.74).
[0109] These six antagonists were included in subsequent studies. The functions of these antagonists were then verified using the mouse testicular interstitial (TM3) cell line. After co-treatment with MEHP (400 μM) for 48 hours, it was found that quercetin (1 μM), taurine (50 μM), carbamate (12.5 μM) and phlorizin (1 μM) significantly alleviated MEHP-induced cell apoptosis ( Figure 6 F) This application fully demonstrates the efficiency and accuracy of the method for screening antagonists of toxic effects of the present invention.
[0110] Preferably, after generating the key gene set of the compound, the method further comprises:
[0111] S6. Obtaining a preset protein interaction network;
[0112] S7, mapping the key gene set to the protein interaction network to construct a dysregulated protein regulatory network;
[0113] S8. clustering the disordered protein regulatory network using a greedy algorithm to identify key functional modules in the disordered protein regulatory network;
[0114] S9. Calculate the node degrees of differentially expressed genes in key functional modules through network topology analysis, and generate a related core gene set based on a number of target differentially expressed genes whose node degrees are greater than a preset threshold;
[0115] In a preferred embodiment of the present invention, the protein-protein interaction (PPI) network from the STRING database is introduced, and genes in the dysregulated gene co-expression module from the transcriptome data are mapped to the PPI network, thereby constructing a specific dysregulated protein regulatory network. A greedy algorithm is used to cluster the network and identify gene regulatory network modules. Through network topology analysis, the genes in the dysregulated network are sorted by node degree and used as DTGs.
[0116] S10. Determine the toxicity antagonist of the target compound based on the relevant core gene set.
[0117] Preferably, the relevant core gene set includes: a second up-regulated gene whose regulatory relationship with the target compound is up-regulated, and a second down-regulated gene whose regulatory relationship with the target compound is down-regulated;
[0118] Determining the toxicity antagonist of the target compound based on the relevant core gene set includes:
[0119] S101, obtaining the gene regulation table of chemical compounds and the gene regulation table of traditional Chinese medicine molecules;
[0120] S102, screening the second up-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a third compound and a third traditional Chinese medicine molecule that have a down-regulation relationship with the second up-regulated gene;
[0121] S103, screening the second down-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a fourth compound and a fourth traditional Chinese medicine molecule that have an up-regulated regulatory relationship with the second down-regulated gene;
[0122] S104. Use the third compound, the third traditional Chinese medicine molecule, the fourth compound, and the fourth traditional Chinese medicine molecule as toxicity antagonists of the target compound.
[0123] In a preferred embodiment of the present invention, similarly, up-regulated DTGs and down-regulated DTGs were used as inputs, and the GSEA algorithm was used to scan CTGRLS and CMTGRLS (screening conditions: number of overlapping genes > 2, FDR < 0.25). According to the GSEA results, the phenomenon in which up-regulated KEGs are enriched in down-regulated CTGRLS or CMGLs (upNES < 0); and the phenomenon in which down-regulated KEGs are enriched in up-regulated CTGRLS or CMTGRLS (dnNES > 0) is defined as an "opposite" pattern.
[0124] If the chemical or traditional Chinese medicine corresponding to CTGRLS or CMTGRLS exhibits an "opposite" pattern, it may reverse the KEG expression pattern and be identified as a potential toxicity antagonist. The target normalized enrichment score (opNES) of the opposite pattern compound is defined as: opNES = min(upNES + dnNES). The opNES is negative; a smaller score indicates a greater ability to reverse the expression pattern.
[0125] Furthermore, based on the above-mentioned screening method for compound toxicity antagonists, the SATN online platform was developed, which is a dtGSEA 2.0 algorithm ( Figure 6 A) interactive web platform for discovering antagonists ( Figure 6 B). The SATN platform consists of two main tools: "ChemGene" and "AntScreen" ( Figure 6 C, D). In ChemGene, users can obtain key target genes from CTD or CTD and DisGeNet. The AntScreen tool is designed to help users quickly identify toxic antagonists. The SATN platform contains two main tools: "ChemGene" and "AntScreen". In ChemGene, users can paste custom chemical lists and disease phenotype lists, and define different species (human or mouse) and databases (CTD, or CTD and DisGeNet) as data sources. Then, appropriate thresholds are selected to obtain key target genes, chemical-gene association files, disease-gene association files, and Venn diagrams. The Antscreen tool is designed to help users quickly identify toxic antagonists.
[0126] The present embodiment provides a method for screening compound toxicity antagonists, by constructing a key target gene set for a target disease and a target compound, and then taking the target disease as an adverse result, according to the key target gene set, determining from an adverse result analysis network several enriched key events enriched with several key target genes, and finally taking the enriched key events and the related key events as target key events, constructing a target adverse result analysis subnetwork for the target compound, so that according to the target adverse result analysis subnetwork, the toxic effect target of the target compound is effectively identified, a core gene set is determined, and finally a toxicity antagonist for treating the target disease caused by the target compound is determined. Therefore, the present invention can quickly locate the core gene set of the current complex multi-target toxic compound, thereby effectively improving the screening efficiency of the compound toxicity antagonist.
[0127] See also Figure 2 , is a schematic structural diagram of a screening device for compound toxicity antagonists provided in one embodiment of the present invention, comprising:
[0128] A key target gene acquisition module is used to acquire a number of key target genes and construct a key target gene set based on the key target genes; wherein the key target genes include: a first key target gene that acts on the target compound, and a second key target gene related to the target disease caused by the target compound;
[0129] An enrichment analysis module is configured to perform enrichment analysis on a pre-constructed adverse outcome analysis network based on the key target gene set and taking the target disease as the adverse outcome, and to determine from the adverse outcome analysis network a number of enriched key events that are enriched with a number of key target genes; wherein the adverse outcome analysis network is a network formed by integrating a number of adverse outcome analysis pathways based on the same key events in a number of different adverse outcome analysis pathways;
[0130] a subnetwork construction module, configured to extract several related key events associated with the enriched key events from the adverse outcome analysis network, and use the enriched key events and the related key events as target key events to construct a target adverse outcome analysis subnetwork for the target compound;
[0131] a gene set determination module, configured to calculate the degree of each target key event in the target adverse outcome analysis subnetwork, take the target key event corresponding to the largest degree as the core key event, and determine the core gene set in the core key event;
[0132] The antagonist screening module is used to determine the toxicity antagonist of the target compound based on the core gene set.
[0133] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.
[0134] Those skilled in the art will clearly understand that for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0135] Another preferred embodiment of the present invention provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for screening a compound toxicity antagonist as described in any one of the above embodiments is implemented.
[0136] The terminal device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0137] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.
[0138] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0139] Another preferred embodiment of the present invention provides a storage medium, which is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.
[0140] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for screening compound toxicity antagonists, characterized in that: include: Acquire a plurality of key target genes, and construct a key target gene set based on the plurality of key target genes; wherein the key target genes include: a first key target gene that acts on the target compound, and a second key target gene related to the target disease caused by the target compound; Taking the target disease as an adverse outcome, an enrichment analysis is performed on a pre-constructed adverse outcome analysis network based on the key target gene set, and a plurality of enriched key events enriched with the plurality of key target genes are determined from the adverse outcome analysis network; wherein the adverse outcome analysis network is a network formed by integrating a plurality of adverse outcome analysis pathways based on the same key events in a plurality of different adverse outcome analysis pathways; Extracting several related key events associated with the enriched key events from the adverse outcome analysis network, and taking the enriched key events and the related key events as target key events to construct a target adverse outcome analysis subnetwork for the target compound; Calculating the degree of each target key event in the target adverse outcome analysis subnetwork, taking the target key event corresponding to the largest degree as the core key event, and determining the core gene set in the core key event; Based on the core gene set, antagonists of the toxicity of the target compound are determined.
2. The method for screening a compound toxicity antagonist according to claim 1, characterized in that: The method of obtaining several key target genes includes: Obtain a preset chemical disease gene association database; A first key target gene associated with the target compound is extracted from the chemical gene association file in the chemical disease gene association database, and a second key target gene associated with the target disease is extracted from the disease gene association file in the chemical disease gene association database.
3. The method for screening a compound toxicity antagonist according to claim 2, characterized in that: Also includes: When it is determined that the first key target gene and the second key target gene do not exist in the chemical disease gene association database, obtaining transcriptome data of the target compound; identifying differentially expressed genes induced by the target compound based on the transcriptome data, and generating a plurality of gene expression profile matrices; A plurality of the gene expression profile matrices are merged to generate a disordered gene expression module of the target compound, and the differentially expressed genes in the disordered gene expression module are used to generate a key gene set for the compound.
4. The method for screening compound toxicity antagonists according to claim 3, characterized in that: After generating the key gene set for the compound, it also includes: Obtain a preset protein interaction network; Mapping the key gene set to the protein interaction network to construct a dysregulated protein regulatory network; Clustering the disordered protein regulatory network using a greedy algorithm to identify key functional modules in the disordered protein regulatory network; By analyzing the network topology, the node degrees of differentially expressed genes in key functional modules are calculated, and a related core gene set is generated based on a number of target differentially expressed genes whose node degrees are greater than a preset threshold; Based on the relevant core gene set, toxicity antagonists of the target compound are determined.
5. The method for screening compound toxicity antagonists according to claim 3, characterized in that: The construction of the adverse outcome analysis network includes: Obtaining a plurality of adverse outcome pathways; wherein the adverse outcome pathways include: molecular initiation events, key events, key event relationships, and adverse outcomes; The same key event is used as the intersection node of each of the adverse result paths, and the key event relationship is used as the directed edge connecting the key events. Several of the adverse result paths are integrated to generate an adverse result analysis network.
6. The method for screening compound toxicity antagonists according to claim 5, characterized in that: The core gene set includes: a first up-regulated gene whose regulatory relationship with the target compound is up-regulated, and a first down-regulated gene whose regulatory relationship with the target compound is down-regulated; Determining the toxicity antagonist of the target compound based on the core gene set includes: Obtain gene regulation tables for chemical compounds and traditional Chinese medicine molecules; Screening the first up-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a first compound and a first traditional Chinese medicine molecule that have a down-regulation relationship with the first up-regulated gene; Screening the first down-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a second compound and a second traditional Chinese medicine molecule that have an up-regulated regulatory relationship with the first down-regulated gene; The first compound, the first traditional Chinese medicine molecule, the second compound, and the second traditional Chinese medicine molecule are used as toxicity antagonists of the target compound.
7. The method for screening a compound toxicity antagonist according to claim 4, characterized in that: The relevant core gene set includes: a second up-regulated gene whose regulatory relationship with the target compound is up-regulated, and a second down-regulated gene whose regulatory relationship with the target compound is down-regulated; Determining the toxicity antagonist of the target compound based on the relevant core gene set includes: Obtain gene regulation tables for chemical compounds and traditional Chinese medicine molecules; Screening the second up-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a third compound and a third traditional Chinese medicine molecule that have a down-regulation regulatory relationship with the second up-regulated gene; Screening the second down-regulated gene in the compound gene regulation table and the traditional Chinese medicine molecule gene regulation table to determine a fourth compound and a fourth traditional Chinese medicine molecule that have an up-regulated regulatory relationship with the second down-regulated gene; The third compound, the third traditional Chinese medicine molecule, the fourth compound, and the fourth traditional Chinese medicine molecule are used as toxicity antagonists of the target compound.
8. A screening device for compound toxicity antagonists, characterized in that: include: A key target gene acquisition module is used to acquire a number of key target genes and construct a key target gene set based on the key target genes; wherein the key target genes include: a first key target gene that acts on the target compound, and a second key target gene related to the target disease caused by the target compound; An enrichment analysis module is configured to perform enrichment analysis on a pre-constructed adverse outcome analysis network based on the key target gene set and taking the target disease as the adverse outcome, and to determine from the adverse outcome analysis network a number of enriched key events that are enriched with a number of key target genes; wherein the adverse outcome analysis network is a network formed by integrating a number of adverse outcome analysis pathways based on the same key events in a number of different adverse outcome analysis pathways; a subnetwork construction module, configured to extract several related key events associated with the enriched key events from the adverse outcome analysis network, and use the enriched key events and the related key events as target key events to construct a target adverse outcome analysis subnetwork for the target compound; a gene set determination module, configured to calculate the degree of each target key event in the target adverse outcome analysis subnetwork, take the target key event corresponding to the largest degree as the core key event, and determine the core gene set in the core key event; The antagonist screening module is used to determine the toxicity antagonist of the target compound based on the core gene set.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for screening a compound toxicity antagonist according to any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the method for screening compound toxicity antagonists according to any one of claims 1 to 7.