A method for predicting active structures of interfering biological pathways

By screening high-quality transcriptome data and utilizing biological pathway similarity assessment and bootstrapping methods, the active structures of potential compounds are identified, overcoming the shortcomings of existing active structure prediction models and improving the accuracy and reliability of new drug development and environmental toxicant identification.

CN116052772BActive Publication Date: 2026-03-31NANJING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-31
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The lack of quality in existing transcriptomics training sets and prediction methods hinders the development and application of active structure prediction models, making it difficult to effectively identify the interfering effects of compounds on biological pathways.

Method used

By collecting and screening high-quality transcriptome gene expression data, gene set enrichment analysis and biological pathway similarity judgment are used. The degree of biological pathway crossover and consistency of regulatory trends between the training set and the test set are evaluated by combining cumulative hypergeometry and cumulative Bernoulli distribution. The active structure of potential compounds is identified by bootstrapping.

Benefits of technology

It improves the accuracy and reliability of active structure prediction, reduces the false positive rate, and effectively supports the development of new drugs and the identification of key toxins in environmental complex systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052772B_ABST
    Figure CN116052772B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting active structures interfering with biological pathways, which comprises the following steps: constructing a compound biological pathway interference database containing clear labels such as cell lines, exposure time and exposure concentration; evaluating the consistency of the biological pathway cross degree and the regulation trend of the training set and the test set through cumulative hypergeometric distribution and cumulative Bernoulli distribution; identifying a batch of potential compounds in the training set; evaluating the occurrence frequency of the molecular descriptors in the potential compounds through cumulative distribution probability; and finally predicting the potential active structures driving the change of the biological pathway through inputting the biological pathway.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of predicting active structures in drug development and environmental complex systems, and more specifically to a method for predicting active structures that interfere with specific biological pathways based on clearly defined biological pathways and compound relationships. Background Technology

[0002] As an intermediate carrier of genetic information, ribonucleic acid (RNA) participates in a series of important life processes, including protein synthesis and gene expression regulation. The development of next-generation sequencing and gene chip technologies has led to an explosion of RNA transcriptome data, greatly promoting the development of multiple fields such as medicine, life sciences, and environmental sciences. A large amount of RNA transcriptome data exists in cells (in vitro) and individuals (in vivo) under different compound treatments (e.g., drugs and pollutants), such as in the CTD (Comparative Toxicogenomics Database) and ConnectivityMap databases. Generally, compounds with similar main active structures lead to similar biological effects or mechanisms of action, thus their transcriptome results tend to be consistent. Transcriptome data effectively provides multidimensional and highly parallel biological fingerprints, giving scientists the opportunity to develop applications similar to quantitative structure-activity relationship (QSAR), i.e., predicting potential active structures leading to RNA changes by inputting RNA information (differentially expressed genes or biological pathways). Obtaining active structures that significantly contribute to specific biological targets and drug functions will effectively promote the design and development of new drugs, greatly saving time and money. Furthermore, identifying the active structures of key toxicants in environmental mixtures through transcriptomics results will support the identification of key toxicants, the recommendation of controlled substances, and the assessment of human and ecological health risks. However, the insufficient quality of existing transcriptomics training sets and prediction methods severely hinders the development and application of active structure prediction models. Summary of the Invention

[0003] 1. The problem to be solved

[0004] To address the problem of insufficient connection between existing RNA information and active structures, the purpose of this invention is to provide a method for predicting the active structures that interfere with biological pathways.

[0005] 2. Technical Solution

[0006] To solve the above problems, the technical solution adopted by the present invention is as follows:

[0007] A method for predicting the active structure that interferes with biological pathways, comprising the following steps:

[0008] 1) Data collection:

[0009] Collect human transcriptome gene expression data;

[0010] Convert gene names to Entrze IDs;

[0011] The data is categorized using biolabels; these biolabels include cell line, exposure time, exposure concentration, and data quality.

[0012] The data is categorized by chemical tags; these tags include molecular structures described by SMILES and molecular structures described by TOXPRINT.

[0013] Among them, for the collected human transcriptome gene expression data, low-quality data are removed as needed based on gene expression intensity and data parallelism to obtain the retained data; in principle, for the gene chip technology used by ConnectivityMap, the screening criteria of qc_pass equal to 1 and tas value greater than 1.5 are adopted.

[0014] For the second-generation sequencing technology that is more commonly used in transcriptomics, the FastQC software is used for sequencing quality control. The screening criteria mainly include the 10th percentile of the sequencing quality statistics being greater than 20, the peak value of the sequencing quality statistics for each sequence being greater than 20, the G / C ratio in the base distribution being less than 20%, the average GC content distribution of the sequence being less than 30%, the N content of the sequence being less than 20%, and the repetitive sequences accounting for less than 50% of the total.

[0015] 2) Enrichment of biological pathways:

[0016] The transcriptome gene expression data collected and retained in step 1) are divided into a training set and a test set; wherein the data volume of the test set shall not be less than 10%;

[0017] The transcriptome gene expression data collected and retained in step 1) were used to enrich biological pathways using gene set enrichment analysis, and the P values ​​obtained from the gene set enrichment analysis were... GSEA Value and NES value;

[0018] Identify significantly enriched biological pathways;

[0019] 3) Assessment of biological pathway similarity:

[0020] The degree of crossover of significantly enriched biological pathways between the training and test sets is determined by the cumulative hypergeometric distribution.

[0021] By using cumulative Bernoulli distribution, we can determine the consistency of the significantly enriched biological pathway regulatory trends between the training set and the test set.

[0022] 4) Integration of similarity results:

[0023] For the cumulative hypergeometric distribution P hypergeometric The value is corrected for the error detection rate to obtain P. FRD1 value;

[0024] P for cumulative Bernoulli distribution bernoulli The value is corrected for the error detection rate to obtain P. FRD2 value;

[0025] If P FRD1 Value, P FRD2 If all values ​​are less than 0.05, the training set data and the test set data are considered similar in terms of biological pathway interference.

[0026] Based on the connections between biological pathways and compounds in the training set, a batch of potential compounds that could cause interference with biological pathways in the test set were further identified.

[0027] 5) Calculation of active structure:

[0028] Based on the number of potential compounds, data samples are obtained from the training set using a bootstrapping method, and the cumulative probability P of each molecular descriptor of a potential compound in a Gaussian distribution is calculated. distribution P distribution Structures described by molecular descriptors with values ​​less than 0.05 are considered to be active structures in the test set that interfere with specific biological pathways.

[0029] Further, in step 2), biological pathways are enriched by Entrze ID and relative gene expression intensity.

[0030] Further, in step 2), the pathway enrichment conditions include:

[0031] Select a biological pathway database, with 4 to 1000 genes enriched for each pathway, and set the species to human.

[0032] The biological pathway database includes any one or more of KEGG, Reactome, and Wikipathway.

[0033] Furthermore, the P obtained from gene set enrichment analysis GSEA Biological pathways with a value less than 0.05 are considered significantly enriched;

[0034] The P GSEA The value is used as a characteristic of biological pathway interference.

[0035] Furthermore, the overall performance of the biological pathway includes both upregulation and downregulation;

[0036] Positive NES values ​​obtained from the gene set enrichment analysis indicate upregulation;

[0037] A negative NES value obtained from the gene set enrichment analysis indicates downregulation.

[0038] Furthermore, in step 3), the cumulative hypergeometric distribution result is represented by P. hypergeometric Value representation;

[0039] The formula for the cumulative hypergeometric distribution is as follows:

[0040]

[0041] N represents the number of background pathways;

[0042] n is the number of biological pathways significantly enriched in the test set, i.e., P GSEA Value < 0.05;

[0043] m is the number of pathways that are significantly enriched together in the training and test sets;

[0044] M represents the number of pathways significantly enriched in the training set.

[0045] Furthermore, the cumulative Bernoulli distribution result is represented by P. bernoulli Value representation;

[0046] P bernoulli The smaller the value, the higher the consistency of the biological pathway regulatory trends that are significantly enriched in both the training set and the test set, which in turn indicates that the two are more similar in terms of biological pathway interference.

[0047] The cumulative Bernoulli formula is as follows:

[0048]

[0049] m is the number of pathways that are significantly enriched together in the test set and the training set;

[0050] k represents the number of biological pathways with consistent regulatory trends that are significantly enriched in both the training and test sets.

[0051] The p parameter is set to 0.5.

[0052] Furthermore, in step 4), based on the connection between biological pathways and compounds in the training set, a batch of potential compounds that cause interference with biological pathways in the test set are further identified. The principle of majority rule is used to judge the similarity of multiple records of the same compound and the same concentration. If the number of similar and dissimilar records is the same, they are also considered dissimilar. For the difference between records of the same compound and different concentrations, the most similar result is retained.

[0053] Further, in step 5), based on the number of potential compounds, the training set is sampled with replacement and without repetition using a bootstrap method. The number of times each molecular descriptor appears in each sampling is calculated and fitted to a Gaussian distribution.

[0054] The number of samplings shall not be less than 1000;

[0055] Count the occurrence frequency of each molecular descriptor in potential compounds and calculate the cumulative probability P of each molecular descriptor in a Gaussian distribution;

[0056] After removing molecule descriptors with a standard deviation less than 1 in the sampling, the cumulative probability P distribution Molecular descriptors with values ​​less than 0.05 represent structures that are considered to be active structures that interfere with biological pathways in the test set.

[0057] Further, in step 5), the pnorm function in R is used to fit the distribution to a Gaussian distribution.

[0058] 3. Beneficial effects

[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0060] (1) The present invention provides a method for predicting the active structure of a specific biological pathway that interferes with the existing transcriptome data. Based on the consistency of the degree of crossover and regulatory trends of biological pathways in the training set and the test set, the method predicts the unknown active structure through known structural compounds (active structures). It has great application value in the fields of new drug development and identification of key toxins in environmental complex systems.

[0061] (2) The present invention provides a method for predicting the active structure of interference with specific biological pathways. Compared with the current method based on the relationship between compounds and genes, this method integrates RNA change trends and intensities through GSEA biological pathway enrichment without losing information, thus constructing a more reliable and effective training set. In addition, this method explicitly sets labels for exposure concentration, exposure time, and cell line, further enhancing data quality through scientific screening of the data.

[0062] (3) The present invention provides a method for predicting the active structures that interfere with specific biological pathways. This method utilizes cumulative hypergeometric distribution and cumulative Bernoulli distribution to assess the consistency between the training and test sets in terms of the degree of crossover and upregulation / downregulation trends in biological pathways. Through explicit chemical labels in the training set data, a batch of potential compounds that may interfere with biological pathways in the test set are effectively identified from two different dimensions. This method fully considers the characteristics and information reflected by the transcriptome, filling the gap in predicting active structures based on biological pathways.

[0063] (4) The present invention provides a method for predicting the active structure of a specific biological pathway by using a bootstrap sampling method to obtain the background distribution of each molecule descriptor in the training set, and identifying the active structure by evaluating the cumulative probability under the curve. This method treats the obtained set of potential compounds as a low-probability and unique event rather than judging the difference from the background, thus greatly reducing the probability of false positives and improving the prediction results of the active structure compared with the method of judging the active structure by significance test. Attached Figure Description

[0064] Figure 1 Flowchart of a method for predicting the active structure that interferes with a specific biological pathway;

[0065] Figure 2 The distribution of significantly enriched biological pathways in the test and training sets;

[0066] Figure 3 The number of potential compounds in biological pathways in the test and training sets;

[0067] Figure 4 Performance of the test set in a biological pathway-based activity structure prediction model;

[0068] Table 1 shows the predicted active structure results for the test set DOSVAL005_MCF7_24H_BRD-K49675259_10. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention; moreover, the various embodiments are not relatively independent and can be combined with each other as needed to achieve better results. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0070] The method for predicting the active structure that interferes with a specific biological pathway provided by the present invention comprises the following basic steps:

[0071] 1) Data collection:

[0072] Collect transcriptome data from different ethnic groups using the same testing platform (meaning those with the same RNA extraction, reverse transcription, library construction, and sequencing protocols);

[0073] The biotags of cell lines, exposure time, exposure concentration, and test gene sets were classified in detail, and data with low response and poor parallelism were removed. All gene names were represented by independent Entrze IDs.

[0074] Compounds with the SMILES pattern were retained, and a 729-bit TOXPRINT molecular descriptor for each compound was calculated to assign a specific chemical label to each data point.

[0075] As described herein, “TOXPRINT” is a set of publicly available references for invariant structural features designed to cover chemical structures in large toxicity databases and regulatory listings. The TOXPRINT chemistries were developed by Altamira LLC (now part of MN-AM) for the CERES (Chemical Evaluation and Risk Assessment System) program of the U.S. FDA CFSAN.

[0076] As described herein, “SMILES” is an abbreviation for Simplified Molecular Input Line Entry System, which is an ASCII-encoded linear symbol primarily used to represent the structure of compounds in a one-dimensional textual language.

[0077] 2) Enrichment of biological pathways:

[0078] This method does not use hard thresholds to identify differentially expressed genes based on transcriptome results. Instead, it selects all genes' Entropy IDs and relative expression levels, and uses Gene Set Enrichment Analysis (GSEA) to enrich biological pathways and obtain P values. GSEA Value and NES value;

[0079] In addition, GSEA uses multiple biological pathway databases compiled by experts, such as KEGG, Reactome, and Wikipathway, and ensures that each biological pathway contains at least 4 genes and at most 1,000 genes, that is, the number of genes is between 4 and 1,000, and the species is set as human race.

[0080] Furthermore, P GSEA Biological pathways with a value less than 0.05 are considered significantly enriched, while the positive or negative NES value indicates the upregulation or downregulation of biological pathways. These two values ​​provide support for the subsequent calculation of the similarity between the training set and the test set.

[0081] During the model development and validation phase, the transcriptome gene expression data collected and retained in step 1) are divided into a training set and a test set; wherein the data volume of the test set shall not be less than 10%. The training set is used for model construction, and the test set is used for model evaluation. Furthermore, if the test set data is already in pathway format during the method application phase, this step is skipped, and pathway similarity determination is performed directly.

[0082] As described herein, “Gene Set Enrichment Analysis (GSEA)” is an enrichment method developed by the BROAD Institute that can include all genes and their expression status for evaluation without identifying differentially expressed genes, and ultimately determine whether a predefined gene set shows statistically significant and consistent differences between two biological states.

[0083] 3) Assessment of biological pathway similarity:

[0084] The degree of crossover of significantly enriched biological pathways between the training and test sets is determined by the cumulative hypergeometric distribution.

[0085] By using cumulative Bernoulli distribution, we can determine the consistency of the significantly enriched biological pathway regulatory trends between the training set and the test set.

[0086] Furthermore, the degree of crossover between significantly enriched biological pathways in the training and test sets was assessed using cumulative hypergeometric distribution, with P0 as the criterion. hypergeometric The value represents P. hypergeometric A smaller value indicates a higher degree of overlap and a greater number of biological pathways enriched by both the training and test sets, suggesting greater similarity in their biological pathway interference. The formula for the cumulative hypergeometric distribution is as follows:

[0087]

[0088] in,

[0089] N represents the number of background pathways;

[0090] n is the number of biological pathways significantly enriched in the test set, i.e., P GSEA Value < 0.05;

[0091] m is the number of pathways that are significantly enriched together in the training and test sets;

[0092] M represents the number of pathways significantly enriched in the training set.

[0093] Since the upregulation and downregulation of biological pathways are mainly determined by the expression trends and levels of genes within the pathway, the consistency of regulatory trends of biological pathways significantly enriched in both the training and test sets can be assessed using the cumulative Bernoulli distribution. The cumulative Bernoulli results are derived from P... bernoulli The value represents P.bernoulli A smaller value indicates a higher consistency in the regulatory trends of biological pathways that are significantly enriched in both the training and test sets, thus suggesting greater similarity in their interference with biological pathways. The cumulative Bernoulli formula is as follows:

[0094]

[0095] in,

[0096] m is the number of pathways that are significantly enriched together in the test set and the training set;

[0097] k represents the number of biological pathways with consistent regulatory trends that are significantly enriched in both the training and test sets.

[0098] The p parameter is set to 0.5.

[0099] 4) Integration of similarity results:

[0100] The similarity of biological pathways between the training and test sets is defined from two different dimensions: the degree of crossover of significantly enriched biological pathways and regulatory trends.

[0101] For the cumulative hypergeometric distribution P hypergeometric The value is corrected for the false discovery rate (FDR) to obtain P. FRD1 Value; P for cumulative Bernoulli distribution bernoulli The value is corrected for the error detection rate to obtain P. FRD2 value;

[0102] If P FRD1 Value, P FRD2 If all values ​​are less than 0.05, the training set data and the test set data are considered similar in terms of biological pathway interference.

[0103] As mentioned here, the "false discovery rate" (FDR) is a common term in statistics. In transcriptome analysis, it is mainly used in the analysis of differentially expressed genes to control the proportion of false positive results in the final analysis.

[0104] Furthermore, based on the connections between biological pathways and compounds in the training set, a batch of potential compounds that cause interference with biological pathways in the test set were identified. If there are multiple biological pathway similarity judgment results at the same concentration, they are retained according to the majority rule. If the number of similar and dissimilar results is the same, they are also considered dissimilar. However, the interference with biological pathways under different concentrations of exposure may vary greatly. Therefore, the most similar result is used first among different concentrations of a compound to maximize the retention of the similarity of potential compounds between concentrations and batches.

[0105] 5) Calculation of active structure:

[0106] The training set is randomly sampled with replacement and without replacement according to the number of potential compounds. The frequency of each molecular descriptor in each sample is calculated and fitted to a Gaussian distribution. The number of sampling cycles with replacement and without replacement is greater than 1000, depending on the computation time, but should not be less than 1000. Taking 1000 sampling cycles with replacement and without replacement as an example, the frequency of each molecular descriptor in the potential compounds is then counted, and the cumulative probability P of each molecular descriptor in the Gaussian distribution is calculated. distribution .

[0107] After removing low-discrepancy molecular descriptors with a standard deviation less than 1 in 1000 samplings, the cumulative probability P distribution The structures described by molecular descriptors with values ​​less than 0.05 are considered to be active structures that cause interference with biological pathways in the test set.

[0108] Furthermore, the distribution is fitted to a Gaussian distribution using the pnorm function in R.

[0109] Based on the above, it is necessary to further clarify that currently, there are no reports in domestic or international literature and patents on predicting active structures through biological pathways. This method first constructs a database of compounds with clearly labeled biological pathway interference, including cell lines, exposure times, and exposure concentrations. It then assesses the consistency of the degree of biological pathway crossover and regulatory trends between the training and test sets using cumulative hypergeometric and cumulative Bernoulli distributions. Next, it identifies a batch of potential compounds in the training set and uses cumulative distribution probabilities to assess the frequency of molecular descriptors appearing in these potential compounds. Finally, it achieves the prediction of potential active structures driving changes in biological pathways through input.

[0110] Since differences in cell lines, exposure time, and exposure concentration all affect RNA information in different ways, predicting active structures using RNA information requires a high-quality, clearly labeled dataset. However, databases like CTD (Content Detection and Treatment) often lack clear chemical and genetic information, making it difficult to maintain consistent experimental conditions with the test set. Furthermore, current methods for predicting active structures primarily model differentially expressed genes in the transcriptome. This rigid thresholding of differences and the inherent fluctuations in transcriptomics experiments can easily lead to bias in the training set, making reliable predictions difficult. Biological pathways, on the other hand, can effectively integrate all gene-level information, reducing error rates and increasing the reliability of data comparisons. Finally, and most importantly, chemical substances often possess other substructures and active structures. Identifying the active structure when multiple structures coexist remains a challenging area of ​​research and experimentation.

[0111] The method of constructing a Gaussian distribution of the substructures of compounds in the training set by bootstrapping and then calculating the cumulative probability can reduce the false positive frequency caused by significance testing and can serve as a new approach to identify active structures.

[0112] To further understand the content of this invention, a detailed description of the invention will be provided in conjunction with the accompanying drawings and embodiments.

[0113] Example 1

[0114] This invention discloses a method for predicting the active structures that interfere with specific biological pathways. The following examples use human breast cancer cells (MCF7 cell line) exposed for 24 hours as the research object. By inputting specific biological pathways, the method predicts unknown active structures that significantly contribute to pathway interference. Figure 1 ).

[0115] (1) Data Collection and Cleaning: Transcriptome data were collected from ConnectivityMap, the world's largest database of interferon-driven gene expression established by the BROAD Institute. The data was accessed on June 1, 2022, and contained approximately 720,216 records. In addition to gene expression data, chemical tags (729-bit TOXPRINT molecular descriptors calculated by SMILES and Chemotyper) and biological tags (cell line, exposure time, exposure concentration, and data quality) were also collected. During the data cleaning process, more reliable data with qc_pass equal to 1 and tas values ​​greater than 1.5 were retained. The exposure time was also set to 24 hours according to the conditions of the study subjects, and the cell line was MCF7. Furthermore, all gene names were converted to Entrez IDs according to the gene name mapping dictionary for different ethnic groups.

[0116] (2) Biological pathway enrichment: The gene expression levels compiled in step 1) were individually enriched using GSEA using the R package clusterProfiler. GSEA used expert-compiled biological pathway databases such as KEGG, Reactome, and Wikipathway, ensuring that each biological pathway contained 4–1000 genes, and the species was set to human. GSEA Biological pathways with a value less than 0.05 were considered significantly enriched. The positive or negative NES value represents the upregulation or downregulation trend of the pathway, and these two values ​​are used as features for subsequent biological pathway similarity assessment. Finally, the data processed in step 1) were randomly divided into a training set (12416 entries) and a test set (1380 entries) in a 9:1 ratio. The Kolmogorov-Smirnov test for the number of significantly enriched biological pathways in the training and test sets showed a p-value. KS A value of 0.2054 indicates that there is no difference in the data distribution between the training and test sets. Figure 2 The training and test sets are well-divided.

[0117] (3) Biological pathway similarity assessment: The R language `phyper` function is called, and the cumulative hypergeometric distribution P is used. hypergeometric The crossover of biological pathways significantly enriched in the training and test sets is evaluated. The pbinom function in R is used to calculate the cumulative Bernoulli distribution of the NES values ​​of biological pathways significantly enriched in both the training and test sets, which shows the consistency between positive and negative values. Since the cumulative Bernoulli distribution only has two cases: consistent upward and downward trends, the probability parameter p is set to 0.5 here.

[0118] (4) Integration of similarity results: The FDR method is used to integrate the P-values ​​of the cumulative hypergeometric distribution. hypergeometric The value is corrected for the error detection rate to obtain P. FRD1 The value of P is obtained by using the FDR method for the cumulative Bernoulli distribution. bernoulli The value is corrected for the error detection rate to obtain P. FRD2 Value; for P FRD1 and P FRD2 Training sets with both values ​​less than 0.05 are considered similar to the test set in terms of biological pathway interference. Based on the connections between biological pathways and compounds in the training set, a batch of potential compounds causing biological pathway interference in the test set were further identified. The principle of majority rule was used to assess the similarity of multiple records for the same compound and the same concentration; if the number of similar and dissimilar records was the same, they were also considered dissimilar. For records of the same compound but different concentrations, the most similar result was retained. The number of potential compounds is shown below. Figure 3 .

[0119] (5) Identification of active structures: If more than 10 potential compounds causing interference with biological pathways in the test set are identified, active structures are further extracted. Using a bootstrap method, 1000 random samples with replacement are taken from the training set according to the number of potential compounds. The frequency of each molecular descriptor in each sample is calculated, and the distribution is fitted to a Gaussian distribution using the pnorm function in R. The frequency of each molecular descriptor in the potential compounds is then counted, and the cumulative probability of each molecular descriptor in the Gaussian distribution is calculated. After removing low-discrepancy molecular structures with a standard deviation less than 1 in the 1000 samples, the cumulative probability P is calculated. distributionMolecular descriptors with a probability less than 0.05 are considered active structures that interfere with biological pathways in the training set. The results are presented using DOSVAL005_MCF7_24H_BRD-K49675259_10 from 1380 test data points. This transcriptome data reflects gene expression in MCF7 cells exposed to 10 μM of the compound BRD-K49675259 for 24 hours. Through this method, 330 molecular descriptors met the criteria (standard deviation greater than 1 in 1000 bootstrapping), 24 had a cumulative probability less than or equal to 0.05, and 7 had a cumulative probability less than or equal to 0.01. The top-ranking molecular descriptors were mainly related to C, O, and N elements, with 4 of them being descriptors of the compound BRD-K49675259. Therefore, the active structures in this test set were effectively predicted (Table 1).

[0120] Table 1. Predicted active structure results of test set DOSVAL005_MCF7_24H_BRD-K49675259_10

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129] Further testing of all data in the test set revealed that 1134 (82%) out of 1380 test set data points contained at least one active structure, while 246 did not contain any active structure. Of the 246 data points where no active structure was identified, 60 were due to having fewer than 10 potential compounds, failing to meet the active structure extraction criteria. Therefore, the active structure prediction model based on biological pathway interference in this invention achieved an accuracy rate of 86%. Figure 4 ).

Claims

1. A method of predicting an active structure of a biological pathway interfering, characterized by, The method comprises the following steps: 1) data collection: Collect and retain human race transcriptome gene expression data; Convert gene names to Entrze IDs; Classify the data by biological labels; the biological labels include cell lines, exposure time, exposure concentration, and data quality; Classify the data by chemical labels; the chemical labels include SMILES described molecular structures and TOXPRINT described molecular structures; 2) biological pathway enrichment: Divide the transcriptome gene expression data collected and retained in step 1) into a training set and a test set; wherein the data amount of the test set accounts for no less than 10%; For the transcriptome gene expression data collected and retained in step 1), biological pathways were enriched using gene set enrichment analysis resulting in P GSEA values and NES values; Identify significantly enriched biological pathways; 3) biological pathway similarity judgment: By accumulating hypergeometric distribution, the degree of intersection of significantly enriched biological pathways between the training set and the test set is judged; By accumulating Bernoulli distribution, the consistency of the significantly enriched biological pathway regulation trend between the training set and the test set is judged; 4) integration of similarity results: The cumulative hypergeometric distribution is given by P hypergeometric The false discovery rate is corrected for values P FRD1 values; A correction for false discovery rate is applied to the cumulative binomial distribution values P bernoulli to obtain values P FRD2 ; If P FRD1 values, P FRD2 values are all less than 0.05, it is considered that the training set data and the test set data are similar in biological pathway interference. Based on the association between biological pathways and compounds in the training set, a batch of potential compounds causing interference of biological pathways in the test set are further identified; The cumulative hypergeometric distribution result is expressed as P hypergeometric values; Wherein, the formula of accumulating hypergeometric distribution is as follows: N is the number of background pathways; n is the number of biological pathways significantly enriched in the test set, i.e. P GSEA p-value < 0.05; m is the number of pathways that are significantly enriched in both the training set and the test set; M is the number of pathways that are significantly enriched in the training set; The cumulative Bernoulli distribution result is represented by P bernoulli values; Wherein, the formula of accumulating Bernoulli is as follows: m is the number of pathways that are significantly enriched in both the training set and the test set; Number of pathways with consistent regulation trend for both training and test sets; Parameter set to 0.5; 5) active structure calculation: According to the number of potential compounds, the number of each molecular descriptor in each sampling is calculated by bootstrap sampling without replacement from the training set, and is fitted into a Gaussian distribution. Wherein, the number of samplings is not less than 1000 times. counting the occurrence times of each molecular descriptor in the potential compounds and calculating the cumulative probability of each molecular descriptor in the Gaussian distribution P; Cumulative probability after removing molecular descriptors with standard deviation less than 1 in the sampling P distribution Structures represented by molecular descriptors with values less than 0.05 were considered active structures responsible for the biological pathway perturbation in the test set.

2. The method of predicting an active structure of a biological pathway according to claim 1, wherein, In step 2), biological pathway enrichment is performed by Entrze ID and gene relative expression intensity. 3.The method of claim 2, wherein, In step 2), the pathway enrichment conditions include: Select a biological pathway database, and the number of genes enriched in each pathway is 4-1000, and the species is set to human species; Wherein, the biological pathway database includes any one or more of KEGG, Reactome and Wikipathway.

4. The method of predicting an active structure of a biological pathway according to claim 3, wherein, Geneset Enrichment Analysis P GSEA Biological pathways with a value less than 0.05 were considered significantly enriched. The P GSEA Values are used as signatures of biological pathway perturbations.

5. The method of claim 3, wherein the method is characterized by, The overall performance of the biological pathway includes up-regulation and down-regulation; The positive value of NES obtained by gene set enrichment analysis indicates up-regulation; The negative value of NES obtained by gene set enrichment analysis indicates down-regulation.

6. The method of predicting an active structure of a biological pathway according to any one of claims 1 to 5, wherein In step 4), based on the association between biological pathways and compounds in the training set, a batch of potential compounds causing interference of biological pathways in the test set are further identified, and the principle of minority serving majority is used for similarity judgment of multiple records of the same compound and the same concentration, and if the number of similar and dissimilar records is consistent, it is also considered to be dissimilar; And for records of the same compound and different concentrations, the most similar result is retained.

7. The method of predicting an active structure of a biological pathway according to claim 6, wherein, In step 5), the R language pnorm function is used to fit the Gaussian distribution.

Citation Information

Patent Citations

  • Prediction method for quantitative activity of endocrine disrupter

    CN119380859A