A method for predicting the quantitative activity of endocrine disruptors

By combining molecular fingerprinting and machine learning, a quantitative prediction model for endocrine disruptors was constructed, which solved the problem that existing models could not quantify the intensity of endocrine disruption and achieved accurate prediction and risk assessment of the endocrine disrupting activity of compounds.

CN119380859BActive Publication Date: 2026-04-17NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2024-11-14
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing endocrine disruptor prediction models lack comprehensive consideration of different types of endocrine disruption effects, cannot provide quantitative information, and have shortcomings in data processing and feature selection, which affect prediction performance and interpretability.

Method used

We employ molecular fingerprinting to extract multi-level structural features of compounds, combine quantitative read-across methods with various machine learning algorithms to construct quantitative prediction models, and investigate the interaction mechanism between endocrine disruptors and nuclear receptors through molecular docking and molecular dynamics simulation techniques, thereby improving the efficiency of quantitative activity prediction of endocrine disruptors.

Benefits of technology

It has enabled accurate quantitative prediction of the intensity and hazard level of endocrine disruption caused by compounds, improved prediction capabilities and risk assessment levels, and provided technical support for safeguarding public health and ecological environment safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380859B_ABST
    Figure CN119380859B_ABST
Patent Text Reader

Abstract

This application discloses a method for predicting the quantitative activity of endocrine disruptors, relating to the field of virtual screening of endocrine disruptors. The method includes: acquiring in vitro experimental data of nuclear receptors and removing duplicate data, as well as removing compound sets that do not contain the simplified molecular linear input canonical (SMILES) representation; using a molecular fingerprinting method to extract the primary, secondary, and tertiary structural features of the compounds. For each compound cluster, a quantitative prediction model based on machine learning or quantitative read-across is constructed to predict the quantitative activity value of the compound. Addressing the low efficiency of existing methods for predicting the quantitative activity of endocrine disruptors, this application extracts multi-level structural features of the compounds and constructs corresponding quantitative prediction models for compound structural clusters of different sizes. Through molecular docking and molecular dynamics simulations, the interaction mechanism between endocrine disruptors and nuclear receptors is studied from a structural biology perspective, thus improving efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual screening of endocrine disruptors, and more specifically, to a method for predicting the quantitative activity of endocrine disruptors. Background Technology

[0002] Endocrine disruptors are exogenous substances that can interfere with the normal function of an organism's endocrine system. They can adversely affect the endocrine system through multiple pathways, leading to various pathological consequences such as hormonal imbalances, reproductive damage, and metabolic abnormalities, seriously threatening human health and ecological environment safety. Traditional research on endocrine disruptors mainly relies on in vitro experiments and animal models, assessing the endocrine-disrupting effects of compounds by detecting changes in relevant biomarkers.

[0003] In recent years, computational toxicology methods, represented by machine learning, have provided a new technical approach for the study of endocrine disruptors. Researchers have attempted to construct qualitative or quantitative predictive models of endocrine disruptors based on compound structures and known activity data, thereby enabling rapid prediction and preliminary screening of the endocrine disrupting potential of unknown compounds. However, existing predictive models of endocrine disruptors still have some limitations: First, most models are built for a single nuclear receptor type, lacking comprehensive consideration of different types of endocrine disrupting effects; second, existing models are mainly qualitative, able to predict whether a compound has endocrine disrupting potential, but unable to provide quantitative information on the intensity or degree of interference, which is crucial for a deeper understanding of endocrine disruption mechanisms and guiding risk management decisions; third, some models still need improvement in data processing, feature selection, and mechanism interpretation, affecting predictive performance and interpretability.

[0004] Therefore, there is an urgent need to develop a comprehensive, quantitative, and interpretable method for predicting endocrine disruptors. This method should be able to integrate experimental data from multiple nuclear receptor pathways, consider the nonlinear relationship between compound structure and activity, and construct a quantitative prediction model to accurately predict the endocrine disrupting intensity and hazard level of compounds. Simultaneously, this method should combine data-driven and mechanism-driven strategies to achieve precise feature extraction and screening at the data level and to reveal the molecular mechanisms of endocrine disrupting effects at the mechanistic level, aiming to significantly improve the predictive ability and risk assessment level of endocrine disruptors, and provide strong technical support for safeguarding public health and ecological environment safety. Summary of the Invention

[0005] 1. Technical problems to be solved

[0006] To address the low efficiency of quantitative activity prediction for endocrine disruptors in existing technologies, this application provides a method for predicting the quantitative activity of endocrine disruptors. It employs molecular fingerprinting to extract multi-level structural features of compounds and classifies them based on tertiary structural features. Building upon this, a strategy combining quantitative read-across methods and multiple machine learning algorithms is used to construct corresponding quantitative prediction models for compound structural clusters of different sizes, reliably predicting the endocrine disrupting activity of unknown compounds. Furthermore, this application introduces molecular docking and molecular dynamics simulation techniques to study the interaction mechanisms between endocrine disruptors and nuclear receptors from a structural biology perspective, thereby improving the efficiency of quantitative activity prediction for endocrine disruptors.

[0007] 2. Technical Solution

[0008] The purpose of this application is achieved through the following technical solution.

[0009] This application provides a method for predicting the quantitative activity of endocrine disruptors, comprising: obtaining in vitro experimental data of nuclear receptors from the ToxCastv3.2 database, and removing duplicate data and compounds represented by the Simplified Molecular Input Line Entry System (SMILES) from the experimental data to obtain a compound information dataset A1; preprocessing the compound information dataset A1 to obtain a quality control compound dataset A2; and using the molecular fingerprinting method to extract the primary, secondary, and tertiary feature structures of the compounds based on the quality control compound dataset A2, wherein: the primary feature structure represents the compound structural fragment that causes the endocrine disrupting effect; the secondary feature structure represents the compound structural fragment used to distinguish whether a compound has an endocrine disrupting effect; and the tertiary feature structure represents the compound structural fragment that distinguishes the type of endocrine disrupting effect of the compound, including agonists, antagonists, and agonist-antagonists.

[0010] Furthermore, this also includes: constructing a compound classification model based on the extracted primary, secondary, and tertiary feature structures to classify compounds in the quality-controlled compound dataset A2 into clusters of compounds with and without endocrine interference effects, as well as those with different interference effect types; specifically, using the extracted primary feature structures, a binary classification model for those with and without endocrine interference effects is constructed; using the extracted secondary feature structures, a multi-classification model for different nuclear receptor types is constructed; using the extracted tertiary feature structures, a multi-classification model for different effect types (agonists, antagonists, and agonist-antagonists) is constructed; the binary and multi-classification models can be constructed based on machine learning algorithms such as Support Vector Machine (SVM), Random Forest (RF), and k-Nearest Neighbors (kNN), and the optimal model is determined through 5-fold cross-validation.

[0011] For each compound cluster, a quantitative prediction model based on machine learning or quantitative read-across is constructed to predict the quantitative activity values of compounds. Among them, when the number of compounds in the compound cluster is greater than a preset threshold (such as 50), a machine learning method is used to construct a regression model based on the extracted molecular fingerprint features. The available machine learning algorithms include support vector regression (SVR), random forest regression (RFR), k-nearest neighbor regression (kNNR), etc. When the number of compounds in the compound cluster is less than or equal to the preset threshold, a quantitative read-across method is adopted. Based on the molecular structure similarity between each test compound and known active compounds, the activity values of the known active compounds are weighted and averaged to obtain the activity prediction value of the test compound. The cross-validation method is used to evaluate the performance of the quantitative prediction model and optimize the parameters to obtain the optimal quantitative prediction model.

[0012] Furthermore, it also includes: selecting representative active compounds from the quality-controlled compound dataset A2. The selection principles include: compounds covering different nuclear receptor types and effect types; selecting high-activity compounds (such as AC50 value lower than 100 nM) and low-activity compounds (such as 100 μM < AC50 < 1 mM); selecting compounds with high chemical structure diversity and representativeness.

[0013] For the representative active compounds, molecular docking and molecular dynamics simulation analysis are carried out to obtain the binding mode and induced conformational change mechanism of the active compounds with nuclear receptors. Obtain the crystal structure of the nuclear receptor corresponding to the selected compounds. When there is no co-crystallization structure of the selected compound and the nuclear receptor, select the apo conformation for docking. Use molecular docking software such as Glide, AutoDock, etc., to obtain the compound-receptor complex structure through semi-flexible docking method. Use molecular dynamics simulation software such as GROMACS, AMBER, etc., to perform molecular dynamics simulation on the docked complex structure, and the simulation duration is not less than 100 ns. Perform clustering analysis on the simulation trajectory to obtain the main clustering centers of the binding mode. Analyze the conformational changes induced by different active compounds in the nuclear receptor, with a focus on the ligand-binding pocket and auxiliary binding regions. Calculate the binding free energy of high-activity and low-activity compounds with the receptor, and analyze the relationship between their activity differences and binding mode and receptor conformational changes. Based on the analysis results, clarify the differences between high-activity and low-activity compounds in terms of binding mode and induced receptor conformational changes, and explain the action mechanism of their activity differences at the molecular level.

[0014] Furthermore, when the number of compounds in a compound cluster exceeds a threshold, a quantitative prediction model for the corresponding compound cluster is constructed using machine learning methods. This includes: standardizing the molecular fingerprint features of the compounds in the corresponding compound cluster using Standard Scaler; scaling the feature data to a standard normal distribution with a mean of 0 and a variance of 1 to eliminate differences in feature dimensions and improve the training effect of the machine learning model; randomly dividing the standardized compound cluster dataset into training and test sets using the train_test_split function; modeling the training set using multiple machine learning regression algorithms, and performing cross-validation and parameter optimization on the model using Grid Search CV to obtain the optimal model for different machine learning regression algorithms; wherein, the multiple machine learning regression algorithms include at least one of linear regression, ridge regression, Lasso regression, elastic network, decision tree, random forest, AdaBoost, gradient boosting, support vector machine, or K nearest neighbors.

[0015] The optimal models of various machine learning regression algorithms are used to make predictions on the test set, and the coefficient of determination (R²) of each optimal model is calculated. The coefficient of determination (R²) is a measure of the goodness of fit of the regression model, reflecting its ability to explain data variation. The value of R² ranges from 0 to 1; the closer to 1, the better the model's fit and prediction performance. Specifically, in this method, for each compound cluster, the R² of the optimal model of each machine learning regression algorithm is calculated using its test set data. The specific steps are as follows: For each sample in the test set, the predicted value is calculated using the optimal model. Calculate the average of the true values ​​of all samples in the test set. Calculate the residual sum of squares (SSres) and the total sum of squares (SStot); substitute them into the R2 calculation formula to obtain the coefficient of determination (R2) of the optimal model on the test set. Repeat the above steps to obtain the R2 values ​​of all optimal models for this compound cluster. Select models with higher R2 values ​​based on a preset threshold (e.g., 0.6), and average their prediction results to obtain a consensus model with more robust and reliable overall prediction performance. Introducing the R2 index and consensus prediction strategy helps to quantitatively evaluate the model's fit and prediction effect from a statistical perspective, improving the objectivity and interpretability of model evaluation. Furthermore, by integrating the prediction results of multiple high-quality models, the robustness of the model can be further improved, reducing the risk of overfitting and underfitting, thus obtaining a final prediction model with better generalization performance. Select the optimal model with a coefficient of determination (R2) greater than the threshold, and perform a weighted average of the prediction results of the selected optimal model with a coefficient of determination (R2) greater than the threshold on the test set to obtain the machine learning consensus prediction model for the corresponding compound cluster, which serves as the quantitative prediction model.

[0016] Furthermore, when the number of compounds in a compound cluster is less than or equal to a threshold, a quantitative prediction model for the corresponding compound cluster is constructed using a quantitative read-across method. This includes: for each test compound in the compound cluster, calculating its molecular structural similarity Sij with all known active compounds within the cluster; calculating the molecular fingerprint using extended connectivity fingerprinting ECFP4 to obtain fingerprint bit vectors Fx and Fy; and calculating the similarity between the two fingerprint vectors using Tanimoto coefficients. Wherein, Sxy represents the molecular structural similarity between the test compound x and the known active compound y, quantifying the structural similarity between the two compounds; x represents the test compound, i.e., the compound whose activity needs to be predicted; y represents the known active compound; Fx represents the molecular fingerprint bit vector of the test compound x, and Fy represents the molecular fingerprint bit vector of the known active compound y; |Fx∩Fy| represents the number of 1s in the two fingerprint vectors; |Fx∪Fy| represents the total number of 1s in the two fingerprint vectors.

[0017] Based on molecular structure similarity, Sij calculates a weighted average of the lgAC50 activity values ​​of known active compounds to obtain the predicted activity value model_ga(pre) of the test compound: Wherein, model_ga(pre) is the predicted model_ga value of the compound, model_ga is the lgAC50 value of the compound, and S is the structural similarity value of the compound. The obtained predicted activity value model_ga(pre) of the test compound is used as the quantitative read-across predicted activity value of the compound, and model_ga(pre) is defined as the "Tanimoto-weighted Activity Coefficient" (TAC) of the compound, as a quantitative indicator to measure the endocrine-disrupting activity of the compound. The larger the TAC value, the stronger the endocrine-disrupting activity of the compound. The TAC values ​​of all test compounds within the compound cluster are calculated to obtain the quantitative read-across prediction results of the compound cluster.

[0018] Furthermore, representative active compounds were selected from the quality control compound dataset A2, including endocrine disrupting compounds belonging to different chemical structure types (such as chemical skeleton, functional group types, substituent properties, etc.) as endocrine disrupting compounds to be studied. The selected compounds should cover the main chemical structure types in the dataset, and compounds with high activity (such as low IC50 or EC50 values) should be selected in each type. Based on known endocrine disrupting compound-nuclear receptor interaction information, the crystal structures of nuclear receptors interacting with the selected endocrine disrupting compounds were obtained from the RCSB protein database. The obtained crystal structures must contain the ligand binding domain (LBD) of the nuclear receptor.

[0019] Molecular docking was employed to simulate the docking of each selected endocrine disrupting compound with the corresponding nuclear receptor crystal structure obtained via LBD. Mainstream docking software such as Autodock, GOLD, and Glide were used, prioritizing docking strategies that considered receptor flexibility. During docking, the LBD active pocket was used as the docking region, and each ligand-receptor pair was run independently at least 10 times, with a different initial conformation each time. After cluster analysis, the conformation cluster with the most binding modes was selected, and its average binding free energy was calculated as the binding free energy of that ligand-receptor pair. Binding free energy refers to the change in the system's free energy when the ligand and receptor form a complex. It reflects the change in thermodynamic stability caused by the ligand binding from the solution environment to a specific site on the receptor protein. The lower the binding free energy (the larger the negative value), the more stable the binding between the ligand and receptor, and the higher the binding affinity. Based on the calculated binding free energy, endocrine-disrupting compounds with binding free energies to the corresponding nuclear receptor LBD below a preset threshold (e.g., -8 kcal / mol) were screened and identified as active compounds with high binding capacity to the nuclear receptor LBD. Combining the screening results of different nuclear receptors, compounds that showed high activity on multiple nuclear receptors were selected as representative active compounds for subsequent analysis.

[0020] Furthermore, representative active compounds were selected from the post-quality control compound dataset A2. Using the AC50 value as the key indicator, dataset A2 was divided into two subsets: high-activity and low-activity. Compounds with AC50 values ​​less than or equal to the threshold were assigned to the high-activity subset, indicating that these compounds produce significant endocrine-disrupting effects even at low concentrations. Compounds with AC50 values ​​greater than the threshold were assigned to the low-activity subset, indicating that these compounds only exhibit significant endocrine-disrupting effects at higher concentrations. The threshold ranged from 0.5 μM to 2 μM. Crystal structure data of nuclear receptors interacting with compounds in the high-activity and low-activity subsets were obtained from the RCSB protein database. The obtained crystal structures must contain the ligand-binding domain (LBD) of the nuclear receptor.

[0021] Using each compound in either the highly active or low-activity subset as a ligand and the corresponding LBD crystal structure of the nuclear receptor as the receptor, ligand-receptor complexes were constructed using molecular docking methods, and the binding free energy ΔG of each complex was calculated. Autodock software was used, prioritizing docking strategies that considered receptor flexibility. Each ligand-receptor pair was run independently 10 times, each run using a different initial conformation. The docking results were clustered, and the conformation cluster with the most binding modes was selected. The binding free energy ΔG of the complex was calculated using methods such as MM-GBSA or MM-PBSA, and the average ΔG was taken as the binding free energy of the ligand-receptor pair. In the highly active compound subset, compounds with a binding free energy ΔG lower than the first threshold (e.g., -10 kcal / mol) for the corresponding nuclear receptor LBD were screened and designated as highly active representative compounds for that nuclear receptor. In the low-activity compound subset, compounds with a binding free energy ΔG higher than the second threshold (e.g., -7 kcal / mol) for the corresponding nuclear receptor LBD were screened and designated as low-activity representative compounds for that nuclear receptor.

[0022] Furthermore, for representative active compounds, molecular docking and kinetic simulation analyses were performed to obtain the binding modes and induced conformational change mechanisms between the active compounds and the nuclear receptors. This included: using the unbound conformation (apo conformation) of the nuclear receptor as the receptor starting structure, assembling the highly active and low-activity representative compounds selected from claim 7 with the corresponding apo conformations of the nuclear receptors to form ligand-receptor complexes. Molecular dynamics (MD) simulations of the ligand-receptor complexes were performed using GROMACS 5.1.2 or later software: the complexes were dissolved in an octahedral water box constructed using the TIP3P water model, and Na+ or Cl- was added to neutralize the net charge of the system; the components of the system were parameterized using the CHARMM27 force field, and the system energy was minimized using the steepest descent method; the system was first pre-equilibrated for 100 ps under the NVT ensemble, and then equilibrium simulations were performed for at least 20 ns under the NPT ensemble, with the temperature maintained at 300 K and the pressure maintained at 1 atmosphere during the simulation.

[0023] The trajectory frames of ligand-receptor complexes obtained from molecular geometry (MD) simulations were analyzed, focusing on the following aspects: The binding modes of highly active and inactive representative compounds with their corresponding endogenous ligands (i.e., the natural ligands of the nuclear receptor) in the LBD were compared, analyzing the binding sites, orientations, and key interacting groups of different active compounds; the number of hydrogen bonds and van der Waals interaction strengths (e.g., contact area) formed between highly active and inactive representative compounds and the nuclear receptor LBD were statistically analyzed, and the contribution of different types of interactions to the stability of ligand-receptor binding was quantitatively assessed; the root mean square deviation (RMSD) of the H12 helix relative to the apo conformation in the trajectory frame was calculated to reveal the conformational rearrangement of H12 induced by ligand binding, thereby inferring the kinetic mechanism of conformational changes induced by active compounds in the nuclear receptor. Molecular graphics software such as PyMOL and VMD were used for visualization analysis of the trajectory frames to visually demonstrate the dynamic process of binding highly active and inactive compounds with the nuclear receptor and the induced conformational changes; quantitative analysis graphs were created using plotting software such as Origin and GraphPadPrism.

[0024] In this application, the apo conformation is used as the receptor initiation structure, which can more realistically simulate the conformational change process induced by ligand binding. In the simulated trajectory analysis, the focus is on the differences in binding modes between highly and less active compounds and endogenous ligands, the type and strength of interactions with the receptor, and the induction of H12 conformational rearrangement. This allows for a deeper understanding of the molecular mechanisms underlying the differences in compound activity from both structural and energy perspectives.

[0025] Furthermore, the acquired experimental data were preprocessed, including: standardizing the raw data using the Z-score method, and identifying and eliminating false positive results that may be caused by factors such as compound cytotoxicity; calculating the Z-score for each compound in the corresponding nuclear receptor experiment, where the Z-score represents the degree of difference between the relative activity value of the compound and the negative control group, and its calculation formula is: Z = (x - μ) / σ, where x is the relative activity value of the compound, and μ and σ are the mean and standard deviation of the relative activity value of the negative control group, respectively; and eliminating compound data with Z-scores less than 3 to exclude potential false positive results where the relative activity value is not significantly different from the background noise.

[0026] Based on the hit-call value, a data quality assessment standard, experimental data were screened in three levels: Level 5: Primary screening data, focusing on compounds with hit-call values ​​other than 0, and further removing low-confidence data with model_prob values ​​other than 1; Level 6: Secondary validation data, focusing on high-quality compound data with hit-call values ​​other than 0 and flags not exceeding 3; Level 7: Dose-response confirmation data, for compounds with hit-call values ​​of 0, removing data with hit_pct values ​​other than 0, retaining only completely inactive compounds. Here, model_prob represents the confidence score given by the compound activity prediction model, flags represent the number of data quality issue markers, and hit_pct represents the percentage of experimental data judged as active out of the total experimental data.

[0027] Furthermore, in vitro experimental data of nuclear receptors were obtained from the ToxCastv3.2 database, and duplicate data and compounds with missing SMILES were removed from the experimental data to obtain compound information dataset A1, which includes: in vitro experimental data containing the following four nuclear receptor pathways selected according to research needs: estrogen receptor α (ER), estrogen receptor β (ERβ), androgen receptor (AR), and thyroid hormone receptor (TR).

[0028] In vitro high-throughput screening data for each nuclear receptor pathway were extracted and sorted according to the compound's AC50 (half-maximal concentration) value; a smaller AC50 value indicates higher potency. Based on this sorting, data deduplication was performed as follows: for multiple measurements of the same compound in the same nuclear receptor experiment, only the set with the lowest AC50 value was retained; for different compounds with identical AC50 values ​​in the same nuclear receptor experiment, only the data of one of the compounds was retained. From the deduplicated data, compounds with clearly defined SMILES (Simplified Molecular Input Line Entry System) structural representations were further screened. SMILES is a standard chemical representation method using ASCII strings to represent molecular structures and is widely used in quantitative structure-activity relationship studies. Compound data lacking SMILES information or containing erroneous SMILES representations were discarded.

[0029] 3. Beneficial effects

[0030] Compared to existing technologies, the advantages of this application are:

[0031] Molecular fingerprinting is used to extract multi-level structural features (primary, secondary, and tertiary) of compounds, and a classification model is constructed based on these features to categorize compounds for endocrine disruption effects. This approach can characterize the relationship between compound structure and endocrine disruption activity from different perspectives. Furthermore, a quantitative prediction model is built for each compound cluster. Depending on the number of compounds, either machine learning or quantitative read-across methods are selected, balancing the model's universality and specificity, effectively improving the prediction efficiency and accuracy of quantitative activities for different types of compounds.

[0032] In the process of constructing the quantitative prediction model, multiple machine learning algorithms are employed in conjunction with cross-validation and parameter optimization strategies to fully explore the feature patterns in the data and select the consensus prediction model with the best performance, thereby improving the robustness and reliability of the prediction results. Simultaneously, for small sample data, a quantitative read-across prediction method based on molecular structure similarity is proposed, which can effectively draw on information from known active compounds and expand the applicability of the model.

[0033] By selecting representative highly active and less active compounds based on molecular docking and kinetic simulation analysis, and revealing the differences in their binding modes with nuclear receptors and the mechanisms of induced conformational changes, the mechanism of endocrine disruption effects of different compounds can be explained at the molecular level, providing a reasonable explanation for the prediction results, and also providing important structural clues for subsequent molecular design and optimization.

[0034] By obtaining in vitro experimental data of four representative nuclear receptors (ERα, ERβ, AR, and TR) from the ToxCast database and employing a reasonable preprocessing strategy (sorting and deduplication based on AC50 values ​​and screening using SMILES representation), a high-quality compound information dataset A1 was obtained. This dataset comprehensively covers compounds with different types of endocrine-disrupting effects and ensures the accuracy and completeness of the data.

[0035] In the data preprocessing stage, a quality control strategy based on the Z-score method and multi-level screening criteria was proposed, which can effectively identify and eliminate false positive data caused by factors such as cytotoxicity, thereby improving data quality. Differential processing was applied to experimental data at different levels by combining indicators such as confidence level, problem labeling, and decision ratio. Attached Figure Description

[0036] Figure 1 This is a schematic diagram illustrating the specific implementation process of this application;

[0037] Figure 2 To demonstrate the predictive performance of a data test set in this application;

[0038] Figure 3 To demonstrate the predictive performance of another data test set in this application;

[0039] Figure 4 This provides the predictive performance of another data test set in this application.

[0040] Figure 5 To demonstrate the predictive performance of another data test set in this application. Detailed Implementation

[0041] The present application will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0042] like Figure 1 As shown, this application discloses a method for quantitative activity prediction of endocrine disruptors based on characteristic structures (taking EAT-type endocrine disruptors as an example), comprising the following steps: Using the R language tcpl package, extracting reporter gene experimental results for androgen receptor (AR), estrogen receptor α (ERα), estrogen receptor β (ERβ), and thyroid hormone receptor (TR) from the ToxCastv3.2 database to obtain an initial dataset of endocrine disruptor compounds. Preprocessing the initial dataset to remove duplicate data and compounds with missing SMILES structures. Based on hit-call values, initially classifying compounds into two categories: active (hit-call = 1) and inactive (hit-call = 0).

[0043] Based on in vitro experimental information from the ToxCastv3.2 database, active compounds are further subdivided into three activity types: agonists, antagonists, and agonist-antagonist, constructing a refined compound classification system. Following the five data quality scoring rules established in this application, the classified compounds are systematically screened: For active compounds (hit-call = 1): data with Z < 3 are removed using the Z-score method to eliminate false positive results caused by cytotoxicity; in level 5, data with model_prob not equal to 1 are removed; in level 6, data with flag count greater than 3 are removed. For inactive compounds (hit-call = 0): in level 5, data with model_prob not equal to 1 are removed; in level 7, data with hit_pct not equal to 0 are removed.

[0044] Analysis of the filtered and removed outlier data mainly covered the following situations: curves based on too low an activity concentration; AC50 values ​​below the experimental minimum concentration; edge responses in the curve; inefficient activity or high noise. Excluding these low-quality data improves the reliability of the dataset. Guided resampling and random noise addition methods were used to assess the reproducibility of AC50 values, providing data confidence information for subsequent modeling.

[0045] This application establishes a hierarchical characteristic structure classification system for endocrine-disrupting active compounds, used to systematically describe and characterize their structural features: Primary characteristic structure: Represents the basic structural skeleton of the compound, such as aromatic rings and aliphatic chains. Secondary characteristic structure: Based on the primary characteristics, structural modifications such as substituents and functional groups are introduced, such as hydroxyl groups, amino groups, and halogens. Tertiary characteristic structure: Further refined into characteristic structural patterns for three activity types: agonists, antagonists, and agonist-antagonists. Specifically, the basic skeleton such as aromatic rings and aliphatic chains in the primary structure provides the basis for hydrophobic interactions and van der Waals forces, which are the main driving forces for the binding of the compound to the receptor. Substituents and functional groups in the secondary structure can form hydrogen bonds, electrostatic interactions, etc., regulating the binding strength and specificity of the compound to the receptor. The tertiary structure reflects the key interaction patterns required for different activity types; for example, agonists need to stabilize the active conformation of the receptor, while antagonists block the activation process of the receptor. By extracting hierarchical characteristic structures, key factors in the interaction between the compound and the receptor can be captured, providing a molecular basis for activity prediction.

[0046] Feature structure extraction and encoding: PaDEL-Descriptor software was used to calculate the PubChem molecular fingerprint of the compounds. The frequency of substructures in active compounds was calculated, and all substructures were ranked. The top 120 most frequent substructures were then selected for substructure proportion analysis, calculating their proportion in the active and inactive compound datasets. Further substructure proportion analysis was used to remove "useless" structures that were high in frequency but only present in a few active compounds. Finally, Pearson correlation coefficients based on Euclidean distance were used to cluster the 120 high-frequency substructures, selecting the fragments with the largest difference in proportion between active and inactive compounds as primary feature structures. Based on the primary features, SARpy software was used to extract secondary feature structures. The chemical structure of the compounds was decomposed, the frequency of each fragment was calculated, and the fragment with the highest frequency was selected as the secondary feature structure. Tertiary feature structures were extracted according to the compound's activity type (agonist, antagonist, agonist-antagonist). All extracted feature structure fragments were SMARTS encoded to generate standardized structural representations.

[0047] Statistical analysis of characteristic structures is conducted, including the frequency of occurrence of characteristic structures at various levels in active compounds and the assessment of their correlation with endocrine-disrupting activity. Differences in characteristic structures among compounds of different activity types are analyzed to uncover specific structural patterns of agonists, antagonists, and agonist-antagonists. A correlation matrix between characteristic structures and activity intensity is constructed to provide characteristic parameters for establishing quantitative structure-activity relationship models.

[0048] Chemical spatial visualization, based on similarity measures of characteristic structures, performs chemical spatial mapping and cluster analysis on active compounds. Dimensionality reduction algorithms such as t-SNE are used to achieve two-dimensional visualization of active compounds, intuitively displaying their structural diversity and distribution patterns. Combined with activity data, the analysis of active clusters and patterns in chemical space guides compound optimization and the identification of novel endocrine disruptors.

[0049] A regression prediction model was constructed based on a three-level feature structure classification. The lgAC50 value was selected as a quantitative indicator to measure the endocrine-disrupting activity of compounds and served as the output variable of the regression prediction model. PubChem molecular fingerprints were used as feature representations of compounds and served as input variables for the regression prediction model. For compounds with multiple "model_ga" values, lower quartile normalization was performed to ensure that the data were compared on the same scale. The IQR (interquartile range) method was used to identify and remove outliers to reduce the impact of extreme values ​​on the model.

[0050] Different modeling strategies are adopted based on the number of compounds in a tertiary structural feature cluster to fully utilize available data and select appropriate modeling methods. The quantitative read-across method (≤50 compounds) is used to predict the endocrine-disrupting activity of compounds when the number of compounds in a tertiary structural feature cluster is small (≤50). Quantitative read-across is based on the principle of compound structural similarity, assuming that structurally similar compounds have similar biological activities. The specific steps are as follows: For the target compound, search for structurally similar compounds within the same tertiary structural feature cluster. Calculate the structural similarity between the target compound and other compounds within the cluster; commonly used similarity metrics include the Tanimoto coefficient and the Dice coefficient. Select several compounds with the highest structural similarity as reference compounds. Based on the known activity data of the reference compounds, a weighted average method is used to predict the activity of the target compound. The weights are usually proportional to the structural similarity, i.e., compounds with higher similarity have a greater weight in the prediction. Repeat the above steps for all compounds within the cluster to obtain the predicted activity value for each compound. Quantitative read-across methods fully leverage the correlation between structural similarity and biological activity, providing a simple and effective prediction method when data is limited. However, this method is limited by its dependence on the availability of reference compounds and the selection of similarity thresholds.

[0051] Machine learning regression algorithms (compound count > 50): When a tertiary feature cluster contains a large number of compounds (> 50), machine learning regression algorithms are used to predict the endocrine-disrupting activities of the compounds. Machine learning methods build predictive models by learning the intrinsic relationship between compound structure and activity. The specific steps are as follows: The compound dataset is randomly divided into training and test sets, typically in an 80% / 20% ratio. For compounds in the training set, a machine learning regression model is constructed using PubChem molecular fingerprints as features and the lgAC50 value as the target variable. Commonly used algorithms include linear regression, decision trees, support vector machines, and random forests. Hyperparameter tuning of the model is performed using methods such as cross-validation, and the best-performing model is selected. The optimized model is applied to the test set to evaluate its predictive performance on unknown data. Commonly used evaluation metrics include root mean square error (RMSE) and coefficient of determination (R^2). For models that meet the performance requirements, they are used to predict the activity values ​​of all compounds within the cluster. Machine learning regression algorithms can automatically learn structure-activity relationships from large amounts of compound data, exhibiting strong predictive and generalization abilities. However, machine learning models have relatively weak interpretability and are quite sensitive to data quality and the choice of feature representation.

[0052] Feature selection based on correlation: To screen features highly correlated with endocrine-disrupting activity, the correlation coefficient between each PubChem molecular fingerprint and "model_ga" is calculated. The specific steps are as follows: For each PubChem molecular fingerprint, its Pearson correlation coefficient with "model_ga" is calculated. The Pearson correlation coefficient measures the linear correlation between two variables, ranging from -1 to 1. The formula for calculating fingerprint similarity is as follows:

[0053]

[0054] in:

[0055] X is the variable value of the fingerprint. This is the average value of the PubChem fingerprint; model_ga is the model_ga value of the compound. This represents the average value of the compound's model_ga. The correlation coefficient ranges from -1 to 1, explained as follows: if r = 1, it indicates a perfect positive correlation between the corresponding PubChem molecular fingerprint and the model_ga value; if r = -1, it indicates a perfect negative correlation; if r = 0, it indicates no linear correlation between the corresponding PubChem molecular fingerprint and the model_ga value. Molecular fingerprints are sorted based on the absolute value of the correlation coefficient. The larger the absolute value, the stronger the correlation between the fingerprint and activity. The top k most correlated molecular fingerprints are selected as a feature subset for subsequent modeling tasks. The value of k can be adjusted based on specific problems and experience, typically selecting 10%–20% of the most correlated features. The selected highly correlated fingerprints are used as input features to construct the predictive model. Correlation analysis effectively identifies molecular fingerprints that significantly influence endocrine-disrupting activity, reducing feature dimensionality and improving the model's explanatory power and generalization ability.

[0056] Dimensionality reduction based on correlation matrices is used to reduce the dimensionality of molecular fingerprint data, eliminating redundant features and improving modeling efficiency and performance. The specific steps are as follows: First, the Pearson correlation coefficients between all molecular fingerprints are calculated to construct a correlation matrix. The correlation matrix is ​​a symmetric matrix where each element represents the correlation strength between two fingerprints. Then, cluster analysis, such as hierarchical clustering or k-means clustering, is performed on the correlation matrix to group highly correlated fingerprints into one cluster. For each cluster, a representative fingerprint is selected as the feature of that cluster, and the remaining fingerprints are considered redundant and removed. The representative fingerprint can be the fingerprint with the strongest correlation within the cluster, or the fingerprint with the highest correlation to "model_ga". The selected representative fingerprints are used as the dimensionality-reduced feature set for subsequent modeling tasks. Cluster analysis based on the correlation matrix effectively eliminates redundant information between molecular fingerprints, reduces feature dimensionality, and improves the computational efficiency and generalization performance of the model. Simultaneously, the most representative fingerprint in each cluster is retained, ensuring the integrity of the feature information.

[0057] Quantitative read-across prediction calculates the structural similarity between compounds. For tertiary feature structure clusters with a small number of compounds (≤50), the quantitative read-across method is used to predict the activity of the compounds. First, the structural similarity between the target compound and other compounds in the cluster is calculated. A commonly used similarity measure is the Tanimoto coefficient. The Tanimoto coefficient is based on the degree of overlap of molecular fingerprints and ranges from [0, 1]. The calculation formula is as follows: Tanimoto(A, B) = (A∩B) / (A∪B), where A and B represent the molecular fingerprints of the two compounds, ∩ represents the number of shared fingerprints, and ∪ represents the total number of fingerprints. The larger the Tanimoto coefficient, the higher the structural similarity between the two compounds.

[0058] The weighted average activity value is predicted by using a weighted average method based on the calculated structural similarity to predict the "model_ga" value of the target compound. The specific steps are as follows: For the target compound, select the k compounds with the highest Tanimoto coefficients within the cluster as reference compounds. The value of k can be adjusted according to the specific problem and experience, typically between 3 and 5. For each reference compound, calculate its contribution to the activity of the target compound based on its known "model_ga" value and Tanimoto coefficient. The contribution value is calculated using the formula: Contribution(i) = Tanimoto(target, i) × model_ga(i), where i represents the i-th reference compound, Tanimoto(target, i) represents the Tanimoto coefficient between the target compound and the i-th reference compound, and model_ga(i) represents the known activity value of the i-th reference compound. The contribution values ​​of all reference compounds are summed to obtain the predicted activity value of the target compound. The prediction formula is: Where model_ga(pre) is the predicted model_ga value of the compound, model_ga is the lgAC50 value of the compound, and S is the structural similarity value of the compound.

[0059] In machine learning modeling, the `train_test_split` function from the scikit-learn library is used to randomly split the dataset into training and test sets. The training set is used for model training and parameter optimization, while the test set is used to evaluate the model's generalization performance. The specific steps are as follows: Represent the feature data and target variable as X and y, respectively. Call the `train_test_split` function, setting the test set proportion to 20% (i.e., `train_size = 0.8`), and setting the `random_state` parameter to a fixed value (e.g., 42) to ensure consistent data splitting results in each run. Save the split datasets as X_train, X_test, y_train, and y_test, respectively, for subsequent modeling and evaluation steps.

[0060] Standardizing feature data to achieve zero mean and unit variance improves model convergence speed and stability. Standardization can be implemented using the `Standard Scaler` class from the scikit-learn library. The specific steps are as follows: Create a `Standard Scaler` object. Call the `fit_transform` method to fit and transform the training set features, obtaining the standardized training features `X_train_scaled`. Call the `transform` method to transform the test set features, obtaining the standardized test features `X_test_scaled`. Note that test set standardization requires using the mean and variance parameters of the training set to maintain data distribution consistency. Use the standardized features in subsequent modeling and prediction steps.

[0061] To comprehensively evaluate the performance of different algorithms on a given dataset, several representative regression models were selected for comparison and optimization. Candidate models included: linear models (linear regression, ridge regression, Lasso regression, elastic network regression); tree models (decision tree, random forest); ensemble models (AdaBoost, gradient boosting); support vector machines (SVR); and K-nearest neighbors (KNeighborsRegressor).

[0062] The optimal parameter combination for each model is found in the parameter search space using grid search and cross-validation. Grid search and cross-validation can be implemented using the Grid Search CV class in the scikit-learn library. The specific steps are as follows: For each candidate model, create a Grid Search CV object, specifying the model, parameter search space, and the number of folds for cross-validation (e.g., 5 folds). Call the `fit` method to perform grid search and cross-validation on the training set, finding the parameter combination with the best average performance on the validation set. Save the optimal parameter combination for subsequent model evaluation and prediction steps.

[0063] The optimized model is used to make predictions on the test set, and multiple evaluation metrics are calculated to quantify the model's performance. Commonly used regression evaluation metrics include: Coefficient of Determination (R-squared): measures the model's ability to explain data variation, ranging from [0, 1], with values ​​closer to 1 indicating a better fit. Root Mean Square Error (RMSE): measures the average deviation between predicted and true values, ranging from [0, +∞), with values ​​closer to 0 indicating a smaller prediction error. Mean Absolute Error (MAE): measures the average absolute deviation between predicted and true values, ranging from [0, +∞), with values ​​closer to 0 indicating a smaller prediction error. Mean Relative Error (MRE): measures the average proportion of prediction error to the true value, ranging from [0, +∞), with values ​​closer to 0 indicating a smaller relative error.

[0064] To further improve the accuracy and robustness of predictions, multiple high-performing models (e.g., R² > 0.6) are integrated to construct a consensus model. The consensus model combines the predictions of multiple base models using a weighted average, with weights determined based on the model's performance on the validation set (e.g., R² value). The specific steps are as follows: Select the k best-performing base models (e.g., k = 3) and predict on the test set to obtain k sets of predicted values. For each base model, calculate its weight based on its R² value on the validation set; the weight is proportional to the R² value. Perform a weighted average of the k sets of predicted values ​​to obtain the final prediction result of the consensus model. Evaluate the performance of the consensus model on the test set and compare it with individual base models to verify the effectiveness of the integration.

[0065] Figures 2 to 5 This figure shows the quantitative prediction results of the model for the interference effects of EAT-type active compounds of different activity types. The R-values ​​for all clusters can be seen in the figure. 2 The values ​​are all above 0.75, and the R values ​​of most clusters are... 2 A value exceeding 0.8 indicates that the dimensionality-reduced molecular fingerprint performs well in predicting EAT-type endocrine interference activities. Specifically, in a cross-sectional comparison of different interference activity types of the same nuclear receptor, agonists show the best predictive performance, while agonist-antagonist predictive performance is relatively poor.

[0066] The advantage of consensus models lies in their ability to combine the predictive power of multiple models, reducing the bias and variance of a single model and improving the stability and reliability of predictions. In practical applications, consensus models have been proven to significantly improve prediction performance, especially when multiple models have comparable performance.

[0067] The foregoing illustrative description of the invention and its embodiments is not restrictive and can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. The accompanying drawings are only one embodiment of the invention, and the actual structure is not limited thereto. No reference numerals in the claims should limit the scope of the claims. Therefore, if a person skilled in the art, inspired by this description, designs a similar structure and embodiment without departing from the spirit of the invention, such design should fall within the scope of protection of this patent. Furthermore, the word "comprising" does not exclude other elements or steps, and the word "a" preceding an element does not exclude the inclusion of "a plurality" of that element. Multiple elements stated in the product claims can also be implemented by a single element through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.

Claims

1. A method for predicting the quantitative activity of endocrine disruptors, comprising: The compound data of the nuclear acceptor in vitro were obtained, and duplicate compound data and compound data that did not contain the simplified molecular linear input canonical SMILES representation were removed to obtain the compound information dataset A1. Preprocessing compound information dataset A1 yields quality-controlled compound information dataset A2. Based on the quality-controlled compound information dataset A2, the molecular fingerprint method was used to extract the primary, secondary, and tertiary feature structures of the compounds. The primary feature structures represent the compound structural fragments that cause endocrine disruption effects; the secondary feature structures represent the compound structural fragments used to distinguish whether a compound has an endocrine disruption effect; and the tertiary feature structures represent the compound structural fragments used to distinguish the types of endocrine disruption effects. Based on the extracted primary, secondary, and tertiary feature structures, a compound classification model is constructed to classify the compounds in the compound information dataset A2 into compound clusters with or without endocrine interference effects and different interference effect types. For each compound cluster, a quantitative prediction model based on machine learning or quantitative read-across is constructed to predict the quantitative activity value of the compound; wherein, when the number of compounds is greater than a preset threshold, the machine learning method is used, and when the number of compounds is less than or equal to the threshold, the quantitative read-across method is used. From the compound information dataset A2, select representative active compounds, including: From the compound information dataset A2, endocrine-disrupting compounds belonging to different chemical structure types were selected as endocrine-disrupting compounds to be studied. Crystal structures of nuclear receptors that interact with selected endocrine disrupting compounds were obtained from the RCSB database. The crystal structures contain the ligand-binding domain (LBD) of the nuclear receptor. Using molecular docking, the selected endocrine disrupting compounds were simulated with the LBDs of the corresponding nuclear receptor crystal structures obtained, and the binding free energy of each endocrine disrupting compound with the corresponding nuclear receptor LBD was calculated. Based on the calculated binding free energy, endocrine-disrupting compounds with binding free energy below the threshold for the corresponding nuclear receptor LBD were screened and identified as active compounds with high binding capacity to the corresponding nuclear receptor LBD, and as representative active compounds of the corresponding nuclear receptor. Representative active compounds were selected from compound information dataset A2, including: Based on the AC50 value of each compound in the quality control compound dataset A2, the compounds are divided into high-activity or low-activity compounds, resulting in high-activity and low-activity compound subsets. Crystal structure data of nuclear receptors that interact with highly active or less active compounds were obtained from the RCSB database; the crystal structure data included the ligand-binding domain (LBD) of the nuclear receptor. Using each highly active or inactive compound as a ligand and the corresponding nuclear receptor LBD crystal structure as the receptor, ligand-receptor complexes were constructed using molecular docking, and the binding free energy of each ligand-receptor complex was calculated. Within the subset of highly active compounds, compounds with binding free energy to the corresponding nuclear receptor LBD below the first threshold were screened and used as representative highly active compounds of the corresponding nuclear receptor. Within the subset of low-activity compounds, compounds with a binding free energy to the corresponding nuclear receptor LBD higher than the second threshold are selected as representative low-activity compounds of the corresponding nuclear receptor; wherein, the first threshold is less than the second threshold. For representative active compounds, molecular docking and kinetic simulation analyses were performed to obtain the binding modes and induced conformational change mechanisms between the active compounds and nuclear acceptors, including: Using the unbound conformation of the nuclear receptor apo as the receptor starting structure, the obtained high-activity representative compound and low-activity representative compound were assembled with the corresponding apo conformation to form ligand-receptor complexes. Equilibrium simulations were performed on the ligand-receptor complex to obtain molecular dynamics trajectory frames of the ligand-receptor complex; Based on molecular dynamics trajectory frames, the differences in binding modes of highly active and inactive representative compounds with their corresponding endogenous ligands in LBD were compared. Endogenous ligands refer to small molecule compounds that naturally bind to each nuclear receptor. High and low activity in statistical molecular dynamics trajectory frames represent hydrogen bonds and van der Waals interactions between compounds and nuclear acceptors (LBDs), and the contributions of hydrogen bonds and van der Waals interactions to ligand-acceptor binding are calculated. The RMSD value of the twelve-helix H12 relative to the apo conformation in the molecular dynamics trajectory frame was calculated, and the conformational rearrangement of H12 under ligand binding was obtained to reveal the dynamic mechanism of conformational changes of nuclear acceptors induced by active compounds.

2. The method for predicting the quantitative activity of endocrine disruptors according to claim 1, characterized in that: After predicting the quantitative activity value of a compound using a quantitative prediction model based on machine learning or quantitative read-across, the following steps are also included: Select representative active compounds from compound information dataset A2; For representative active compounds, molecular docking and kinetic simulation analysis were performed to obtain the binding mode of active compounds with nuclear acceptors and the mechanism of induced conformational change.

3. The method for predicting the quantitative activity of endocrine disruptors according to claim 2, characterized in that: When the number of compounds in a compound cluster exceeds a preset threshold, a quantitative prediction model for the corresponding compound cluster is constructed using machine learning methods, including: The molecular fingerprint characteristics of compounds in the corresponding compound clusters are standardized. The standardized dataset of compound clusters is randomly divided into training and testing sets; Multiple machine learning regression algorithms are used to model the training set, and Grid Search CV is used to perform cross-validation and parameter optimization on the model to obtain the optimal model for different machine learning regression algorithms; among them, multiple machine learning regression algorithms include at least one of linear regression, ridge regression, Lasso regression, elastic network, decision tree, random forest, AdaBoost, gradient boosting, support vector machine or K nearest neighbor; The optimal models using various machine learning regression algorithms were used to make predictions on the test set, and the determination coefficients of each optimal model were calculated. ; Choose the coefficient of determination The optimal model with a coefficient of determination greater than the threshold will select the coefficient of determination. The prediction results of the optimal model with values ​​greater than the threshold on the test set are weighted and averaged to obtain the machine learning consensus prediction model for the corresponding compound cluster, which serves as the quantitative prediction model.

4. The method for predicting the quantitative activity of endocrine disruptors according to claim 2, characterized in that: When the number of compounds in a compound cluster is less than or equal to a threshold, a quantitative prediction model for the corresponding compound cluster is constructed using a quantitative read-across method, including: Within the corresponding compound cluster, based on the molecular structural similarity between each test compound and each known active compound, the structural similarity value S between the test compound and each known active compound is calculated; The predicted activity of the test compound is obtained by weighting the lgAC50 activity values ​​of known active compounds based on the structural similarity value S. ; The predicted activity values ​​of the test compounds obtained The predicted activity values ​​of the corresponding compounds were used as quantitative read-across values; and the predicted activity values ​​of the test compounds were used as... The Tanimoto coefficient is used as the predicted value for the corresponding compound to indicate the strength of the compound's endocrine-disrupting activity. The predicted endocrine-disrupting activities of all compounds within a compound cluster are calculated to obtain the quantitative read-across prediction results for the corresponding compound cluster.

5. The method for predicting the quantitative activity of endocrine disruptors according to claim 4, characterized in that: Predicted activity values ​​of the test compound : in, For the prediction of compounds value, denoted as lgAC50, where S is the structural similarity value of the compound; i represents the target compound, and k represents the reference compound.

6. The method for predicting the quantitative activity of endocrine disruptors according to any one of claims 2 to 5, characterized in that: The acquired experimental data were preprocessed, including: Data with Z-scores less than 3 were screened and removed to exclude false positives caused by cytotoxicity; where the Z-score represents the statistical significance of the relative activity of the compound in the corresponding nuclear receptor experiment compared with the negative control group. Hit-call value was selected as the evaluation criterion, and quality assessment was carried out at three levels, from level 5 to level 7. Level 5 represents the primary screening experimental data, level 6 represents the secondary validation experimental data, and level 7 represents the dose-response relationship confirmation experimental data. for Compound data: In level 5, delete Data that is not equal to 1, among which, Indicates the confidence level of a compound's activity; In level 6, delete data with more than 3 flags, where flags represent data quality issues. For compounds with hit-call=0: In level 5, delete Data that is not equal to 1; In level 7, delete Non-zero data, among which, This indicates the percentage of experimental data that were determined to be active out of the total experimental data.

7. The method for predicting the quantitative activity of endocrine disruptors according to claim 6, characterized in that: Nuclear receptor in vitro experimental data were obtained from the ToxCastv3.2 database, and duplicate data and compounds with missing SMILES were removed to obtain compound information dataset A1, which includes: Based on the nuclear receptor pathway titers demonstrated by the compounds in the high-throughput screening experiments of the ToxCast group, in vitro experimental data of nuclear receptors including estrogen receptor ERα, estrogen receptor ERβ, androgen receptor AR, and thyroid hormone receptor TR were selected. The acquired in vitro experimental data were sorted in ascending order of AC50 value, and duplicate data were removed. After removing duplicate data from the in vitro experimental data, compounds that do not contain SMILES structural information are removed, resulting in compound information dataset A1.