Chemical reproductive toxicity prediction method fusing aop and multi-scale molecular features

CN122599097APending Publication Date: 2026-08-18SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610850934.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

这种缺乏生物学证据链的预测方式,导致模型在处理结构新颖化合物或具有多重毒性路径的物质时,泛化能力差、假阴性率高,且由于缺乏机理解释性,无法为化学品的结构优化与安全设计提供指导建议

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122599097A_ABST
    Figure CN122599097A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of chemical reproductive toxicity prediction method of fusing AOP with multiscale molecular characteristics, belong to the field of computational toxicology, bioinformatics and environmental risk technology, by for each compound molecule, using Morgan algorithm generates extended connectivity fingerprint, and based on extended connectivity fingerprint, construct three-dimensional docking panorama fingerprint and polymorphic residue contact matrix;Cascade machine learning model of fusing AOP key event and physical contact matrix is constructed, and the cascade machine learning model of fusing AOP key event and physical contact matrix is output reproductive toxicity prediction probability and ruling result.The present application innovatively introduces polymorphic non-covalent interaction matrix based on full receptor protein sequence and macroscopic docking binding energy.Through the precise matrix coding of the physical contact state such as hydrogen bond, hydrophobic interaction, packing and salt bridge between thousands of ligands and receptor amino acid residues, the present model truly gives artificial intelligence with three-dimensional space physical vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of technology, and in particular to a method for predicting the reproductive toxicity of chemicals that integrates AOP with multi-scale molecular characteristics. Background Technology

[0002] Rapid and accurate assessment of the reproductive toxicity of chemicals is a key bottleneck in the current regulation of environmental health and industrial chemicals. Existing assessment methods mainly face the following technical challenges: First, traditional reproductive toxicity assessments rely on in vivo experiments using rodents as models. These experiments are extremely time-consuming (usually lasting several months), costly, and involve animal ethics issues, making them unsuitable for the market regulatory needs of rapid iteration and high-throughput screening of current industrial chemicals.

[0003] Secondly, existing computational toxicology models (such as traditional QSAR models) generally suffer from a "black box" effect. These models are often based on linear or nonlinear relationships between chemical structure descriptors and toxic endpoints, frequently neglecting the essential toxicological mechanisms by which chemicals induce damage in the body (such as activation of specific receptors, mitochondrial oxidative stress, etc.). This predictive approach, lacking a chain of biological evidence, results in poor generalization ability and high false-negative rates when dealing with structurally novel compounds or substances with multiple toxic pathways. Furthermore, due to the lack of mechanistic explanation, they cannot provide guidance for the structural optimization and safety design of chemicals.

[0004] Finally, current computer-aided risk assessment methods lack quantitative analysis of target molecular interaction mechanisms. Due to the lack of precise docking and kinetic assessment of key amino acid residues in the active pockets of ligands and receptor proteins (such as key hotspot sites in the 3ERT protein), existing methods cannot effectively distinguish between structurally similar but significantly toxicologically different homologues, leading to high uncertainty in environmental risk warnings. Therefore, developing an intelligent prediction system that integrates a multi-level toxicological pathway (AOP) framework and fuses molecular structural features with physical docking interaction energies has become a common technical challenge urgently needing to be addressed in the fields of environmental toxicology and computational prediction. Summary of the Invention

[0005] This invention overcomes the shortcomings of the prior art and provides a method for predicting the reproductive toxicity of chemicals by integrating AOP with multi-scale molecular characteristics.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The first aspect of this invention provides a method for predicting the reproductive toxicity of chemicals by integrating AOP (Aspect-Oriented Programming) with multi-scale molecular features, comprising the following steps: Acquire multi-source toxicological data, and perform structural standardization, cleaning and deduplication, and sample equalization using a downsampling algorithm on the multi-source toxicological data; For each compound molecule, the Morgan algorithm is used to generate an extended connectivity fingerprint, and a three-dimensional docking panoramic fingerprint and a contact matrix of polymorphic residues are constructed based on the extended connectivity fingerprint. A cascaded machine learning model integrating AOP key events and physical contact matrix is ​​constructed, and the reproductive toxicity prediction probability and adjudication result are output using the cascaded machine learning model integrating AOP key events and physical contact matrix. The SHAP algorithm is used to calculate the expected marginal contribution of the feature subset, and the top-ranking ECFP fingerprint positions are reverse-engineered into specific chemical fragments, and the chemical structure is traced. Input a list of untested chemicals, extract two-dimensional fingerprints, deduce mechanism probabilities, and concurrently execute a large number of molecular docking and interaction matrices in the computing environment. Based on fusion characteristics, give a hazard rating and output a list of chemicals with a high probability of toxicity.

[0007] Furthermore, in the chemical reproductive toxicity prediction method that integrates AOP and multi-scale molecular features, multi-source toxicological data information is acquired, and the multi-source toxicological data information is subjected to structure standardization, deduplication, and sample equalization using a downsampling algorithm. Specifically: High-throughput in vitro screening data related to the core critical events of reproductive toxicity AOP were obtained from the toxicology database, and simplified molecular linear input canonical strings for all compounds were obtained based on the high-throughput in vitro screening data. The simplified linear input strings of all the compounds were processed using an open-source cheminformatics toolkit to remove inorganic salt ions, neutralize molecular charges, remove heavy metals and organometallic complexes, and unify the expression forms of tautomers. For contradictory test results for the same compound, a "majority voting method" or a safety conservatism principle prioritizing positive results is used to remove duplicates. A downsampling algorithm is implemented on the in vitro MIE and KE datasets. Under the condition that the number of negative samples is much larger than the number of positive samples, a random sampling algorithm without replacement is used to extract a subset from the negative sample set with an equal number of positive samples, and then merge them to construct a balanced training set with a ratio of 1:1.

[0008] Furthermore, in the chemical reproductive toxicity prediction method that integrates AOP and multi-scale molecular features, for each compound molecule, the Morgan algorithm is used to generate an extended connectivity fingerprint, and a three-dimensional docking panoramic fingerprint and polymorphic residue contact matrix are constructed based on the extended connectivity fingerprint, specifically: The Morgan algorithm is introduced. For each compound molecule, the Morgan algorithm sets the search radius to two chemical bonds and iteratively searches for the structural information of the polymer atoms and their neighborhood environment, with each heavy atom in the molecule as the center. The aggregated information is folded and mapped to a fixed-length binary bit vector of 1024 bits by a hash function. Each bit of the binary bit vector corresponds to the presence or absence state of a specific chemical microenvironment or substructure fragment, forming a two-dimensional topological feature representation. The crystal structure of the human estrogen receptor ligand binding domain was obtained from a protein database. The water of crystallization was removed and polar hydrogen atoms were added using the AutoDockTools program. The Gasteiger bias charge was calculated and converted into the docking-specific PDBQT format. The collected toxicity validation set and candidate ligand molecules SMILES were used to generate 3D conformations using the ETKDG algorithm, and after force field optimization, they were converted into PDBQT format, and the docking grid points were set to cover the binding pocket of the human estrogen receptor protein. Using the AutoDockVina docking engine, the minimum binding free energy of all ligands and receptors is calculated. The detection results of ligands are integrated. When a ligand and a receptor make effective contact with a certain amino acid residue, the corresponding position in the matrix is ​​recorded as 1, otherwise it is recorded as 0. Finally, the interaction states of all compounds with full-length protein residues are transformed and flattened into a binary interaction matrix.

[0009] Furthermore, in the chemical reproductive toxicity prediction method that integrates AOP (Aspect-Oriented Events) and multi-scale molecular features, a cascaded machine learning model integrating AOP key events and physical contact matrices is constructed. This cascaded machine learning model outputs the reproductive toxicity prediction probability and decision result, specifically as follows: A two-stage cascaded architecture is adopted to deeply stitch together the structure, mechanism probability and high-dimensional physical contact matrix. In the first stage, the AOP mechanism response classifier is constructed with 1024 bits as input to train random forest classification models for predicting human estrogen receptor activation and mitochondrial damage respectively. The random forest classification model outputs a probability vector of compound-induced human estrogen receptor activation and mitochondrial damage through the predict_proba function. In the second stage, a final reproductive toxicity prediction ensemble model is constructed. For each ligand molecule in the dataset, perform lateral stitching of the feature space to generate a panoramic fusion feature vector, and input the panoramic fusion feature vector into the second-level random forest classifier; Parameters were set to address the class imbalance of in vivo reproductive toxicity data, and the model was fine-tuned through 5-fold cross-validation. The final output of the reproductive toxicity prediction probability and the adjudication result is the model.

[0010] Furthermore, in the chemical reproductive toxicity prediction method that integrates AOP and multi-scale molecular features, the SHAP algorithm is used to calculate the expected marginal contribution of feature subsets. The top-ranking ECFP fingerprint sites are then reverse-engineered into specific chemical fragments, and chemical structure tracing is performed. Specifically: The SHAP algorithm is introduced to calculate the expected marginal contribution of the feature subset, quantify the contribution of each feature to the prediction of reproductive toxicity, and rank the contributions. The two-dimensional topological features of the pre-defined ranking contribution are reverse-engineered to specific chemical fragments. The toxic functional groups are highlighted in color in the molecular skeleton by the RDKit visualization engine, and the source is traced by extracting residues with significant SHAP values.

[0011] Furthermore, in the chemical reproductive toxicity prediction method that integrates AOP and multi-scale molecular features, a list of untested chemicals is input, two-dimensional fingerprints are extracted, mechanism probabilities are deduced, and a large number of molecular docking and interaction matrices are executed concurrently within the computational environment. Based on the fusion features, a hazard rating is given, and a list of chemicals with a high probability of toxicity is output, specifically: Input a massive list of untested chemicals from the existing chemical substance catalog, extract two-dimensional fingerprints, deduce mechanism probabilities, and concurrently execute a large number of molecular docking and interaction matrix transformations within the computing environment; The model provides a hazard rating based on fusion characteristics, outputs a list of chemicals with a high probability of toxicity, and classifies pathogenic mechanisms into single receptor-mediated, mitochondrial damage-mediated, and dual-target synergistic destructive types based on Boolean logic.

[0012] This invention addresses the shortcomings of the prior art and has the following beneficial effects: (1) Breaking through the limitations of two dimensions, the robustness and accuracy of the prediction model are significantly improved. Traditional computational toxicology (QSAR) models rely heavily on two-dimensional molecular fingerprints, making it difficult to distinguish homologues exhibiting the "activity cliff" phenomenon (i.e., molecules with highly similar structures but vastly different toxicities). This invention innovatively introduces a "polymorphic non-covalent interaction matrix (PLIF)" based on the entire receptor protein sequence and macroscopic docking binding energies. By precisely encoding the physical contact states—such as hydrogen bonds, hydrophobic interactions, π-π stacking, and salt bridges—between thousands of ligand and receptor amino acid residues using a 0 / 1 matrix, this model truly endows artificial intelligence with "three-dimensional spatial physical vision." In independent test sets, the ensemble model incorporating multi-scale features consistently maintained a ROCAUC index above 0.80. This result demonstrates that by introducing underlying physicochemical interaction principles, the false positive and false negative rates introduced by purely statistical models are significantly reduced, making toxicity predictions more accurate and robust.

[0013] (2) Breaking the "black box effect" and realizing a two-way "white box" panoramic mechanism analysis of prediction results Existing deep learning and ensemble learning models are often considered "black boxes" due to their extremely complex internal structures, making them difficult for toxicology experts and regulatory agencies to accept. This invention applies the SHAP cooperative game theory algorithm to successfully achieve bidirectional source tracing analysis of feature contributions: on the one hand, "downward source tracing": accurately extracting the Top 10 chemical structural warning words (StructuralAlerts) that contribute the most to reproductive toxicity, and using RDKit to achieve highlighted visualization of toxic molecular fragments; on the other hand, "upward target anchoring": through the analysis of SHAP values ​​in the high-dimensional interaction matrix, the system automatically identifies the "key hotspot amino acid residues" in the receptor protein that truly drive toxicity. This bidirectional verification chain of "specific toxic chemical fragments + specific amino acid residue interactions" provides a textbook-level molecular toxicology explanation for AI prediction results, filling the technical gap in existing technologies that "know what but not why." Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained from these drawings without creative effort.

[0015] Figure 1 The SHAP force graph used for toxicity feature importance analysis in machine learning is shown. Figure 2 The first ROC curve for machine learning is shown; Figure 3 A bar chart showing toxicology-related statistics is provided. Figure 4 The second ROC curve for machine learning is shown; Figure 5 A flowchart of the overall methodology for predicting the reproductive toxicity of chemicals by integrating AOP with multi-scale molecular features is shown. Detailed Implementation

[0016] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0017] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0018] Example 1 This invention aims to establish an intelligent prediction and risk assessment system for the reproductive toxicity of chemicals based on the adverse outcome pathway (AOP) mechanism by deeply integrating computational toxicology, structural biology, and multi-scale machine learning technologies. This system will systematically address the core technical challenges currently faced in the field of chemical safety, such as long assessment cycles, unclear mechanisms, and limited prediction accuracy. Specific objectives of the invention cover the following aspects: First, this invention aims to overcome the over-reliance on traditional animal experiments in existing chemical reproductive toxicity assessments. By establishing a high-throughput computer-aided screening method, it is possible to rapidly and cost-effectively classify the toxicity of tens of thousands of industrial chemicals in my country's existing chemical substance inventory (IECSC), thereby effectively reducing the ethical burden on animals in environmental risk assessments and greatly shortening the decision-making cycle from chemical research and development to market entry or environmental emission assessment.

[0019] Secondly, this invention aims to overcome the limitations of traditional quantitative structure-activity relationship (QSAR) models as "black box" predictions. Existing prediction models mostly focus on the apparent statistical correlation between molecular structural fragments and toxicity, lacking a deep characterization of the essential toxicological mechanisms of chemical-induced damage. This invention, by introducing two key events (KEs)—ESR1 receptor activation and mitochondrial dysfunction—and combining them with SHAP attribution analysis, can clearly elucidate the specific biochemical pathways and toxic molecular fragments that chemicals interfere with the reproductive system. This "mechanism-driven" prediction model can provide a solid chain of biological evidence for environmental risk early warning, making the AI ​​prediction results interpretable by experts.

[0020] Furthermore, this invention aims to refine the molecular-level analysis of toxic mechanisms. By integrating molecular docking technology, the physical binding affinity between candidate toxicants and key active residues of the ESR1 receptor (such as RES_353, RES_394, etc.) is quantified, achieving cross-scale feature fusion from two-dimensional molecular fingerprints to three-dimensional spatial interactions. This aims to solve the technical challenge of traditional models failing to distinguish homologues with similar structures but significantly different toxicities, thereby greatly improving the sensitivity and specificity for identifying potent endocrine disruptors and reproductive toxins.

[0021] Finally, this invention aims to provide a practical and intelligent decision-making tool for national chemical environmental risk prevention and control. By conducting a comprehensive screening of my country's existing chemical substance inventory, it achieves accurate identification and graded early warning of potentially high-risk chemicals. This not only optimizes the risk prevention and control priorities of regulatory authorities but also guides the green substitution and structural optimization of chemicals, resulting in significant social and environmental benefits.

[0022] like Figure 5 As shown, the core of this patent application is to disclose a rapid prediction method for the reproductive toxicity of chemicals based on the fusion of adverse outcome pathway (AOP) mechanisms and multi-scale molecular features. This technical solution aims to address the shortcomings of traditional toxicity prediction models, such as their singular structure, lack of biological mechanism explanation, and incomplete analysis of receptor three-dimensional physical interactions. This invention constructs a cross-scale cascaded machine learning framework by coupling two-dimensional topological fingerprints, three-dimensional molecular docking binding energies, polymorphic protein-ligand full residue interaction matrices (including hydrogen bonds, hydrophobicity, π-π stacking, etc.), and multi-target pathway activation probabilities.

[0023] Specifically, the present invention provides a rapid prediction method for the reproductive toxicity of chemicals based on the fusion of AOP mechanism and multi-scale molecular features, comprising the following steps S1 to S5 performed sequentially: S1. Construction, standardization, cleaning, and sample equalization of multi-source toxicology datasets This step aims to provide high-quality, low-noise benchmark training data for subsequent machine learning models.

[0024] Data Collection and Classification: High-throughput screening (HTS) data related to key events (KEs) of reproductive toxicity adverse events (AOPs) were acquired from toxicology databases (such as PubChem and ToxCast). Specifically, this included large-scale datasets (covering thousands to tens of thousands of compounds) of estrogen receptor (ESR1) agonist / antagonist activity characterizing molecular initiation events (MIEs), and mitochondrial membrane potential (MMP) damage datasets characterizing subsequent key events (KEs). Simultaneously, animal experimental validation datasets with clearly defined in vivo reproductive toxicity endpoints (Adverse Outcomes, AOs) were collected.

[0025] Structural normalization and deduplication: Obtain simplified molecular linear input canonical (SMILES) strings for all compounds. Canonicalize the SMILES using open-source cheminformatics toolkits (such as RDKit). Specific cleaning rules include: removing inorganic salt ions from compounds, neutralizing molecular charges, removing heavy metals and organometallic complexes, and standardizing the expression forms of tautomers. For contradictory data with multiple test results for the same compound, deduplication is performed using a majority voting method or a safe conservatism principle prioritizing positive results.

[0026] Majority-Class Downsampling: Due to extreme class imbalance in in vitro toxicity screening data (the number of inactive negative samples is often more than 10 times that of positive samples), direct training can lead to severe majority class bias in the classifier. This invention implements a downsampling algorithm on the in vitro MIE and KE datasets: when the number of negative samples is much greater than the number of positive samples, a subset with an equal number of positive samples is extracted from the negative sample set using a random sampling algorithm without replacement, and these subsets are merged to construct a balanced training set with a 1:1 ratio.

[0027] S2. Characterization and Extraction of Multi-Dimensional, Multi-Scale Molecular Feature Space This step aims to map the chemical structure of the compound and the microscopic physical properties of its receptor interaction into a computer-recognizable high-dimensional mathematical vector. The feature space contains two-dimensional structural features and a high-resolution three-dimensional interaction fingerprint matrix. The optimal docking complex's three-dimensional coordinates are analyzed. An interaction detection algorithm is developed to systematically detect whether non-covalent physicochemical contacts with the following specific properties have occurred: Two-dimensional topological structure feature representation (ECFP fingerprint): Extended-Connectivity Fingerprints are generated using the Morgan algorithm. For each compound molecule, with each heavy atom in the molecule as the center, a search radius of 2 chemical bonds (equivalent to ECFP4) is set, and the structural information of the aggregated atoms and their neighborhood environment is iteratively aggregated. The aggregated information is folded and mapped into a fixed-length 1024-bit binary bit vector using a hash function. Each bit of this vector corresponds to the presence or absence state of a specific chemical microenvironment or substructure fragment.

[0028] Construction of a 3D docking panoramic fingerprint and polymorphic residue contact matrix (PLIF): To completely overcome the shortcomings of 2D fingerprints in terms of spatial conformation and target binding specificity, this invention innovatively introduces a multiple non-covalent interaction matrix based on the whole amino acid sequence. Specifically, it includes the following steps: Macromolecular receptor preparation: Obtain the crystal structure of the human estrogen receptor (ESR1) ligand-binding domain (e.g., PDBID:3ERT) from the protein database (PDB). Use programs such as AutoDockTools to remove water of crystallization, add polar hydrogen atoms, calculate Gasteiger bias charge, and convert to the docking-specific PDBQT format.

[0029] Batch ligand 3D conformation generation: The collected SMILES of more than 2,000 toxicity validation sets and candidate ligand molecules were used to generate 3D conformations using the ETKDG algorithm, and then converted into PDBQT format after force field optimization.

[0030] Spatial grid search and docking scoring: A docking grid (GridBox) was set to cover the binding pocket of the ESR1 protein (preferred coordinates: Center_X=22.63, Y=5.40, Z=22.05, size 20 ų). Using the AutoDockVina docking engine, the minimum binding free energy (BindingAffinity) for all 2000+ ligands and receptors was calculated.

[0031] Generate a giant 0 / 1 protein-ligand interaction fingerprint (PLIF) matrix for each ligand and every amino acid residue in the full sequence of the receptor protein. Hydrogen bonds: the distance between the donor and acceptor atoms is <3.5 Å and the donor-hydrogen-acceptor angle is >130°; Hydrophobic interactions: The distance between the aliphatic / aromatic carbon of the ligand and the hydrophobic side chain carbon of the protein is <4.5 Å; π-π stacking: The distance between the centroids of the aromatic rings is <6.0 Å and the angle between the normal vectors conforms to the T-type or parallel stacking model; Salt bridges and electrostatic interactions: The center-to-center distance between pairs of atoms with opposite charges is <4.0 Å.

[0032] The detection results of the above 2000+ ligands were integrated. If a ligand and a certain amino acid residue of the receptor made effective contact in any of the above-mentioned ways, the corresponding position in the matrix was recorded as 1; otherwise, it was recorded as 0. Finally, the interaction states of all compounds with the full-length protein residues were transformed and flattened into a giant 0 / 1 binary interaction matrix.

[0033] It should be noted that traditional computational toxicology (QSAR) models rely heavily on two-dimensional molecular fingerprints, making it difficult to distinguish homologues exhibiting the "activity cliff" phenomenon (i.e., molecules with highly similar structures but vastly different toxicities). This invention innovatively introduces a "polymorphic non-covalent interaction matrix (PLIF)" based on the entire receptor protein sequence and macroscopic docking binding energy. By precisely encoding the physical contact states between thousands of ligand and receptor amino acid residues using a 0 / 1 matrix, such as hydrogen bonds, hydrophobic interactions, π-π stacking, and salt bridges, this model truly endows artificial intelligence with "three-dimensional spatial physical vision." In independent test sets, the ensemble model incorporating multi-scale features consistently maintained a ROCAUC index above 0.80. This result demonstrates that by introducing underlying physicochemical interaction principles, the false positive and false negative rates introduced by purely statistical models are significantly reduced, making toxicity predictions more accurate and robust.

[0034] S3: A cascaded machine learning model integrating AOP key events and physical contact matrices A two-tier cascaded architecture is adopted to deeply stitch together the structural and mechanistic probabilities with the high-dimensional physical contact matrix.

[0035] Level 1: The AOP mechanism response classifier is constructed using 1024 bits as input to train random forest classification models that predict ESR1 activation (MIE) and mitochondrial damage (KE) respectively. The model outputs probability vectors of compound-induced MIE and KE through the predict_proba function, thereby realizing the quantitative expression of toxicological empirical knowledge.

[0036] Level 2: Constructing the final reproductive toxicity (AO) prediction ensemble model. For each ligand molecule in the dataset, a lateral feature concatenation is performed. The panoramic fusion feature vector contains four dimensions: structural fingerprint (1024 bits); pathway response probability (2 bits); macroscopic binding energy scalar (1 bit, normalized); and microscopic polymorphic interaction 0 / 1 matrix (hundreds of bits, representing the panoramic contact state between the ligand and whole protein residues at the hydrogen bonding, hydrophobicity, and π-π stacking levels). This high-dimensional input (up to several thousand dimensions) is fed into the second-level random forest classifier. Parameters are set to handle class imbalance in in vivo reproductive toxicity data; `class_weight='balanced'` and `max_depth=15` are set to prevent overfitting risks caused by the high-dimensional matrix. After 5-fold cross-validation optimization, the model finally outputs the reproductive toxicity prediction probability and the final decision.

[0037] It should be noted that, as Figure 1 , Figure 2 and Figure 4 As shown, existing deep learning and ensemble learning models are often considered "black boxes" due to their extremely complex internal structures, making them difficult for toxicology experts and regulatory agencies to accept. This invention applies the SHAP cooperative game theory algorithm to successfully achieve bidirectional source tracing analysis of feature contributions: on the one hand, "downward source tracing structure" accurately extracts the Top 10 chemical structure warning words (StructuralAlerts) that contribute the most to reproductive toxicity, and uses RDKit to achieve highlighted visualization of toxic molecular fragments; On the other hand, "anchoring the target upwards": by analyzing the SHAP value of the high-dimensional interaction matrix, the system automatically identifies the "key hotspot amino acid residues" in the receptor protein that truly drive toxicity. This two-way verification chain of "specific toxic chemical fragments + specific amino acid residue interactions" provides a textbook-level molecular toxicology explanation for the AI's prediction results, filling the technical gap in existing technologies where "we know what happens but not why."

[0038] S4. Bidirectional Attribution Analysis and Target Confirmation Based on Cooperative Game Theory Prediction Model: This step overcomes the "black box" limitation, achieving bidirectional interpretation of structural warning terms and core amino acid residues. Specifically, it includes: Feature Marginal Contribution Calculation (SHAP Algorithm): The SHAP algorithm is applied to analyze the trained fusion model. By calculating the expected marginal contribution of feature subsets, the impact of each feature on the prediction of reproductive toxicity (Shapley value) is quantified.

[0039] Bidirectional tracing of structural fragments and key residues: (1) Chemical structure tracing: For the ECFP fingerprint bits that contribute the most, reverse the process to identify specific chemical fragments and use the RDKit visualization engine to highlight toxic functional groups in the molecular skeleton (StructuralAlerts).

[0040] (2) Since the model is input with the 0 / 1 contact matrix of all amino acids, by extracting the source features of residues with significant SHAP values, the “key hot spot amino acids” that truly drive toxicity can be accurately identified. This proves that the model has not only learned to look at the shape of molecules, but also learned to make rational toxicity judgments based on whether “the ligand has formed hydrogen bonds or π-π stacking with a specific residue of the receptor”.

[0041] It should be noted that this invention not only has assessment and early warning functions, but also significant industrial guidance value. Through the "high-risk structure warning words" and "critical residue contact matrix" output by this invention, materials scientists and pharmaceutical chemists can directly obtain clear structural modification guidelines. For example, if the system identifies that a certain type of nitrogen-containing group readily undergoes strong π-π stacking with specific amino acids of the ESR1 receptor, thereby causing toxicity, researchers can specifically modify or replace this group in the early stages of molecular design for novel pesticides, plasticizers, or industrial solvents to eliminate physical interactions. Therefore, this invention provides a powerful reverse design basis for "Green and Sustainable Chemistry."

[0042] 5. High-throughput mechanism-driven early warning based on China's existing chemical substance inventory (IECSC) The "AOP-Structure-Panoramic Force" fusion system, which has been constructed and validated, will be deployed in actual environmental chemical regulatory operations. It will input a massive list of untested chemicals from the China International Chemical Substances Inventory (IECSC). The system will automatically extract two-dimensional fingerprints, deduce mechanism probabilities, and concurrently perform large-scale molecular docking and interaction matrix (0 / 1 matrix) transformations within the computational environment. The final model will provide a hazard rating based on the fusion characteristics and output a list of chemicals with a high probability of toxicity. Furthermore, it can classify their pathogenic mechanisms based on Boolean logic into: single receptor-mediated, mitochondrial damage-mediated, and dual-target synergistic destructive mechanisms, providing detailed and interpretable data support for my country's chemical risk prioritization and green chemical substitution design.

[0043] It should be noted that real-world chemical reproductive toxicity is often not caused by a single pathway. This invention, based on the AOP framework, outputs the independent probabilities of occurrence for "ESR1 receptor activation (MIE)" and "mitochondrial dysfunction (KE)" at the first level of the model. Through this cascade mechanism, this invention can not only determine whether a chemical is "toxic," but also finely classify it based on Boolean logic into "receptor-only mediated," "mitochondrial damage-only mediated," and "dual synergistic toxicity." This effect provides data support for revealing the complex network toxicological effects of environmental pollutants across multiple targets and pathways, greatly deepening the scientific understanding of the diversity of chemical toxicity mechanisms.

[0044] It should be noted that the existing chemical lists in my country and globally contain a vast number of chemical substances with "data-poor" information. Relying entirely on traditional live animal reproductive toxicity testing (such as the OECD guidelines) would not only consume hundreds of millions of dollars and involve extremely long testing cycles, but also face serious animal welfare and ethical controversies. The intelligent virtual screening system provided by this invention can concurrently perform feature extraction and mechanism deduction in a computer cluster, completing high-throughput risk screening of thousands or even tens of thousands of chemical substances in a very short time. The application of this invention puts into practice the toxicological alternative approach (3Rs principle) from the source, providing regulatory authorities with an efficient and low-cost early warning tool for developing priority control chemical lists.

[0045] In addition, this method may also include the following for predicting toxicity: Based on the three-dimensional interaction fingerprint matrix, molecular structure modes of three-dimensional interaction fingerprint matrix are constructed, and physicochemical property modes are constructed by combining the molecule's logP, molecular weight, number of hydrogen bond donors, number of hydrogen bond acceptors, and number of rotatable bonds. Activity data of assay endpoints related to key DART pathways were extracted from the ToxCast database to construct in vitro bioactivity modalities. Pharmacokinetic and toxicokinetic parameters were calculated using a pre-trained ADMET prediction model to form the ADMET modalities. Based on chemical names, CAS numbers and similar information, PubMed and patent databases are retrieved. The toxicity information described in the literature is extracted through the BioBERT variant of the BERT model, and document-level embedding representations are generated. These representations are then aggregated into chemical-level knowledge representations through attention weighting to form a literature knowledge modality. The embedding vectors of the five modalities are first mapped to a unified semantic space by independent encoders, and then the contribution weights of each modality are dynamically calculated by a fusion gating network. The final joint representation is the weighted sum of the embeddings of each modality. A Bayesian neural network framework is constructed, a standard normal prior is applied to the network weights, and Monte Carlo dropout is used to approximate Bayesian inference. During the inference phase, the same input is forward propagated, and a different dropout mask is applied each time. The prediction results follow an approximate posterior distribution, the mean of the prediction is the final toxicity probability, and the final toxicity probability is divided into three confidence intervals.

[0046] It should be noted that at least 24 physicochemical descriptors, including logP, molecular weight, number of hydrogen bond donors, number of hydrogen bond acceptors, number of rotatable bonds, topological polar surface area, water solubility (logS), and pKa, were calculated and input into a fully connected network after standardization (Z-score normalization) to form a physicochemical property modality. AC50 values ​​and maximum response values ​​of 23 assay endpoints related to estrogen receptor activation, androgen receptor antagonism, steroidogenesis, cell cycle regulation, oxidative stress response, and DNA damage response were extracted from the ToxCast database to construct an activity feature vector, forming an in vitro bioactivity modality. For missing data, matrix factorization or k-nearest neighbor imputation was used for filling. At least 15 pharmacokinetic and toxicokinetic parameters, including Caco-2 permeability, plasma protein binding rate, CYP450 enzyme inhibition, half-life, and oral bioavailability, were calculated using a pre-trained ADMET prediction model (such as ADMETLab 3.0) to form an ADMET modality. Furthermore, a distance-based application domain evaluation method is employed: the cosine similarity between the chemical to be predicted and the training set in the chemical feature space is calculated, and if the similarity is below a threshold, the system prompts "out of model application domain". The model provides a calibrated uncertainty estimate based on the compound's position in the chemical space, providing a statistical basis for accepting or rejecting the prediction.

[0047] It should be noted that the three confidence intervals are [0, 0.1] for high confidence (green), (0.1, 0.2] for medium confidence (yellow), and >0.2 for low confidence (red) and are output to the user. This method can integrate five modalities of data: molecular structure, physicochemical properties, in vitro biological activity, ADMET, and literature knowledge, to construct a multimodal reproductive toxicity prediction framework. Moreover, the cross-modal attention-weighted fusion gating network can adaptively adjust the contribution of each modality, avoiding the feature redundancy problem of simple splicing and improving the accuracy of source tracing.

[0048] For model learning, alternative solutions can also be set up, including: For the SMILES expression of the input chemical molecule, the RDKit calculation engine is used to generate multiple low-energy three-dimensional conformations. Each conformation is optimized by force field to obtain a stable spatial structure with minimized energy. Based on molecular docking simulation, the dominant conformation with the lowest binding free energy to the key target protein is selected from multiple conformations as the three-dimensional standard conformation, and each atom is encoded as a node and the bond is encoded as an edge. It should be noted that the node feature vector includes at least 32 dimensions of features, such as: atom type (one-hot encoding), hybridization state, formal charge, aromaticity, whether it is a chiral center, hydrogen bond donor / acceptor ability, van der Waals radius, and polarizability. Bonds are encoded as edges, and edge features include at least 8 dimensions, such as: bond type (single, double, triple, or aromatic), bond length (three-dimensional distance), whether it is a conjugated bond, and stereochemical configuration (E / Z configuration).

[0049] A multi-layer graph attention mechanism is adopted. During the node update process, the hidden state of the node is calculated by aggregating the information of its neighboring nodes. After message passing, all node features are concatenated after global average pooling and global max pooling to obtain a molecular graph-level representation vector, which is input to the multi-layer perceptron classifier and outputs the DART toxicity probability. Pre-training was performed using graph contrastive learning, which employed the SimCLR framework. Two enhanced versions of the same molecule with different conformations were considered positive sample pairs, while different molecules were considered negative sample pairs. The InfoNCE loss function was optimized to enable the model to learn a robust 3D molecular structure embedding representation to conformational perturbations.

[0050] It should be noted that by simultaneously introducing three-dimensional molecular conformation information (including chiral centers, bond lengths, and spatial orientation) and force field optimization energy information into the graph neural network framework for reproductive toxicity prediction, the accuracy of the prediction is improved. The introduction of a contrastive learning pre-training strategy enables the model to learn universal three-dimensional molecular structure representations in an unlabeled large-scale chemical space, significantly reducing the dependence on labeled data. Furthermore, the visualization of the attention mechanism can reveal the local atomic regions and chemical substructures that contribute the most to toxicity, providing a certain degree of interpretability for the model.

[0051] Example 2 This embodiment, based on the Python 3.9 programming environment and utilizing computational modules such as RDKit (v2025.9), Scikit-learn, SHAP, and AutoDockVina 1.2.3, implements a method for rapid prediction and mechanism tracing of the reproductive toxicity of chemicals. The specific implementation steps are as follows: Step S1: Construction and Downsampling Equalization of Multi-Source Toxicology Datasets This embodiment first obtains modeling data from public databases. In vitro toxicity data are obtained from the High Throughput Screening (HTS) project in the PubChem database, including the AID743075 dataset characterizing estrogen receptor activation (MIE) and the AID720637 dataset characterizing mitochondrial membrane potential damage (KE); at the same time, reproductive toxicity endpoint (AO) datasets validated by live animal experiments are also collected.

[0052] The SMILES strings of all compounds were standardized using the RDKit library, including: using parsed molecules, removing solvent and inorganic salt ions, and applying dehydrogenation to unify tautomers. To address the extreme class imbalance problem in the in vitro HTS dataset (i.e., negative samples far outnumber positive samples), this embodiment specifically introduces a majority-class downsampling strategy: by calling the Pandas library's Chem.MolFromSmilessample function, a global random seed is set, and negative samples are randomly sampled without replacement, ensuring their number is strictly equal to that of positive samples. This ultimately constructs a 1:1 ratio training set, completely eliminating the prior distribution bias of the classifier. random_state=42.

[0053] Step S2: Concurrent extraction of multi-scale molecular features and generation of interaction fingerprint (PLIF) To comprehensively characterize the physicochemical and spatial properties of chemical substances, this embodiment extracts two-dimensional topological fingerprints and three-dimensional physical interaction matrices: (1) Two-dimensional topological features (ECFP): For each SMILES sequence after the above processing, apply the RDKit algorithm, set the search radius parameter AllChem.GetMorganFingerprintAsBitVectradius=2, fold the bit width, and calculate to generate a 1024-bit binary extended connectivity fingerprint vector (nBits=1024).

[0054] (2) Three-dimensional docking and macroscopic binding energy extraction: The crystal structure of the ESR1 protein ligand binding domain (PDBID: 3ERT) was downloaded from the PDB database. AutoDockTools was used to remove water, add polar hydrogen, and perform Gasteiger charge assignment, saving the data in PDBQT format. The ETKDG algorithm of RDKit was used to generate three-dimensional conformations of over 2000 ligand molecules. The AutoDockVina docking engine was called, setting the center coordinates of the search grid (GridBox) to Center (X=22.63, Y=5.40, Z=22.05), and the grid size to 20Å×20Å×20Å. Semi-flexible docking was performed, and the minimum binding free energy (BindingAffinity) of the optimal binding conformation was extracted.

[0055] (3) This is the core feature of obtaining the ligand binding mechanism in this embodiment. A distance detection algorithm was developed based on the three-dimensional coordinates of the generated receptor-ligand optimal docking complex to evaluate the spatial contact state between the ligand and the 247 amino acid residues of the entire 3ERT protein sequence. The set physicochemical judgment threshold was: the generation of the microscopic polymorphic residue 0 / 1 contact matrix (PLIF). ①: The distance between the ligand donor / acceptor heteroatom and the protein acceptor / donor atom is <3.5 Å, and the bond angle is >130°; hydrogen bonding. ②: The minimum distance between the ligand and the aliphatic or aromatic carbon atom of the protein is <4.5 Å; hydrophobic interaction. ③: The distance between the aromatic ring of the ligand and the centroid of the aromatic ring of the protein residues (such as PHE, TYR, TRP, etc.) is <6.0 Å; π-π stacking interaction. ④: The center-to-center distance between atoms with opposite charges is <4.0 Å. Salt bridge / electrostatic interaction. For more than 2,000 ligands, if a ligand satisfies any of the above-mentioned interaction conditions with a certain amino acid residue of 3ERT, the characteristic bit of that residue is recorded as 1; otherwise, it is recorded as 0. Finally, the contact state of each ligand with respect to 247 residues is mapped to a eigenvector of a 247-bit binary interaction matrix containing 0 / 1 bits.

[0056] Step S3: Construction of a two-level prediction model integrating AOP mechanism probability and interaction matrix (1) Using 1024-bit first-level (AOP path probability prediction) As input, a random forest classifier is trained to predict ESR1 activation and mitochondrial damage. The model calls the output `predict_proba` to determine the mechanism activation probability of each ligand. and.

[0057] (2) The features are horizontally concatenated to form a panoramic feature vector with a dimension of 1274: Level 2 (multi-feature fusion toxicity prediction). This is then input into the final random forest reproductive toxicity classification model. Hyperparameters are set during training: number of decision trees. `n_estimators=100`, maximum tree depth to prevent overfitting from ultra-high dimensional features, minimum number of split samples `max_depth=15` `min_samples_split=5`, and these settings are configured. After 5-fold cross-validation, the ROCAUC score on the independent test set is greater than 0.80. `class_weight='balanced'` Step S4: Bidirectional source tracing analysis of the black-box model based on the SHAP algorithm This embodiment uses the SHAP (Shapley Additive ex Planations) interpreter to analyze the trained fusion model: like Figure 1 , Figure 2 and Figure 4 As shown, on the one hand, the top 10 ECFP fingerprint bits with the highest SHAP value contribution (such as Bit_849, Bit_314, Bit_356) are extracted and reversed to specific chemical fragments using the RDKitDrawMorganBits function, clearly highlighting the nitrogen-containing aromatic groups and carbonyl oxidized structures that cause toxicity (StructuralAlerts). On the other hand, feature columns that significantly contribute to SHAP were extracted, and several core hotspot residues of the 3ERT protein (such as HIS and PHE at specific positions) were precisely located. This proves that the model's predictions do not rely on spurious statistical associations, but rather on the precise alignment of the well-defined chemical backbone with the receptor binding site.

[0058] Step S5: High-throughput screening application for China's existing chemical substance inventory (IECSC) like Figure 4 As shown, a list of compounds SMILES from the IECSC catalog is obtained and input into the aforementioned validated computer system for virtual screening. The model automatically and concurrently performs feature calculations and inferences. For high-risk toxicants with a positive prediction category (Predicted_Tox=1), this embodiment classifies and statistically analyzes their pathogenic mechanisms based on Boolean logic: if only and , it is classified as a single ESR1-mediated type; if and , it is classified as a single mitochondrial damage type; if both are greater than 0.5, it is classified as a dual-target synergistic type. Finally, the system outputs a high-risk reproductive toxicant warning list containing toxicity probability, binding energy, and pathogenic mechanism classification, and visualizes the mechanism distribution in the form of a bar chart, providing a practical solution for environmental screening and source substitution of industrial chemicals in my country.

[0059] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0060] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0061] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0062] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0063] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0064] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for predicting the reproductive toxicity of chemicals by integrating AOP and multi-scale molecular characteristics, characterized in that, Includes the following steps: Acquire multi-source toxicological data, and perform structural standardization, cleaning and deduplication, and sample equalization using a downsampling algorithm on the multi-source toxicological data. For each compound molecule, the Morgan algorithm is used to generate an extended connectivity fingerprint, and a three-dimensional docking panoramic fingerprint and a contact matrix of polymorphic residues are constructed based on the extended connectivity fingerprint. A cascaded machine learning model integrating AOP key events and physical contact matrix is ​​constructed, and the reproductive toxicity prediction probability and adjudication result are output using the cascaded machine learning model integrating AOP key events and physical contact matrix. The SHAP algorithm is used to calculate the expected marginal contribution of the feature subset, and the top-ranking ECFP fingerprint positions are reverse-engineered into specific chemical fragments, and the chemical structure is traced. Input a list of untested chemicals, extract two-dimensional fingerprints, deduce mechanism probabilities, and concurrently execute molecular docking and interaction matrices within the computing environment. Based on fusion characteristics, give a hazard rating and output a list of chemicals with a high probability of toxicity. For each compound molecule, an extended connectivity fingerprint is generated using the Morgan algorithm, and a three-dimensional docking panoramic fingerprint and a contact matrix of polymorphic residues are constructed based on the extended connectivity fingerprint, specifically as follows: The Morgan algorithm is introduced. For each compound molecule, the Morgan algorithm sets the search radius to two chemical bonds and iteratively searches for the structural information of the polymer atoms and their neighborhood environment, with each heavy atom in the molecule as the center. The aggregated information is folded and mapped to a fixed-length binary bit vector of 1024 bits by a hash function. Each bit of the binary bit vector corresponds to the presence or absence state of a specific chemical microenvironment or substructure fragment, forming a two-dimensional topological feature representation. The crystal structure of the human estrogen receptor ligand binding domain was obtained from a protein database. The water of crystallization was removed and polar hydrogen atoms were added using the AutoDockTools program. The Gasteiger bias charge was calculated and converted into the docking-specific PDBQT format. The collected toxicity validation set and candidate ligand molecules SMILES were used to generate 3D conformations using the ETKDG algorithm, and after force field optimization, they were converted into PDBQT format, and the docking grid points were set to cover the binding pocket of the human estrogen receptor protein. Using the AutoDockVina docking engine, the minimum binding free energy of all ligands and receptors is calculated. The detection results of ligands are integrated. When a ligand and a receptor make effective contact with a certain amino acid residue, the corresponding position in the matrix is ​​recorded as 1, otherwise it is recorded as 0. Finally, the interaction states of all compounds with full-length protein residues are transformed and flattened into a binary interaction matrix.

2. The method for predicting the reproductive toxicity of chemicals by integrating AOP and multi-scale molecular characteristics according to claim 1, characterized in that, Acquire multi-source toxicological data, and perform structural standardization, cleaning and deduplication, and sample equalization using a downsampling algorithm on the multi-source toxicological data, specifically as follows: High-throughput in vitro screening data related to the core critical events of reproductive toxicity AOP were obtained from the toxicology database, and simplified molecular linear input canonical strings for all compounds were obtained based on the high-throughput in vitro screening data. The simplified linear input strings of all the compounds were processed using an open-source cheminformatics toolkit to remove inorganic salt ions, neutralize molecular charges, remove heavy metals and organometallic complexes, and unify the expression forms of tautomers. For contradictory test results for the same compound, a "majority voting method" or a safety conservatism principle prioritizing positive results is used to remove duplicates. A downsampling algorithm is implemented on the in vitro MIE and KE datasets. Under the condition that the difference between the number of negative samples and the number of positive samples is greater than a preset threshold, a subset with the same number of positive samples is extracted from the negative sample set using a random sampling algorithm without replacement, and then merged to construct a balanced training set with a ratio of 1:

1.

3. The method for predicting the reproductive toxicity of chemicals by integrating AOP and multi-scale molecular characteristics according to claim 1, characterized in that, A cascaded machine learning model integrating AOP key events and physical contact matrices is constructed. This model outputs reproductive toxicity prediction probabilities and adjudication results, specifically: A two-stage cascaded architecture is adopted to deeply stitch together the structure, mechanism probability and high-dimensional physical contact matrix. In the first stage, the AOP mechanism response classifier is constructed with 1024 bits as input to train random forest classification models for predicting human estrogen receptor activation and mitochondrial damage respectively. The random forest classification model outputs a probability vector of compound-induced human estrogen receptor activation and mitochondrial damage through the predict_proba function. In the second stage, a final reproductive toxicity prediction ensemble model is constructed. For each ligand molecule in the dataset, perform lateral stitching of the feature space to generate a panoramic fusion feature vector, and input the panoramic fusion feature vector into the second-level random forest classifier; Parameters were set to address the class imbalance of in vivo reproductive toxicity data, and the model was fine-tuned through 5-fold cross-validation. The final output of the reproductive toxicity prediction probability and the adjudication result is the model.

4. The method for predicting the reproductive toxicity of chemicals by integrating AOP and multi-scale molecular characteristics according to claim 1, characterized in that, The SHAP algorithm is used to calculate the expected marginal contribution of the feature subset. The top-contributing ECFP fingerprints are then reverse-engineered into specific chemical fragments, and their chemical structures are traced. Specifically: The SHAP algorithm is introduced to calculate the expected marginal contribution of the feature subset, quantify the contribution of each feature to the prediction of reproductive toxicity, and rank the contributions. The two-dimensional topological features of the pre-defined ranking contribution are reverse-engineered to specific chemical fragments. The toxic functional groups are highlighted in color in the molecular skeleton by the RDKit visualization engine, and the source is traced by extracting residues with significant SHAP values.

5. The method for predicting the reproductive toxicity of chemicals by integrating AOP and multi-scale molecular characteristics according to claim 1, characterized in that, Input a list of untested chemicals, extract two-dimensional fingerprints, deduce mechanism probabilities, and concurrently execute large-scale molecular docking and interaction matrices within the computational environment. Based on fusion characteristics, provide hazard ratings and output a list of chemicals with a high probability of toxicity, specifically: Input a massive list of untested chemicals from the existing chemical substance catalog, extract two-dimensional fingerprints, deduce mechanism probabilities, and concurrently execute a large number of molecular docking and interaction matrix transformations within the computing environment; The model provides a hazard rating based on fusion characteristics, outputs a list of chemicals with a high probability of toxicity, and classifies pathogenic mechanisms into single receptor-mediated, mitochondrial damage-mediated, and dual-target synergistic destructive types based on Boolean logic.