Atherosclerosis risk assessment method and system based on harmful outcome path
By dynamically mapping the biological conduction hierarchy of AOP to the computable graph structure of GNN, an end-to-end atherosclerosis risk assessment model is constructed, which solves the problems of missing multi-level causal data and insufficient predictive ability in existing technologies and achieves high-precision risk assessment.
Patent Information
- Application Number
- CN202511140543.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies make it difficult to effectively quantify the early prediction of atherosclerosis risk by exogenous chemicals, and existing machine learning methods make it difficult to construct multi-level causal network structures, resulting in low accuracy in AS risk assessment.
A graph neural network (GNN) is used to dynamically map the biological conduction hierarchy of the adverse outcome pathway (AOP) into a computable graph structure, and to construct an end-to-end quantitative prediction model from the molecular structure characteristics of chemical substances to the risk of atherosclerosis. Risk assessment is performed by extracting molecular structure features and constructing causal pathways.
It achieves high-precision prediction of atherosclerosis risk in the absence of data, provides an explainable mechanism pathway map, breaks through the mechanism modeling bottleneck of traditional models, and improves the accuracy of early risk assessment.
Smart Images

Figure CN120636828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of computational chemistry and computational toxicology, and in particular to a method and system for assessing the risk of atherosclerosis based on harmful outcome pathways. Background Art
[0002] Cardiovascular disease is the leading cause of death worldwide, with atherosclerosis (AS) as its primary pathology. Predicting AS risk has long relied on biochemical markers (such as low-density lipoprotein-cholesterol (LDL)) and statistical models (such as the Framingham score). While these methods can identify traditional risk factors (age, dyslipidemia, etc.), they struggle to quantify the impact of exogenous chemicals (such as environmental pollutants and medications), making them incapable of early prediction of AS risk.
[0003] The mechanism of action of chemical substances inducing AS involves a cross-scale toxic cascade effect, from receptor activation at the molecular level (such as nuclear receptor binding), gene expression perturbation at the cellular level (such as lipid metabolic pathways), to lipid accumulation at the tissue level, ultimately triggering the formation of vascular plaques.
[0004] Existing computational chemistry models attempt to integrate chemical substance data through machine learning (such as random forests and XGBoost) to build chemical substance toxicity prediction models, but ignore the causal cascade relationship between toxic effects, resulting in a lack of mechanistic explainability in the prediction results and low model accuracy.
[0005] In recent years, the adverse outcome pathway (AOP) framework has provided a new approach for modeling the mechanism of AS. It defines the causal chain of "molecular initiating event (MIE) → key events (KEs) → adverse outcome (AO)" to analyze the toxicity pathway. However, the AOP prediction model faces a dual dilemma: First, the heterogeneity of multi-level data (molecular structure characteristics, gene expression, physiological indicators, and disease phenotypes) leads to large data gaps, making it difficult to construct a complete chain of data for the serial AOP framework. Second, existing machine learning methods are difficult to directly model the multi-level causal network structure of AOPs, especially for the toxicity mechanism of multi-organ collaborative processes such as AS. Existing methods have not yet achieved end-to-end quantitative prediction from chemical exposure to plaque formation.
[0006] Therefore, the existing AOP-based AS risk assessment technology has bottlenecks such as insufficient early prediction capabilities of chemical substances and lack of cross-level causal data and models. Summary of the Invention
[0007] In order to solve the above problems, the present invention proposes an atherosclerosis risk assessment method and system based on adverse outcome pathways. By dynamically mapping the biological conduction hierarchy of AOPs into a computable graph structure of a graph neural network (GNN), end-to-end quantitative prediction from the molecular structure characteristics of chemical substances to AS disease risks is achieved, providing a high-precision computing tool for drug safety evaluation and environmental pollutant screening.
[0008] According to some embodiments, the present invention adopts the following technical solutions: Atherosclerosis risk assessment methods based on adverse outcome pathways include: Extract the molecular structure characteristics of the chemical substances to be tested; The extracted molecular structure features are input into the trained risk scoring model to obtain the risk score of atherosclerosis; Among them, the risk scoring model is based on molecular structural characteristics, with the interference between chemical substances and biological macromolecules as the molecular initiating event MIE, the changes in gene expression of nuclear receptor downstream regulatory pathways, tissue hormones or marker levels as key events KEs, and the risk of atherosclerosis as the harmful outcome AO. A causal path based on the harmful outcome path MIE→KEs→AO is constructed, and the constructed causal path is input into the graph neural network to obtain the risk score of atherosclerosis.
[0009] According to some embodiments, the present invention adopts the following technical solutions: Atherosclerosis risk assessment system based on adverse outcome pathways, including: The feature extraction module is configured to: extract the molecular structure features of the chemical substance to be tested; The risk scoring module is configured to: input the extracted molecular structure features into the trained risk scoring model to obtain a risk score for atherosclerosis; Among them, the risk scoring model is based on molecular structural characteristics, with the interference between chemical substances and biological macromolecules as the molecular initiating event MIE, the changes in gene expression of nuclear receptor downstream regulatory pathways, tissue hormones or marker levels as key events KEs, and the risk of atherosclerosis as the harmful outcome AO. A causal path based on the harmful outcome path MIE→KEs→AO is constructed, and the constructed causal path is input into the graph neural network to obtain the risk score of atherosclerosis.
[0010] According to some embodiments, the present invention adopts the following technical solutions: A computer program product includes a computer program, which implements the method for assessing atherosclerosis risk based on harmful outcome pathways when the computer program is executed by a processor.
[0011] According to some embodiments, the present invention adopts the following technical solutions: A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method for assessing the risk of atherosclerosis based on a harmful outcome pathway is implemented.
[0012] According to some embodiments, the present invention adopts the following technical solutions: An electronic device comprises: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the atherosclerosis risk assessment method based on the harmful outcome pathway.
[0013] Compared with the prior art, the present invention has the following beneficial effects: This method requires only the chemical structure (SMILES code) to initiate predictions. Addressing the common problem of missing AOP data, it integrates a machine learning model prediction infill with a GNN multi-task loss function to maintain an R² of above 0.7 even with significant data loss, addressing the data-dependent challenge of early chemical prediction. The GNN inter-layer weight matrix explicitly quantifies the intensity of event transmission (e.g., the contribution of gene expression to lipid accumulation), providing an interpretable mechanistic pathway map and overcoming the mechanistic modeling bottleneck of traditional models. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0015] Figure 1 This is a flow chart of the method of Example 1.
[0016] Figure 2 This is the AOP framework diagram of Example 1. Figure 3 This is the GNN model structure diagram of Example 1. DETAILED DESCRIPTION The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0017] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0018] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "comprising" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0019] Example 1 In one embodiment of the present invention, a method for assessing the risk of atherosclerosis based on a harmful outcome pathway is provided, comprising: Step S1: extracting the molecular structure characteristics of the chemical substance to be tested; Furthermore, the chemical substance is represented by a SMILES code, and the molecular structure characteristics are calculated by a molecular fingerprint algorithm. In this embodiment, the molecular structure characteristics adopt Morgan fingerprints.
[0020] Step S2: Inputting the extracted molecular structure features into the trained risk scoring model to obtain the risk score of atherosclerosis; Among them, the risk scoring model is based on molecular structural characteristics, with the interference between chemical substances and biological macromolecules as the molecular initiating event MIE, the changes in gene expression of nuclear receptor downstream regulatory pathways, tissue hormones or marker levels as key events KEs, and the risk of atherosclerosis as the harmful outcome AO. A causal path based on the harmful outcome path MIE→KEs→AO is constructed, and the constructed causal path is input into the graph neural network to obtain the risk score of atherosclerosis.
[0021] Furthermore, the risk scoring model includes a causal path construction sub-model and a graph neural network scoring model; The causal pathway construction sub-model predicts MIE and KEs based on molecular structure characteristics, and uses the predicted values to construct the causal pathway of MIE→KEs→AO. The graph neural network scoring model is a graph neural network that performs risk scoring on the constructed causal path.
[0022] Furthermore, the prediction of MIE and KEs based on molecular structure features is to construct a regression model using a machine learning algorithm and fill in the predicted values.
[0023] Furthermore, the topological structure of the causal path is fixed to a linear chain structure, with nodes arranged in the order of MIE, KEs, and AO; The graph neural network scoring model consists of a GraphSAGE layer and a Transformer attention layer. The GraphSAGE layer is used to transmit and weightedly aggregate information of adjacent nodes on the event path to simulate the causal transmission process between events. The graph attention layer is used to model global dependencies and capture interactions between remote nodes. Furthermore, the risk scoring model uses weighted losses of MIE, KEs, and AO for model optimization training.
[0024] As an embodiment, the atherosclerosis risk assessment method based on the harmful outcome pathway of the present invention dynamically maps the biological conduction level of AOP to the computable graph structure of GNN, thereby realizing end-to-end quantitative prediction from the molecular structure characteristics of chemical substances to AS disease risk, providing a high-precision calculation tool for drug safety evaluation and environmental pollutant screening. From the perspective of model construction, Figure 1 The specific implementation process is as follows: (1) Establish an adverse outcome pathway (AOP) framework: define molecular initiating events (MIEs), key events (KEs) and adverse outcomes (AOs), and establish Figure 2 The causal path of MIE→KEs→AO is shown; (2) AOP data collection and processing: Collect biochemical test experimental data and corresponding chemical substances of MIE, KEs and AO in the AOP framework, extract molecular structure features based on the SMILES code of chemical substances, and use the extracted molecular structure features to process missing values of other data (i.e., MIE, KEs, AO) through regression model (i.e., causal path construction sub-model); (3) Construction of the graph neural network (GNN) model framework, namely the graph neural network scoring model: Based on the MIE→KEs→AO path of AOP, the graph structure of the GNN model input layer is constructed. GraphSAGE and Transformer attention are used as the core architecture of the GNN model. The output layer is an independent multi-task of MIE, KEs and AO, with shared feature extraction; (4) GNN prediction model training optimization and evaluation: The processed AOP data are divided into data sets and then GNN model training is carried out. The Huber loss function is used to optimize the AO task. During the training process, the model parameters are adjusted based on the multi-task prediction accuracy. Finally, the atherosclerosis risk value of the new chemical substance is predicted.
[0025] The following is a detailed description of data preparation, model building and training: 1. Data preparation and regression model construction Atherosclerotic AS risk data of several chemical substances were obtained from the CTD database, and the unique digital identifiers (CIDs) and SMILES codes of the chemical substances were extracted from the PubChem database according to the chemical names.
[0026] The molecular initiating event (MIE) uses PXR binding affinity and PXR gene expression to characterize the interference between chemical substances and biological macromolecules. The key events (KEs) include KE1 and KE2, which use lipid metabolism gene expression and liver lipid content, respectively, to characterize changes in gene expression, tissue hormone or marker levels in the downstream regulatory pathways of nuclear receptors. The adverse outcome (AO) is a risk score extracted from the atherosclerosis (AS) risk data of the CTD database to characterize the risk of atherosclerosis.
[0027] There are two sources of data for MIE: PXR binding affinity is calculated based on molecular docking, and PXR gene expression is directly obtained from the PubChem database.
[0028] Specifically, the chemical substances were converted into .pdbqt files based on their SMILES codes. Based on the binding pocket of the pregnane X receptor (PXR) protein crystal structure, the binding site coordinates and box size were set for molecular docking, and batch molecular docking of the chemical substances with the pregnane X receptor (PXR) was performed. After molecular docking, the binding affinity of the multiple binding conformations (poses) generated for each chemical (4915) was calculated using the Vina scoring function, and the lowest binding affinity was selected as the PXR binding affinity data for MIE.
[0029] Get biological test data of chemical substances from PubChem database, including multiple Excel files: MIE: From the Excel file of PXR gene expression (Gene ID 8856), find the corresponding column named acvalue by the CID of the chemical substance and use it as the PXR gene expression value of MIE; KE1: Lipid metabolism gene expression, including CYP3A4 gene expression (Gene ID 1576) and CYP3A5 gene expression (Gene ID 1577). The KE1 value is obtained by searching the corresponding columns named Activity Value and AcValue in the Excel files for CYP3A4 and CYP3A5 gene expression, using the chemical substance CID identifier. KE2: Liver lipid content, including HDL (Assay ID 172026) and LDL (Assay ID 172041). The KE2 value is obtained by searching the Standard Value column for each chemical in the Excel file for HDL and LDL using the chemical's CID. The molecular structure characteristics of chemical substances, namely Morgan fingerprints, are extracted based on SMILES codes. Together with the above-mentioned CID, SMILES, MIE, KEs, and AO data, a data set is formed. The missing MIE, KE1, and KE2 data in the data set are filled using a machine learning regression model. The regression model can be constructed based on the LinearRegression, LassoCV, SVR, KNeighborsRegressor, RandomForestRegressor, LGBMRegressor, XGBRegressor, and MLPRegressor machine learning algorithms.
[0030] The data set is divided into training set and test set, and based on the test data values, outlier processing is performed using methods such as histogram and QQ graph.
[0031] According to RMSE and Evaluation indicators, screening the best algorithm, further using network search to tune hyperparameters, to obtain the optimal regression model, as follows: PXR gene expression: RandomForestRegressor model, training set =0.894; CYP3A4 gene expression: RandomForestRegressor model, training set =0.905; CYP3A5 gene expression: RandomForestRegressor model, training set =0.808; Liver HDL content: RandomForestRegressor model, training set =0.720; Liver LDL content: RandomForestRegressor model, training set =0.855; After filling the missing data in the dataset through the regression model constructed above, a GNN training dataset that runs through all levels of the AOP framework is finally formed. Each substance in the training dataset contains SMILES code, PXR binding affinity, PXR gene expression, CYP3A4 gene expression, CYP3A5 gene expression, liver HDL content, liver LDL content and AS risk data.
[0032] 2. GNN model construction and training GNN model architecture, such as Figure 3 Shown, including: Input layer: A chain topology is used to directly simulate the causal path mechanism in AOP. Each graph has four nodes, each representing a chemical substance: [MIE→KE1→KE2→AO]. In the [MIE→KE1→KE2→AO] path, the values of other quantities in the path are filled in, the quantity to be predicted is set to empty, and the empty value is input into the GNN for prediction.
[0033] Core GNN architecture: 5 GraphSAGE layers + 3 Transformer attention layers: First, a projection layer is used to uniformly map the input dimensions to a high-dimensional space. Then, five stacked GraphSAGE layers are used to gradually aggregate the node's neighbor information to enhance the expressiveness of local structural features. A three-layer Transformer is introduced, each layer containing a multi-head attention mechanism and a feedforward network to model global dependencies and capture interactions between remote nodes. LayerNorm and residual connections are integrated internally to help stabilize the training process, accelerate convergence, and alleviate internal covariate shift. The ReLU activation function is added after each fully connected layer, and Dropout regularization is set to prevent overfitting.
[0034] Output layer: Set up a multi-task model, namely MIE, KE1, KE2 and AS risk. The four independent linear tasks correspond to each level of AOP, share the feature extraction backbone network, and finally output AS risk.
[0035] GNN model training: The model ultimately optimizes the prediction of AS risk, using Huber Loss as the loss function to improve robustness to outliers. The AdamW optimizer is used, combined with the CosineAnnealingLR learning rate scheduler. AMP automatic mixed precision (torch.amp.autocast) is enabled during training to improve training efficiency, and GradScaler is used to dynamically scale gradients to prevent accuracy loss. Each epoch traverses all sample graphs, going through GNN encoding → AO prediction → loss calculation → backpropagation → parameter update.
[0036] Based on the above AOP data and GNN architecture, the model performs as follows on the training set: Prediction of PXR binding affinity The root mean square error (RMSE) was 0.9817 and 0.2464 respectively; is 0.8887, RMSE is 3.7413; AOrisk prediction The p-value is 0.8924 and the RMSE is 3.2290, which shows that the model has a strong fitting ability for the training data.
[0037] Further verify the AOrisk prediction performance on the test set, The value of the proposed method reaches 0.7159, indicating that the model has good generalization ability and risk prediction effectiveness. This result verifies that the proposed method has high reliability in actual risk screening tasks, and is particularly suitable for risk warning scenarios of new chemical substances that lack systematic toxicological data.
[0038] Example 2 In one embodiment of the present invention, an atherosclerosis risk assessment system based on a harmful outcome pathway is provided, comprising: The feature extraction module is configured to: extract the molecular structure features of the chemical substance to be tested; The risk scoring module is configured to: input the extracted molecular structure features into the trained risk scoring model to obtain a risk score for atherosclerosis; Among them, the risk scoring model is based on molecular structural characteristics, with the interference between chemical substances and biological macromolecules as the molecular initiating event MIE, the changes in gene expression of nuclear receptor downstream regulatory pathways, tissue hormones or marker levels as key events KEs, and the risk of atherosclerosis as the harmful outcome AO. A causal path based on the harmful outcome path MIE→KEs→AO is constructed, and the constructed causal path is input into the graph neural network to obtain the risk score of atherosclerosis.
[0039] Example 3 In one embodiment of the present invention, a computer program product is provided, comprising a computer program. When the computer program is executed by a processor, the method for assessing the risk of atherosclerosis based on a harmful outcome pathway is implemented.
[0040] Example 4 In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method for assessing the risk of atherosclerosis based on a harmful outcome pathway is implemented.
[0041] Example 5 In one embodiment of the present invention, an electronic device is provided, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device implements the method for assessing atherosclerosis risk based on harmful outcome pathways.
[0042] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0043] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0044] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A method for assessing atherosclerosis risk based on a harmful outcome pathway, characterized in that: include: Extract the molecular structure characteristics of the chemical substances to be tested; The extracted molecular structure features are input into the trained risk scoring model to obtain the risk score of atherosclerosis; Among them, the risk scoring model is based on molecular structural characteristics, with the interference between chemical substances and biological macromolecules as the molecular initiating event MIE, the changes in gene expression of nuclear receptor downstream regulatory pathways, tissue hormones or marker levels as key events KEs, and the risk of atherosclerosis as the harmful outcome AO. A causal path based on the harmful outcome path MIE→KEs→AO is constructed, and the constructed causal path is input into the graph neural network to obtain the risk score of atherosclerosis.
2. The method for assessing atherosclerosis risk based on a harmful outcome pathway according to claim 1, wherein: The chemical substances are characterized by SMILES codes, and the molecular structure characteristics are calculated using a molecular fingerprint algorithm.
3. The method for assessing atherosclerosis risk based on adverse outcome pathways according to claim 1, wherein: The risk scoring model includes a causal path construction sub-model and a graph neural network scoring model; The causal pathway construction sub-model predicts MIE and KEs based on molecular structure characteristics, and uses the predicted values to construct the causal pathway of MIE→KEs→AO. The graph neural network scoring model is a graph neural network that performs risk scoring on the constructed causal path.
4. The method for assessing atherosclerosis risk based on adverse outcome pathways according to claim 3, wherein: The prediction of MIE and KEs based on molecular structure features is to construct a regression model using a machine learning algorithm and fill in the predicted values.
5. The method for assessing atherosclerosis risk based on harmful outcome pathways according to claim 3, wherein: The topology of the causal path is fixed as a linear chain structure, with nodes arranged in the order of MIE, KEs, and AO; The graph neural network scoring model consists of a GraphSAGE layer and a Transformer attention layer. The GraphSAGE layer is used to transmit and weightedly aggregate information of adjacent nodes on the event path to simulate the causal transmission process between events. The graph attention layer is used to model global dependencies and capture interactions between remote nodes.
6. The method for assessing atherosclerosis risk based on adverse outcome pathways according to claim 1, wherein: The risk scoring model uses weighted losses of MIE, KEs, and AO for model optimization training.
7. A system for assessing the risk of atherosclerosis based on a harmful outcome pathway, characterized in that: include: The feature extraction module is configured to: extract the molecular structure features of the chemical substance to be tested; The risk scoring module is configured to: input the extracted molecular structure features into the trained risk scoring model to obtain a risk score for atherosclerosis; Among them, the risk scoring model is based on molecular structural characteristics, with the interference between chemical substances and biological macromolecules as the molecular initiating event MIE, the changes in gene expression of nuclear receptor downstream regulatory pathways, tissue hormones or marker levels as key events KEs, and the risk of atherosclerosis as the harmful outcome AO. A causal path based on the harmful outcome path MIE→KEs→AO is constructed, and the constructed causal path is input into the graph neural network to obtain the risk score of atherosclerosis.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for assessing atherosclerosis risk based on harmful outcome pathways according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method for assessing atherosclerosis risk based on a harmful outcome pathway according to any one of claims 1 to 6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the atherosclerosis risk assessment method based on the harmful outcome pathway as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Social media content propagation prediction method and device
CN115759482A
Intelligent construction method for optimal control list of 5-hydroxytryptamine reuptake inhibitor
CN116705143A
Method and device for determining course video, equipment and storage medium
CN117253386A
Attention mechanism driven graph neural network power grid state prediction system
CN117743799A
Load prediction method and device based on GraphSAGE and Transform combined model
CN118472931A