Reaction site prediction method and device based on chemical and physical prior driving

By employing a chemical physics prior-driven approach, molecular data is represented using molecular diagrams, SMILES sequences, and three-dimensional conformations to generate atomic embedding and mixing distance features. By combining charge difference and quantum chemical prior regularized attention weights, the problem of insufficient utilization of topological information in existing methods is solved, and high-precision and interpretable reaction site prediction is achieved.

CN120877896APending Publication Date: 2025-10-31烟台国工智能科技有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510983169.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing methods lack full utilization of topological information in chemical structure diagrams in molecular reaction modeling, have weak reaction context modeling capabilities, poor model interpretability, and difficulty in accurately predicting reaction sites.

Method used

A chemical physics-based prior-driven approach is adopted to represent molecular data through molecular graphs, SMILES sequences, and three-dimensional conformations. Atom embeddings are generated using message-passing neural networks, and mixed distance features and graph position encoding are calculated. Attention scores are calculated by incorporating charge difference, and dual-task joint training is performed. Attention weights are regularized using quantum chemistry priors to generate thermal icons to annotate highly active reaction sites.

Benefits of technology

It enhances the predictability and automated design capabilities of molecular reactions, improves prediction accuracy, reduces the need for training data, and generates heatmaps with high accuracy and mechanistic interpretability, making it suitable for drug active site screening and catalytic reaction design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877896A_ABST
    Figure CN120877896A_ABST
Patent Text Reader

Abstract

The invention discloses a reaction site prediction method and device based on chemical and physical prior driving, and the method comprises the steps: extracting set features through the multi-modal input of a fusion molecular map, an SMILES sequence and a three-dimensional conformation; generating atomic embedding by using a message passing neural network, and calculating a mixed feature fusing a topological path and a three-dimensional distance; combining the key type weight to construct a graph position code of chemical environment correction; injecting the mixed distance and the charge difference into a Transform attention mechanism, and explicitly modeling an inter-atomic long-range electron effect; a model is jointly trained through double tasks of comparative learning and mask prediction, the comparative learning adopts a directional negative sample to enhance generalization, and mask prediction synchronously recovers an atom type and a charge transfer matrix; and finally, injecting quantum chemistry priori constraint attention weights such as a Fuzzy well function, outputting an atomic-scale reaction activity probability, generating a thermodynamic diagram, and realizing high-precision and interpretable active site labeling. According to the method, the drug design and reaction mechanism analysis efficiency can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computational chemistry, specifically to a method and apparatus for predicting reaction sites based on chemical physics priors. Background Technology

[0002] Molecular reaction modeling is a core technology in fields such as drug synthesis and materials design. Accurately identifying active sites (i.e., atoms or bonds that may undergo chemical changes) in reactant molecules is crucial for new drug development, materials design, and reaction mechanism research. Traditional methods typically rely on the following two approaches:

[0003] One approach is the rule-based method based on expert experience. This type of method relies on manually set response rules or hand-constructed response templates to search for matching structures that conform to the rules in a database. Although intuitive and highly interpretable, it has poor generalization ability, struggles to cover novel or complex structures, and is highly dependent on expert knowledge.

[0004] Second, methods based on quantum chemical calculations. Quantum chemical calculations can predict reactive sites with relatively high accuracy, for example, by calculating parameters such as molecular orbitals, charge distribution, and Fukui functions. However, these methods are computationally intensive and inefficient, making them difficult to apply to large-scale molecular screening or high-throughput reaction prediction tasks.

[0005] Currently, with the development of deep learning technology, more and more research is applying machine learning models (such as graph neural networks and recurrent neural networks) to chemical structure data analysis. In recent years, the Transformer model, due to its superior performance in natural language processing (NLP), has been gradually introduced into molecular sequence representation (such as SMILES), reaction prediction, and molecular property prediction tasks. Transformers can model long-distance dependencies and have significant advantages in structural characterization. However, existing methods still have the following shortcomings in applying the Transformer architecture to molecular reaction modeling: first, they lack full utilization of topological information in chemical structure diagrams; second, their reaction context modeling ability is weak, making it difficult to distinguish active sites in different environments within the same molecule; and third, the model interpretability is poor, failing to provide effective explanations of structure-activity relationships.

[0006] Therefore, how to invent a reaction site prediction method based on chemical physics priors that can improve the predictability and automated design capability of molecular reactions has become an urgent problem to be solved. Summary of the Invention

[0007] To address this, the present invention provides a reaction site prediction method and apparatus based on chemical physics priors, which solves the problems of existing technologies lacking full utilization of topological information in chemical structure diagrams, having weak reaction context modeling capabilities, and poor model interpretability.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a reaction site prediction method based on chemical physics prior driving, comprising:

[0009] Molecular data is represented and input using molecular diagrams, SMILES sequences, and three-dimensional conformations; defined features of the molecular data are extracted; and atomic embeddings are generated using a message-passing neural network based on these defined features.

[0010] Based on the atomic embedding, the hybrid distance feature of the fused topological shortest path and three-dimensional Euclidean distance is calculated and obtained;

[0011] Based on the aforementioned mixed distance features, and combined with bond type weights, a chemical environment-corrected graph position encoding is constructed.

[0012] Based on the graph position encoding, the hybrid distance features and charge difference are incorporated into the attention score calculation formula to obtain an optimized attention mechanism;

[0013] Based on the optimized attention mechanism, the reaction site prediction model is jointly trained in two tasks to obtain the trained reaction site prediction model.

[0014] The attention weights are regularized based on quantum chemical priors, and the atomic-level reaction activity probability is predicted using the trained reaction site prediction model. A heatmap is generated based on the atomic-level reaction activity probability, and highly active reaction sites are labeled using the heatmap.

[0015] As a preferred embodiment of a reaction site prediction method driven by prior knowledge of chemical physics, the defined features include: atom type, bond type, charge, and quantum interaction energy characteristics;

[0016] The expression for the atomic embedding is:

[0017]

[0018] In the formula, Let i be the feature vector of atom i at the l-th layer of MPNN; Let i be the feature vector of atom i in the (l-1)th layer; W is the eigenvector of neighboring atom j in layer l-1; (l) Let N be the weight matrix of the l-th layer; N(i) be the set of neighbor nodes of atom i; h ij E represents the eigenvector of the edge between atom i and its neighbor atom j; quantum (i,j) represents the quantum interaction energy between atom i and atom j obtained by DFT calculation; ReLU is the linear rectified function; MLP is the multilayer perceptron.

[0019] As a preferred embodiment of the reaction site prediction method based on chemical physics priors, the formula for calculating the mixing distance feature is as follows:

[0020]

[0021] In the formula, It is a mixed distance feature; This is the topological shortest path; α represents the three-dimensional Euclidean distance; α is the weighting coefficient.

[0022] As a preferred embodiment of the reaction site prediction method based on chemical physics priors, the expression for the graph location encoding is:

[0023]

[0024] In the formula, GPE ij The resulting vector encodings of the graph positions of atoms i and j; It is a mixed distance feature; It is a chemical environment correction factor.

[0025] As a preferred embodiment of the reaction site prediction method driven by prior chemical physics, in the process of incorporating the mixing distance feature and the charge difference into the attention score calculation formula, the expression of the attention score calculation formula is as follows:

[0026]

[0027] In the formula, Attention(Q,K,V) represents the calculation result of the multi-head self-attention mechanism; Q,K,V are the query, key, and value matrices, respectively; α is the hyperparameter; q i q j d represents the charges of atoms i and j, respectively; k Let K be the dimension of the key matrix K; T is the transpose symbol.

[0028] As a preferred embodiment of the reaction site prediction method based on chemical physics priors, in the process of performing the dual-task joint training of the reaction site prediction model based on the optimized attention mechanism, the dual-task joint training includes: a contrastive learning task and a mask prediction task.

[0029] The loss function expression for the contrastive learning task is:

[0030]

[0031] In the formula, L contrast The loss function for contrastive learning; sim(x) + ,x +) represents the similarity between positive samples; τ represents the temperature coefficient; sim(x) + ,x - s(x) represents the similarity between positive and negative samples; + ,x - () represents the chemical similarity between positive and negative samples calculated based on RDKit fingerprints;

[0032] The loss function expression for the mask prediction task is:

[0033]

[0034] In the formula, L mask Let be the loss function for mask prediction; BCE be the binary cross-entropy loss function; λ be the hyperparameter; MSE be the mean squared error loss function; and T be the charge transfer probability matrix. This is the true matrix.

[0035] As a preferred approach for reaction site prediction based on chemical physics priors, the expression for regularizing the attention weights according to quantum chemical priors is as follows:

[0036]

[0037] In the formula, L explain Let A be the loss function used to constrain the attention weights; A is the attention weight matrix learned by the model; A chem This is a theoretical interaction matrix based on chemical experience; ||·|| F For Frobenius van; A chem,ij Let be the theoretical interaction matrix between atom i and atom j; γ, η are trainable parameters.

[0038] As a preferred embodiment of the reaction site prediction method driven by prior chemical physics, the expression for the atomic-level reaction activity probability is:

[0039] p i =σ(W p ·h i +b p )

[0040] In the formula, p i σ represents the probability that an atom is an active site; σ is the Sigmoid function; W p b p h represents the trainable weight matrix and bias vector. i is the eigenvector of the atom.

[0041] This invention also provides a reaction site prediction device based on chemical physics prior-driven methods, comprising:

[0042] The molecular data input and processing module is used to represent and input molecular data through molecular diagrams, SMILES sequences, and three-dimensional conformations; extract predetermined features from the molecular data; and generate atom embeddings based on the predetermined features using a message-passing neural network.

[0043] The hybrid distance feature calculation module is used to calculate and obtain the hybrid distance feature that combines the topological shortest path and the three-dimensional Euclidean distance based on the atomic embedding.

[0044] The graph position coding construction module is used to construct a chemical environment-corrected graph position code based on the mixed distance features and bond type weights.

[0045] An attention mechanism optimization module is used to incorporate the hybrid distance features and charge difference into the attention score calculation formula based on the graph position encoding to obtain an optimized attention mechanism.

[0046] The model training module is used to perform dual-task joint training on the response site prediction model based on the optimized attention mechanism to obtain the trained response site prediction model.

[0047] The reaction site prediction module is used to regularize the attention weights based on quantum chemical priors and predict the atomic-level reaction activity probability using the trained reaction site prediction model; generate a heatmap based on the atomic-level reaction activity probability; and label highly active reaction sites using the heatmap.

[0048] As a preferred embodiment of a reaction site prediction device driven by prior knowledge of chemical physics, the set features in the molecular data input and processing module include: atom type, bond type, charge and quantum interaction energy characteristics;

[0049] The expression for the atomic embedding is:

[0050]

[0051] In the formula, Let i be the feature vector of atom i at the l-th layer of MPNN; Let i be the feature vector of atom i in the (l-1)th layer; W is the eigenvector of neighboring atom j in layer l-1; (l) Let N be the weight matrix of the l-th layer; N(i) be the set of neighbor nodes of atom i; h ij E represents the eigenvector of the edge between atom i and its neighbor atom j; quantum (i,j) represents the quantum interaction energy between atom i and atom j obtained by DFT calculation; ReLU is the linear rectified function; MLP is the multilayer perceptron.

[0052] As a preferred embodiment of a reaction site prediction device driven by prior chemical physics, the formula for calculating the mixing distance feature in the mixing distance feature calculation module is as follows:

[0053]

[0054] In the formula, It is a mixed distance feature; This is the topological shortest path; α represents the three-dimensional Euclidean distance; α is the weighting coefficient.

[0055] As a preferred embodiment of a reaction site prediction device driven by prior chemical physics, the graph location encoding expression in the graph location encoding construction module is:

[0056]

[0057] In the formula, GPE ij The resulting vector encodings of the graph positions of atoms i and j; It is a mixed distance feature; It is a chemical environment correction factor.

[0058] As a preferred embodiment of a reaction site prediction device driven by prior chemical physics, in the attention mechanism optimization module, during the process of incorporating the mixing distance feature and the charge difference into the attention score calculation formula, the expression of the attention score calculation formula is as follows:

[0059]

[0060] In the formula, Attention(Q,K,V) represents the calculation result of the multi-head self-attention mechanism; Q,K,V are the query, key, and value matrices, respectively; α is the hyperparameter; q i q j d represents the charges of atoms i and j, respectively; k Let K be the dimension of the key matrix K; T is the transpose symbol.

[0061] As a preferred embodiment of a reaction site prediction device driven by chemical physics priors, in the model training module, during the dual-task joint training of the reaction site prediction model based on the optimized attention mechanism, the dual-task joint training includes: a contrastive learning task and a mask prediction task.

[0062] The loss function expression for the contrastive learning task is:

[0063]

[0064] In the formula, Lcontrast The loss function for contrastive learning; sim(x) + ,x + ) represents the similarity between positive samples; τ represents the temperature coefficient; sim(x) + ,x - s(x) represents the similarity between positive and negative samples; + ,x - () represents the chemical similarity between positive and negative samples calculated based on RDKit fingerprints;

[0065] The loss function expression for the mask prediction task is:

[0066]

[0067] In the formula, L mask Let be the loss function for mask prediction; BCE be the binary cross-entropy loss function; λ be the hyperparameter; MSE be the mean squared error loss function; and T be the charge transfer probability matrix. This is the true matrix.

[0068] As a preferred embodiment of a reaction site prediction device driven by chemical physics priors, the expression for regularizing the attention weights based on quantum chemical priors in the reaction site prediction module is as follows:

[0069]

[0070] In the formula, L explain Let A be the loss function used to constrain the attention weights; A is the attention weight matrix learned by the model; A chem This is a theoretical interaction matrix based on chemical experience; ||·|| F For Frobenius van; A chem,ij Let be the theoretical interaction matrix between atom i and atom j; γ, η are trainable parameters.

[0071] As a preferred embodiment of a reaction site prediction device driven by prior knowledge of chemical physics, the expression for the atomic-level reaction activity probability in the reaction site prediction module is as follows:

[0072] p i =σ(W p ·h i +b p )

[0073] In the formula, p i σ represents the probability that an atom is an active site; σ is the Sigmoid function; W p b p h represents the trainable weight matrix and bias vector. i is the eigenvector of the atom.

[0074] This invention has the following advantages: It represents and inputs molecular data through molecular graphs, SMILES sequences, and three-dimensional conformations; it extracts and obtains defined features from the molecular data; based on these features, it generates atomic embeddings using a message-passing neural network; based on these atomic embeddings, it calculates a hybrid distance feature that integrates the topological shortest path and the three-dimensional Euclidean distance; based on this hybrid distance feature, it constructs a chemical environment-corrected graph position code by combining bond type weights; based on this graph position code, it integrates the hybrid distance feature and charge difference into the attention score calculation formula to obtain an optimized attention mechanism; based on this optimized attention mechanism, it performs dual-task joint training on the reaction site prediction model to obtain a trained reaction site prediction model; it regularizes the attention weights according to quantum chemical priors and predicts the output atomic-level reaction activity probability using the trained reaction site prediction model; based on the atomic-level reaction activity probability, it generates a heatmap; and it labels highly active reaction sites using the heatmap. This invention constructs atomic-level mixed distance features through multimodal data fusion (molecular graph topology, SMILES sequences, and 3D conformations), overcoming the limitations of single characterization and comprehensively capturing the spatial and electronic properties of chemical structures. It innovatively introduces bond type correction factors and charge difference coupling mechanisms, establishing a chemical physics-driven attention weight allocation rule within the Transformer architecture to explicitly model long-range interatomic interactions (such as the conjugation effect in aromatic systems and charge attraction of polar groups), thus addressing the insensitivity of traditional graph neural networks to electron transfer. It also designs directional negative sample contrastive learning and electron transfer matrix prediction. The dual-task framework generates mechanism-level negative samples by perturbing key functional groups and forces the model to synchronously recover the atomic type and charge density distribution, making the learning process deeply consistent with the essential laws of electron rearrangement in chemical reactions. Finally, it combines quantum chemical priors such as the Fukui function to regularize the attention heatmap, transforming black-box prediction into a mechanism-verifiable atomic-level reaction path. This invention can improve prediction accuracy and reduce training data requirements, making the prediction results highly consistent with the reaction energy barrier distribution calculated by DFT, providing an industrial-grade tool with both high accuracy and mechanism interpretability for drug active site screening and catalytic reaction design. Attached Figure Description

[0075] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0076] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0077] Figure 1 This is a flowchart illustrating the reaction site prediction method based on chemical physics prior driving provided in Embodiment 1 of the present invention.

[0078] Figure 2 This is a schematic diagram of the architecture of the reaction site prediction device based on chemical physics prior driving provided in Embodiment 2 of the present invention. Detailed Implementation

[0079] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0080] Example 1

[0081] See Figure 1 Embodiment 1 of the present invention provides a reaction site prediction method based on chemical physics prior driving, comprising the following steps:

[0082] S1. Represent and input molecular data using molecular diagrams, SMILES sequences, and three-dimensional conformations; extract predetermined features from the molecular data; and generate atom embeddings based on the predetermined features using a message-passing neural network.

[0083] S2. Based on the atomic embedding, calculate and obtain the hybrid distance feature that combines the topological shortest path and the three-dimensional Euclidean distance;

[0084] S3. Based on the aforementioned mixed distance features and combined with bond type weights, construct a chemical environment-corrected graph position encoding;

[0085] S4. Based on the graph position encoding, the hybrid distance features and charge difference are incorporated into the attention score calculation formula to obtain the optimized attention mechanism;

[0086] S5. Based on the optimized attention mechanism, perform dual-task joint training on the reaction site prediction model to obtain the trained reaction site prediction model.

[0087] S6. Regularize the attention weights according to quantum chemical priors, and predict the atomic-level reaction activity probability using the trained reaction site prediction model; generate a heatmap based on the atomic-level reaction activity probability; and label the highly active reaction sites using the heatmap.

[0088] In this embodiment, in step S1, molecular data is represented and input using molecular diagrams, SMILES sequences, and three-dimensional conformations; predetermined features of the molecular data are extracted; and atomic embeddings are generated using a message passing neural network based on the predetermined features.

[0089] Specifically, the input representation of molecular data is achieved through three complementary forms: molecular graphs, SMILES sequences, and three-dimensional conformations. Molecular graphs construct a topological structure with atoms as nodes and chemical bonds as edges. Node features include atom type (e.g., C, O, N), charge (normalized to the [-1, 1] interval), and hybridization state (sp). 2 sp 3 The SMILES sequence encodes the bond type (single bond, double bond, etc.), conjugation state, and stereochemical information, as well as implicit hydrogen atoms. The SMILES sequence is parsed into a labeled sequence through a syntax tree. For example, carboxylic acid (C(=O)O) is broken down into a sequence of atoms, bonds, and brackets ['C','(','=','O',')','O'], preserving the chemical symbols and structural logic.

[0090] Specifically, molecular structures and their reactivity labels are obtained from public databases (such as PubChem and ChEMBL) or experimental data. Molecules are stored as SMILES strings, which are parsed into molecular graph structures using the RDKit toolkit. Physicochemical features of atoms and bonds (such as atom type, hybridization state, and bond type) are extracted, and corresponding SMILES syntax tree marker sequences and three-dimensional conformations are generated. For reactivity data, reaction site atoms (such as catalytic sites and redox centers) need to be labeled, and atomic-level binary labels are constructed. To enhance data diversity, data augmentation strategies such as SMILES enumeration and stereoisomer generation are employed, and negative samples are generated through operations such as random masking and bond breaking.

[0091] In this embodiment, atomic embeddings are generated through a message-passing neural network based on the defined features.

[0092] Specifically, to integrate the advantages of the three representations, an improved message-passing neural network (MPNN) is designed to generate atom embeddings: through multiple rounds of iterative message passing, each atom aggregates local chemical environment information from its neighboring nodes and edge features, and the update formula is as follows:

[0093]

[0094] In the formula, Let i be the feature vector of atom i at the l-th layer of MPNN; Let i be the feature vector of atom i in the (l-1)th layer; W is the eigenvector of neighboring atom j in layer l-1; (l) Here, h is the weight matrix of the l-th layer, used to perform linear transformations on the atomic features; N(i) is the set of neighboring nodes of atom i; ij E represents the eigenvector of the edge between atom i and its neighbor atom j; quantum (i,j) is the quantum interaction energy between atom i and atom j calculated by DFT, used to capture the electron cloud overlap effect; ReLU is the linear rectified function, used to introduce nonlinearity, with the formula ReLU(x)=max(0,x); MLP is the multilayer perceptron, used to process neighbor node information.

[0095] In this embodiment, in step S2, based on the atomic embedding, the hybrid distance feature of the fused topological shortest path and the three-dimensional Euclidean distance is calculated and obtained;

[0096] Specifically, the 3D conformation is used to generate a low-energy conformation via RDKit, and information such as 3D Euclidean distance and dihedral angle is extracted and combined with the topological shortest path to form a hybrid distance feature. The calculation formula for the hybrid distance feature is as follows:

[0097]

[0098] In the formula, It is a mixed distance feature; This is the topological shortest path; α represents the three-dimensional Euclidean distance; α is the weighting coefficient.

[0099] In this embodiment, in step S3, a chemical environment-corrected graph position code is constructed based on the mixed distance features and combined with bond type weights;

[0100] Specifically, to address the challenge of modeling deserialized molecular graph structures, a chemical environment-corrected graph position encoding (GPE) is constructed, mapping the shortest path length between atoms to position vectors:

[0101]

[0102] In the formula, GPE ij The resulting vector encodings of the graph positions of atoms i and j; It is a mixed distance feature; The chemical environment correction factor is a weighted coefficient related to the type of chemical bond and conjugation state between atoms i and j, such as w=1 for single bonds, w=0.5 for double bonds, w=0.3 for triple bonds, and w=0.6 for aromatic bonds.

[0103] In this embodiment, in step S4, based on the graph position encoding, the mixed distance features and charge difference are incorporated into the attention score calculation formula to obtain the optimized attention mechanism;

[0104] Specifically, the expression for the attention score calculation formula is as follows:

[0105]

[0106] In the formula, Attention(Q,K,V) represents the calculation result of the multi-head self-attention mechanism; Q,K,V are the query, key, and value matrices, respectively; α is the hyperparameter; q i q j d represents the charges of atoms i and j, respectively; k Let K be the dimension of the key matrix K; T is the transpose symbol.

[0107] In this embodiment, in step S5, based on the optimized attention mechanism, the reaction site prediction model is jointly trained in two tasks to obtain the trained reaction site prediction model.

[0108] In this embodiment, the reaction site prediction model is constructed based on the PyTorch framework. The molecular graph encoder adopts an MPNN architecture, the message passing function is designed as a two-layer MLP, the aggregation function is a weighted summation, and the update function uses a gated recurrent unit (GRU). The Transformer encoder has 6 layers, each containing 8 attention heads. The graph position encoding (GPE) maps the shortest path length to a 128-dimensional vector through a single-layer MLP. In the multi-task learning module, the temperature coefficient τ of the contrastive learning loss is set to 0.1, and the masking ratio for the mask prediction task is 15%. The output layer maps atom embeddings to activity probabilities through a fully connected network, and the threshold of the sigmoid function is determined by optimization using the validation set F1 score.

[0109] Specifically, this invention constructs a dual-task joint training framework for reaction mechanism awareness, and performs the dual-task joint training on the reaction site prediction model; the dual-task joint training includes: a contrastive learning task and a mask prediction task;

[0110] In contrastive learning, a targeted negative sample generator is used to perturb key functional groups, and a chemical similarity penalty is introduced through a loss function.

[0111] The loss function expression for the contrastive learning task is:

[0112]

[0113] In the formula, L contrast The loss function for contrastive learning; sim(x)+ ,x + ) represents the similarity between positive samples; τ is a temperature coefficient used to adjust the smoothness of the similarity distribution; sim(x) + ,x - s(x) represents the similarity between positive and negative samples; + ,x - The chemical similarity between positive and negative samples calculated based on RDKit fingerprints is used to suppress invalid negative samples.

[0114] The loss function expression for the mask prediction task is as follows:

[0115]

[0116] In the formula, L mask Here, BCE is the loss function for the mask prediction task; BCE is the binary cross-entropy loss function used to calculate the probability p of predicting the atom type. type Compared with the true probability The differences between them; λ is a hyperparameter used to balance the prediction loss of atom type and the prediction loss of charge transfer probability matrix; MSE is the mean squared error loss function used to calculate the predicted charge transfer probability matrix T and the true matrix. The difference between them; T is the charge transfer probability matrix; This is the true matrix.

[0117] By restoring the atom type and predicting the charge transfer probability matrix, the electron density rearrangement was simulated, enabling dual-task collaborative learning.

[0118] In this embodiment, the Adam optimizer is used to train the reaction site prediction model with an initial learning rate of 3e-4, dynamically adjusted using a cosine annealing strategy. The training data is divided into training, validation, and test sets in an 8:1:1 ratio, with a batch size of 32. To prevent overfitting, random dropout (DropEdge, probability 0.2) is added to the molecular graph edges, and Dropout (probability 0.3) is applied to the Transformer output embedding. During training, the contrastive learning and mask prediction tasks are jointly optimized using a weighted loss (weight ratio 1:0.5), with a training cycle of 100 epochs. An early stopping mechanism (patience = 10) is triggered based on the validation set loss.

[0119] In this embodiment, in step S6, the attention weights are regularized according to quantum chemical priors, and the atomic-level reaction activity probability is predicted and output through the trained reaction site prediction model; a heat map is generated based on the atomic-level reaction activity probability; and highly active reaction sites are labeled through the heat map.

[0120] Specifically, to address the disconnect between traditional methods and chemical mechanisms, this invention injects quantum chemical prior knowledge into the attention mechanism and applies chemical regularization constraints to the attention weight matrix by using a loss function that constrains the attention weights.

[0121] The loss function is:

[0122]

[0123] In the formula, L explain Let A be the loss function used to constrain the attention weights; A is the attention weight matrix learned by the model; A chem This is a theoretical interaction matrix based on chemical experience; ||·|| F The Frobenius norm is used to measure the difference between two matrices; A chem,ij Let be the theoretical interaction matrix between atom i and atom j; γ and η are trainable parameters used to adjust the calculation of the theoretical interaction matrix.

[0124] By applying regularization constraints to the attention weights, the model-generated reaction paths are aligned with classical reaction mechanisms. The final active sites are output with atomic-level probabilities.

[0125] p i =σ(W p ·h i +b p )

[0126] In the formula, p i σ represents the probability that an atom is an active site; σ is the Sigmoid function; W p b p h represents the trainable weight matrix and bias vector. i is the eigenvector of the atom.

[0127] Using RDKit, atomic activity probabilities are mapped to 3D molecular structures, generating heatmaps that can be exported as PNG or interactive HTML files. Attention heatmaps are plotted using Matplotlib, with high-weight atom pairs labeled with color-gradient lines.

[0128] The technical specifications of this invention compared to traditional GNNs are shown in Table 1:

[0129] Technical indicators Traditional GNN This invention Long-range dependency modeling weak Strong (GPE+Transformer) Training data requirements high Low (contrast-based learning enhancement) Explainability Low High (Attention Heatmap)

[0130] Table 1 Comparison of Technical Indicators

[0131] In summary, this invention represents and inputs molecular data through molecular graphs, SMILES sequences, and three-dimensional conformations; extracts and obtains defined features from the molecular data; generates atomic embeddings based on these defined features using a message-passing neural network; calculates hybrid distance features that combine topological shortest paths and three-dimensional Euclidean distances based on these atomic embeddings; constructs a chemical environment-corrected graph position code based on these hybrid distance features and bond type weights; integrates the hybrid distance features and charge difference into the attention score calculation formula based on the graph position code to obtain an optimized attention mechanism; performs dual-task joint training on the reaction site prediction model based on the optimized attention mechanism to obtain a trained reaction site prediction model; regularizes the attention weights according to quantum chemical priors and predicts the output atomic-level reaction activity probability using the trained reaction site prediction model; generates a heatmap based on the atomic-level reaction activity probability; and labels highly active reaction sites using the heatmap. This invention constructs atomic-level mixed distance features through multimodal data fusion (molecular graph topology, SMILES sequences, and 3D conformations), overcoming the limitations of single characterization and comprehensively capturing the spatial and electronic properties of chemical structures. It innovatively introduces bond type correction factors and charge difference coupling mechanisms, establishing a chemical physics-driven attention weight allocation rule within the Transformer architecture to explicitly model long-range interatomic interactions (such as the conjugation effect in aromatic systems and charge attraction of polar groups), thus addressing the insensitivity of traditional graph neural networks to electron transfer. It also designs directional negative sample contrastive learning and electron transfer matrix prediction. The dual-task framework generates mechanism-level negative samples by perturbing key functional groups and forces the model to synchronously recover the atomic type and charge density distribution, making the learning process deeply consistent with the essential laws of electron rearrangement in chemical reactions. Finally, it combines quantum chemical priors such as the Fukui function to regularize the attention heatmap, transforming black-box prediction into a mechanism-verifiable atomic-level reaction path. This invention can improve prediction accuracy and reduce training data requirements, making the prediction results highly consistent with the reaction energy barrier distribution calculated by DFT, providing an industrial-grade tool with both high accuracy and mechanism interpretability for drug active site screening and catalytic reaction design.

[0132] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0133] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0134] Example 2

[0135] See Figure 2 Embodiment 2 of the present invention also provides a reaction site prediction device based on chemical physics prior driving, comprising:

[0136] The molecular data input and processing module 001 is used to represent and input molecular data through molecular diagrams, SMILES sequences, and three-dimensional conformations; extract predetermined features from the molecular data; and generate atomic embeddings based on the predetermined features through a message-passing neural network.

[0137] The hybrid distance feature calculation module 002 is used to calculate and obtain the hybrid distance feature that fuses the topological shortest path and the three-dimensional Euclidean distance based on the atomic embedding.

[0138] Graph position coding construction module 003 is used to construct a chemical environment-corrected graph position code based on the mixed distance features and bond type weights.

[0139] Attention mechanism optimization module 004 is used to incorporate the hybrid distance features and charge difference into the attention score calculation formula based on the graph position encoding to obtain an optimized attention mechanism;

[0140] The model training module 005 is used to perform dual-task joint training on the reaction site prediction model based on the optimized attention mechanism to obtain the trained reaction site prediction model.

[0141] The reaction site prediction module 006 is used to regularize the attention weights based on quantum chemical priors and predict the atomic-level reaction activity probability through the trained reaction site prediction model; generate a heat map based on the atomic-level reaction activity probability; and label highly active reaction sites through the heat map.

[0142] In this embodiment, the set features in the molecular data input and processing module 001 include: atom type, bond type, charge and quantum interaction energy features;

[0143] The expression for the atomic embedding is:

[0144]

[0145] In the formula, Let i be the feature vector of atom i at the l-th layer of MPNN; Let i be the feature vector of atom i in the (l-1)th layer; W is the eigenvector of neighboring atom j in layer l-1; (l) Let N be the weight matrix of the l-th layer; N(i) be the set of neighbor nodes of atom i; h ij E represents the eigenvector of the edge between atom i and its neighbor atom j; quantum (i,j) represents the quantum interaction energy between atom i and atom j obtained by DFT calculation; ReLU is the linear rectified function; MLP is the multilayer perceptron.

[0146] In this embodiment, the formula for calculating the mixed distance feature in the mixed distance feature calculation module 002 is as follows:

[0147]

[0148] In the formula, It is a mixed distance feature; This is the topological shortest path; α represents the three-dimensional Euclidean distance; α is the weighting coefficient.

[0149] In this embodiment, the graph position encoding expression in the graph position encoding construction module 003 is:

[0150]

[0151] In the formula, GPE ij The resulting vector encodings of the graph positions of atoms i and j; It is a mixed distance feature; It is a chemical environment correction factor.

[0152] In this embodiment, in the attention mechanism optimization module 004, during the process of integrating the mixed distance features and the charge difference into the attention score calculation formula, the expression of the attention score calculation formula is:

[0153]

[0154] In the formula, Attention(Q,K,V) represents the calculation result of the multi-head self-attention mechanism; Q,K,V are the query, key, and value matrices, respectively; α is the hyperparameter; q i q j d represents the charges of atoms i and j, respectively; k Let K be the dimension of the key matrix K; T is the transpose symbol.

[0155] In this embodiment, in the model training module 005, during the process of performing the dual-task joint training of the reaction site prediction model based on the optimized attention mechanism, the dual-task joint training includes: a contrastive learning task and a mask prediction task.

[0156] The loss function expression for the contrastive learning task is:

[0157]

[0158] In the formula, L contrast The loss function for contrastive learning; sim(x) + ,x + ) represents the similarity between positive samples; τ represents the temperature coefficient; sim(x) + ,x - s(x) represents the similarity between positive and negative samples; + ,x - () represents the chemical similarity between positive and negative samples calculated based on RDKit fingerprints;

[0159] The loss function expression for the mask prediction task is:

[0160]

[0161] In the formula, L mask Let be the loss function for mask prediction; BCE be the binary cross-entropy loss function; λ be the hyperparameter; MSE be the mean squared error loss function; and T be the charge transfer probability matrix. This is the true matrix.

[0162] In this embodiment, the expression for regularizing the attention weights based on quantum chemical priors in the reaction site prediction module 006 is as follows:

[0163]

[0164] In the formula, L explain Let A be the loss function used to constrain the attention weights; A is the attention weight matrix learned by the model; A chem This is a theoretical interaction matrix based on chemical experience; ||·|| F For Frobenius van; A chem,ij Let be the theoretical interaction matrix between atom i and atom j; γ, η are trainable parameters.

[0165] In this embodiment, the expression for the atomic-level reaction activity probability in the reaction site prediction module 006 is:

[0166] p i =σ(Wp ·h i +b p )

[0167] In the formula, p i σ represents the probability that an atom is an active site; σ is the Sigmoid function; W p b p h represents the trainable weight matrix and bias vector. i is the eigenvector of the atom.

[0168] It should be noted that the information interaction and execution process between the modules of the above system are based on the same concept as the method embodiment in Embodiment 1 of this application, and the resulting technical effects are the same as those in the method embodiment of this application. For details, please refer to the description in the method embodiment shown above in this application, and it will not be repeated here.

[0169] Example 3

[0170] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium storing program code for a reaction site prediction method based on chemical physics priors. The program code includes instructions for executing the reaction site prediction method based on chemical physics priors as described in Embodiment 1 or any possible implementation thereof.

[0171] Computer-readable storage media can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives, SSDs).

[0172] Example 4

[0173] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0174] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor can call the program instructions to execute the reaction site prediction method based on chemical physics prior driving in Embodiment 1 or any possible implementation thereof.

[0175] Specifically, a processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.

[0176] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable system. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0177] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using program code executable by a computing system, thereby storing them in a storage system for execution by the computing system. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0178] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A reaction site prediction method based on chemical physics priors, characterized in that, include: Molecular data is represented and input using molecular diagrams, SMILES sequences, and three-dimensional conformations; The molecular data is obtained by extracting specific features; Based on the defined features, atomic embeddings are generated using a message-passing neural network. Based on the atomic embedding, the hybrid distance feature of the fused topological shortest path and three-dimensional Euclidean distance is calculated and obtained; Based on the aforementioned mixed distance features, and combined with bond type weights, a chemical environment-corrected graph position encoding is constructed. Based on the graph position encoding, the hybrid distance features and charge difference are incorporated into the attention score calculation formula to obtain an optimized attention mechanism; Based on the optimized attention mechanism, the reaction site prediction model is jointly trained in two tasks to obtain the trained reaction site prediction model. The attention weights are regularized based on quantum chemical priors, and the atomic-level reaction activity probability is predicted using the trained reaction site prediction model. A heatmap is generated based on the atomic-level reaction activity probability, and highly active reaction sites are labeled using the heatmap.

2. The reaction site prediction method based on chemical physics prior driving according to claim 1, characterized in that, The defined features include: atom type, bond type, charge, and quantum interaction energy characteristics; The expression for the atomic embedding is: In the formula, Let i be the feature vector of atom i at the l-th layer of MPNN; Let i be the feature vector of atom i in the (l-1)th layer; W is the eigenvector of neighboring atom j in layer l-1; (l) Let N be the weight matrix of the l-th layer; N(i) be the set of neighbor nodes of atom i; h ij Let be the feature vector of the edge between atom i and its neighbor atom j; E quantum (i,j) represents the quantum interaction energy between atom i and atom j obtained by DFT calculation; ReLU is the linear rectified function; MLP is the multilayer perceptron.

3. The reaction site prediction method based on chemical physics prior driving according to claim 2, characterized in that, The formula for calculating the hybrid distance feature is as follows: In the formula, It is a mixed distance feature; This is the topological shortest path; α represents the three-dimensional Euclidean distance; α is the weighting coefficient.

4. The reaction site prediction method based on chemical physics prior driving according to claim 3, characterized in that, The expression for the image position encoding is: In the formula, GPE ij The resulting vector encodings of the graph positions of atoms i and j; It is a mixed distance feature; It is a chemical environment correction factor.

5. The reaction site prediction method based on chemical physics prior driving according to claim 4, characterized in that, In the process of incorporating the mixed distance feature and the charge difference into the attention score calculation formula, the expression of the attention score calculation formula is as follows: In the formula, Attention(Q,K,V) represents the calculation result of the multi-head self-attention mechanism; Q,K,V are the query, key, and value matrices, respectively; α is the hyperparameter; q i q j d represents the charges of atoms i and j, respectively; k Let K be the dimension of the key matrix K; T is the transpose symbol.

6. The reaction site prediction method based on chemical physics prior driving according to claim 5, characterized in that, During the dual-task joint training of the reaction site prediction model based on the optimized attention mechanism, the dual-task joint training includes: a contrastive learning task and a mask prediction task. The loss function expression for the contrastive learning task is: In the formula, L contrast The loss function for contrastive learning; sim(x) + ,x + ) represents the similarity between positive samples; τ represents the temperature coefficient; sim(x) + ,x - s(x) represents the similarity between positive and negative samples; + ,x - () represents the chemical similarity between positive and negative samples calculated based on RDKit fingerprints; The loss function expression for the mask prediction task is: In the formula, L mask Let be the loss function for mask prediction; BCE be the binary cross-entropy loss function; λ be the hyperparameter; MSE be the mean squared error loss function; and T be the charge transfer probability matrix. This is the true matrix.

7. The reaction site prediction method based on chemical physics prior driving according to claim 6, characterized in that, The expression for regularizing the attention weights based on quantum chemical priors is as follows: In the formula, L explain Let A be the loss function used to constrain the attention weights; A is the attention weight matrix learned by the model; A chem This is a theoretical interaction matrix based on chemical experience; ||·|| F For Frobenius van; A chem,ij Let be the theoretical interaction matrix between atom i and atom j; γ, η are trainable parameters.

8. The reaction site prediction method based on chemical physics prior driving according to claim 7, characterized in that, The expression for the atomic-level reactivity probability is: p i =σ(W p ·h i +b p ) In the formula, p i σ represents the probability that an atom is an active site; σ is the Sigmoid function; W p b p h represents the trainable weight matrix and bias vector. i is the eigenvector of the atom.

9. A reaction site prediction device based on chemical physics prior-driven methods, employing the reaction site prediction method based on chemical physics prior-driven methods according to any one of claims 1-8, characterized in that, include: The molecular data input and processing module is used to represent and input molecular data through molecular diagrams, SMILES sequences, and three-dimensional conformations; The molecular data is obtained by extracting specific features; Based on the defined features, atomic embeddings are generated using a message-passing neural network. The hybrid distance feature calculation module is used to calculate and obtain the hybrid distance feature that combines the topological shortest path and the three-dimensional Euclidean distance based on the atomic embedding. The graph position coding construction module is used to construct a chemical environment-corrected graph position code based on the mixed distance features and bond type weights. An attention mechanism optimization module is used to incorporate the hybrid distance features and charge difference into the attention score calculation formula based on the graph position encoding to obtain an optimized attention mechanism. The model training module is used to perform dual-task joint training on the response site prediction model based on the optimized attention mechanism to obtain the trained response site prediction model. The reaction site prediction module is used to regularize the attention weights based on quantum chemical priors and predict the atomic-level reaction activity probability using the trained reaction site prediction model; generate a heatmap based on the atomic-level reaction activity probability; and label highly active reaction sites using the heatmap.

10. The reaction site prediction device based on chemical physics prior driving according to claim 9, characterized in that, In the molecular data input and processing module, the set features include: atom type, bond type, charge, and quantum interaction energy characteristics; The expression for the atomic embedding is: In the formula, Let i be the feature vector of atom i at the l-th layer of MPNN; Let i be the feature vector of atom i in the (l-1)th layer; W is the eigenvector of neighboring atom j in layer l-1; (l) Let N be the weight matrix of the l-th layer; N(i) be the set of neighbor nodes of atom i; h ij E represents the eigenvector of the edge between atom i and its neighbor atom j; quantum (i,j) represents the quantum interaction energy between atom i and atom j obtained by DFT calculation; ReLU is the linear rectified function; MLP is the multilayer perceptron.

Citation Information

Cited By

  • Molecular property prediction method based on multi-mode gating and comparative learning

    CN121075484A

  • A molecular property prediction method based on multimodal gating and contrastive learning

    CN121075484B

  • Molecular attribute prediction method and system based on nuclear power charge sorting and KAN fusion

    CN121281694A