Drug combination risk prediction method and device based on multi-source feature fusion and comparative learning, equipment and medium
Through the multi-source feature fusion and contrast learning method, the molecular graph neural network and graph convolution network are used to extract drug features, and combined with the contrast learning mechanism, the problems of risk level quantification and data imbalance in drug combination risk assessment are solved, and accurate risk prediction and feature correlation mining are achieved.
Patent Information
- Application Number
- CN202510740403.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing risk assessment methods of drug combination therapy cannot accurately quantify risk levels, and it is difficult to meet the clinical accurate demand for risk grading assessment. Moreover, the problems of unbalanced data distribution and insufficient semantic alignment of characteristics have not been effectively solved by traditional prediction methods.
Using a method based on multi-source feature fusion and contrast learning, the molecular graph features of the drug are extracted through a ternary message delivery mechanism, combined with the graph convolution network and Morgan molecular fingerprint similarity characteristics, and the contrast learning mechanism is used to mine feature associations and differences to predict the risk level of the drug combination.
It realizes accurate prediction of drug combination risk levels, effectively deals with data imbalance problem, improves the generalization ability and feature correlation mining ability of the model, and improves the accuracy of risk assessment.
Smart Images

Figure CN120260732A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug combination prediction, and in particular, to a drug combination risk prediction method, device, equipment and medium based on multi-source feature fusion and contrast learning. Background Art
[0002] As an important treatment means for modern medicine to deal with complex diseases, combination drug therapy realizes synergistic effect and reduces the single drug dose through the combination of multiple drugs. However, its application is severely restricted by the risk of drug-drug interactions. Accurately evaluating the risk level of drug combinations is crucial for clinical medication safety. However, traditional prediction methods can only judge whether there is an interaction between drugs, and cannot quantify the difference in risk levels, making it difficult to meet the precise requirements of clinical risk grading assessment. Although existing clinical guidelines have paid attention to the risk assessment of combined drug use, there is a lack of effective technical means to achieve quantitative prediction of risk levels.
[0003] In the field of deep learning, existing drug interaction prediction methods are mainly divided into chemical structure methods and graph representation learning methods. The chemical structure method is based on the molecular similarity hypothesis, and uses a deep neural network to mine the similarity between molecular structures to predict the interaction type. However, it relies on artificial feature engineering and is difficult to model complex network relationships. The graph representation learning method models drugs or associated entities as graph structures and uses graph neural networks to capture complex relationships between drugs. However, existing models are mostly limited to binary classification discrimination, unable to effectively handle risk level prediction tasks, and lacking effective solutions to problems such as data distribution imbalance and insufficient feature semantic alignment.
[0004] There are mainly three defects in the existing technology: First, the prediction paradigm is limited to the binary classification framework and cannot achieve multi-level evaluation of risk levels. Second, rare drug interaction events lead to a serious imbalance in data distribution. Traditional methods are prone to overfitting the majority class and ignoring the tail class, resulting in insufficient recall rate for rare events. Finally, existing models mostly adopt simple feature splicing or weighted average strategies, lacking in-depth mining of semantic alignment of different features and multi-source feature associations, and it is difficult to comprehensively evaluate the risk of drug combinations. These problems restrict the clinical application value of drug combination risk level prediction technology. Summary of the Invention
[0005] The present invention provides a drug combination risk prediction method, device, equipment and medium based on multi-source feature fusion and contrast learning to improve at least one of the above technical problems.
[0006] In the first aspect, the present invention provides a drug combination risk prediction method based on multi-source feature fusion and contrast learning, which includes steps S1 to S8.
[0007] S1. Obtain a drug data set.
[0008] S2. Obtain a feature map according to the drug dataset.
[0009] S3. According to the feature graph, the molecular graph neural network based on the ternary message passing mechanism is used to perform adaptive feature extraction of the molecular graph structure to obtain the molecular graph features of the drug. The molecular graph features are defined as structural features. .
[0010] S4. Input the molecular graph features generated by TrimNet into the multi-layer cascaded graph convolutional network to extract high-order topological features in the drug network and obtain relationship features. .
[0011] S5. Calculate similarity features based on the drug dataset, Morgan molecular fingerprint and Tanimoto coefficient , and extract interactive embedding features .
[0012] S6. Based on the structural features and the similarity features, potential associations and differences between the features are mined through a comparative learning mechanism to obtain structural comparative learning features. Comparative learning of features with similarity .
[0013] S7: perform feature fusion and feature alignment on the structural feature, the relationship feature, the structural contrast learning feature, the similarity contrast learning feature, and the interactive embedding feature to obtain a fused feature. .
[0014] S8. According to the fusion features, a classification model is used to determine the risk level of the drug combination and obtain the drug combination risk.
[0015] As a further solution of the present invention, according to the feature graph, adaptive feature extraction of the molecular graph structure is performed through a molecular graph neural network based on a ternary message passing mechanism to obtain the molecular graph features of the drug, specifically including: The molecular graph neural network based on the ternary message passing mechanism performs the following steps in the message passing phase: According to the feature graph, node pairs and their connecting edges are projected into a unified feature space through a triplet attention network, and the interaction strength is calculated through a nonlinear transformation. In the formula, is the interaction strength, represents the LeakyReLU activation function, U is the learnable weight vector, Represents transposition, and are two different trainable parameter matrices, and respectively represent time steps time nodes and nodes hidden states of is the node and nodes edge features between represents vector concatenation
[0016] Generate an attention distribution by normalization according to the interaction intensity . In the formula is the attention distribution is the exponential function with the natural constant e as the base is the node neighborhood set of is the node and nodes interaction intensity
[0017] Through the hidden state of the node and the edge features weighted sum to obtain the aggregated message of the node In the formula . In the formula represents the aggregated message of the node of is the attention coefficient and are two different trainable weight matrices is the node hidden state of represents the Hadamard product
[0018] According to the aggregated message, use the multi-head attention mechanism to generate multiple groups of messages in parallel and concatenate them to obtain multi-dimensional features . In the formula is the multi-dimensional feature represents concatenation is the attention coefficient calculated by the th attention head is the first weight matrix of the input linear transformation is the second weight matrix of the input linear transformation
[0019] Use the gated recurrent unit as the node update function to fuse the previously extracted message with the current integrated feature, realize the integration of temporal features, introduce layer normalization to alleviate gradient anomalies, and gradually optimize the node representation through T iterations to output high-order features . In the formula For high - order features, LN represents layer normalization, and GRU represents gated recurrent unit, is the currently integrated feature.
[0020] The molecular graph neural network based on the ternary message passing mechanism performs the following steps in the read - out stage: Use LSTM to update the hidden state. . Where, is the hidden state, represents the LSTM network, represents the aggregated feature at the previous time step.
[0021] According to the hidden state, calculate the attention weights of the nodes. . Where, is the attention weight, is the softmax activation function, is the final hidden state of node i.
[0022] According to the attention weights, aggregate the node features. . Where, is the aggregated node feature, is the total number of nodes.
[0023] After aggregating the node features, through T iterations, finally obtain the molecular graph features of the drug. . Where, is the molecular graph feature of the drug, is the hidden state at the
[0024] Define the molecular graph feature as the structural feature.
[0025] As a further solution of the present invention, input the molecular graph features generated by TrimNet into a multi - level cascaded graph convolutional network to extract high - order topological features in the drug network and obtain relationship features, specifically including: Input the molecular graph features generated by TrimNet into a two - level cascaded graph convolutional network to extract relationship features. The extraction model of the relationship features is: . Where, is the output of the first - layer graph convolutional network, is the output of the second - layer graph convolutional network, is the relationship feature, represents the residual connection, represents the first - layer graph convolutional operation, represents the second - layer graph convolutional operation, is the initial graph representation output by TrimNet, Let EdgeIndexSet be the edge index set, and AttnWeight represent the attention mechanism.
[0026] As a further solution of the present invention, the steps of single-layer graph convolution are as follows: Use the molecular graph features of the drug as the initial features for graph convolution. . In the formula, is the feature after convolution, is the graph convolution, is the convolutional layer, is the initial feature, is the edge index set, is the degree matrix, is the adjacency matrix with self-loops, is the trainable parameter matrix.
[0027] Process the feature after convolution through batch normalization, ReLU activation, and residual connection. . In the formula, is the output feature, is dropout, is the ReLU activation function, is batch normalization, is the dimensionality projection function.
[0028] Calculate the node attention weights through a shallow perception network based on the processed features. . In the formula, is the node attention weight, is the tanh activation function, is the attention parameter matrix.
[0029] As a further solution of the present invention, based on the drug dataset, calculate the similarity features based on the Morgan molecular fingerprint and the Tanimoto coefficient, and extract the interaction embedding features, specifically including: Obtain the Morgan molecular fingerprints of the drug combinations according to the drug dataset.
[0030] Calculate the similarity between two drug molecular fingerprints based on the Morgan molecular fingerprints of the drug combinations to obtain the similarity features. . In the formula, is the similarity, and respectively represent the fingerprint vectors of two drug molecules.
[0031] Introduce the identity matrix as the structured prior information, and map the identity matrix to the target embedding space through the linear projection layer to generate a set of orthogonal basis vectors. . In the formula, is the orthogonal basis vector, is the identity matrix, represents a linear projection layer, represents a real number, is the total number of nodes, is the embedding space dimension of the projection.
[0032] Project the similarity feature into the target embedding space, and then perform weighted mixing with the orthogonal basis vectors to obtain a mixed embedding matrix. . Wherein, is the new similarity feature projected into the embedding space, is the mixed embedding matrix, is a learnable weight parameter.
[0033] For the drug combination and , extract the mixed embedding vectors corresponding to the drug combination from the mixed embedding matrix and . Wherein, and respectively represent the mixed embedding vectors of drugs and drug , and are the mixed embedding matrices of the two drug sets respectively.
[0034] According to the mixed embedding vectors and , capture the synergistic effect of the drug pair through the Hadamard product, and apply the L2 normalization operation to obtain the interaction embedding feature. . Wherein, is the interaction embedding feature, represents the L2 normalization operation, represents the Hadamard product, represents the L2 norm.
[0035] As a further solution of the present invention, according to the structural feature and the similarity feature, through a contrastive learning mechanism, potential associations and differences between the features are mined to obtain a structural contrastive learning feature and a similarity contrastive learning feature, specifically including: Given the original model parameters of the heterogeneous graph transformer, calculate the gradient sign direction of its corresponding loss function. The calculation model of the gradient sign direction is: . Wherein, is the gradient sign direction, is the sign function, represents the gradient direction, represents the gradient, are the original model parameters, is the corresponding loss function.
[0036] Introduce a loss weight factor to dynamically adjust the noise intensity according to the current loss value. . Wherein, is the loss weight factor, is the current loss value, is the smoothing factor.
[0037] Perturb the parameters of the GNN layer of the model according to the loss weight factor, and the perturbed parameters are: . Wherein, represents the perturbed GNN layer parameters, is a learnable weight parameter used to balance the contributions of the gradient direction and Gaussian noise, is the gradient sign direction, is a random noise obeying the Gaussian distribution.
[0038] Use the heterogeneous graph transformer and its perturbed version as double encoders for encoding to obtain the original view and the perturbed view. . Wherein, and respectively represent the original views generated by inputting the structural features and similarity features into the heterogeneous graph transformer, and respectively represent the perturbed views generated by the perturbed heterogeneous graph transformer, ( ) is the original encoder, ( ) is the encoder after applying the mixed perturbation, is the graph representation.
[0039] Map the representations of the original view and the perturbed view to the latent space through a non-linear projection head. . Wherein, and are respectively the latent feature representations obtained by non-linearly transforming the structural features and similarity features of the original view through the MLP, and are respectively the latent feature representations obtained by non-linearly transforming the structural features and similarity features of the perturbed view after perturbation through the MLP, represents the non-linear projection head, is the non-linear projection head corresponding to the original encoder, is the non-linear projection head corresponding to the perturbed encoder, is the multi-layer perceptron, and are respectively two different learnable parameter matrices.
[0040] Obtain the structural contrast learning features and similarity contrast learning features according to the latent feature representations mapped to the latent space. In the formula, is the structural contrast learning feature, is the similarity contrast learning feature, is the average pooling.
[0041] As a further solution of the present invention, feature fusion and feature alignment are performed on the structural feature, the relational feature, the structural contrast learning feature, the similarity contrast learning feature, and the interaction embedding feature to obtain a fused feature, which specifically includes: Concatenate the structural feature and the relational feature to obtain a new drug structural feature . . In the formula, is the drug structural feature, is the concatenated feature, represents vector concatenation, is the feature after residual projection, and are two different learnable parameters respectively, represents the ReLU activation function, is the structural feature, is the relational feature, is the non-linear bias vector.
[0042] Fuse the structural contrast learning feature and the similarity contrast learning feature in a dual-path fusion manner to obtain a contrast learning feature .
[0043] Align the structural feature , the contrast learning feature and the previous interaction embedding feature to obtain an aligned feature. . In the formula, represents the aligned feature, FFN is a two-layer feed-forward network, is the normalization operation, is the attention weight matrix, is the Softmax activation function, represents transpose, is the query, is the key, is the value, is the hidden layer dimension, are respectively 's projection matrix.
[0044] According to the aligned feature, introduce a learnable affine transformation to dynamically adjust the fused feature to obtain the finally output fused feature . In the formula, and They are different trainable recalibration parameters respectively.
[0045] As a further embodiment of the present invention, the drug dataset comprises SMILES sequences of drug combinations.
[0046] In the second aspect, the present invention provides a drug combination risk prediction device based on multi-source feature fusion and comparative learning, which includes a data set acquisition module, a feature graph construction module, a structural feature extraction module, a relationship feature extraction module, an interactive embedding feature extraction module, a comparative learning module, a feature fusion module and a risk prediction module.
[0047] The data set acquisition module is used to acquire the drug data set.
[0048] The feature map construction module is used to obtain a feature map according to the drug data set.
[0049] The structural feature extraction module is used to extract the molecular graph structure adaptively based on the feature graph through a molecular graph neural network based on a ternary message passing mechanism to obtain the molecular graph features of the drug. The molecular graph features are defined as structural features. .
[0050] The relational feature extraction module is used to input the molecular graph features generated by TrimNet into the multi-layer cascaded graph convolutional network to extract high-order topological features in the drug network and obtain relational features. .
[0051] An interactive embedding feature extraction module for calculating similarity features based on the drug dataset based on Morgan molecular fingerprints and Tanimoto coefficients , and extract interactive embedding features .
[0052] The contrastive learning module is used to mine the potential associations and differences between the features through the contrastive learning mechanism according to the structural features and the similarity features, and obtain the structural contrastive learning features. Comparative learning of features with similarity .
[0053] A feature fusion module is used to fuse and align the structural features, the relationship features, the structural contrast learning features, the similarity contrast learning features, and the interactive embedding features to obtain a fused feature. .
[0054] The risk prediction module is used to determine the risk level of the drug combination based on the fusion features and use a classification model to obtain the drug combination risk.
[0055] In a third aspect, the present invention provides a drug combination risk prediction device based on multi-source feature fusion and contrast learning, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a drug combination risk prediction method as described in any paragraph of the first aspect.
[0056] In a fourth aspect, the present invention provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute a drug combination risk prediction method as described in any paragraph of the first aspect.
[0057] By adopting the above technical solutions, the present invention can achieve the following technical effects: The drug combination risk prediction method based on multi-source feature fusion and contrast learning in the embodiments of the present invention can accurately predict the risk level, effectively handle data imbalance, and deeply mine feature associations. And it has good generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for the specific embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some specific embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0059] Figure 1 is a flowchart of the drug combination risk prediction method.
[0060] Figure 2 is a logical block diagram of the drug combination risk prediction method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention.
[0062] Example 1. Please refer to Figures 1 to 2 , the first embodiment of the present invention provides a drug combination risk prediction method based on multi-source feature fusion and contrast learning. It can be executed by a drug combination risk prediction device based on multi-source feature fusion and contrast learning (hereinafter referred to as: drug combination risk prediction device). Specifically, it is executed by one or more processors in the drug combination risk prediction device to implement steps S1 to S8.
[0063] S1. Obtain a drug dataset. Among them, the drug dataset contains the SMILES sequences of drug combinations. The SMILES sequence converts the atomic and chemical bond information in the molecule into a string through specific characters and rules.
[0064] S2. Obtain a feature map according to the drug dataset.
[0065] Specifically, use the RDKit tool for the SMILES sequence and preprocess it into a feature map with node features, edge features, and an adjacency matrix. RDKit is an open-source toolkit for chemoinformatics. Based on the 2D and 3D molecular operations of compounds, it uses machine learning methods for compound descriptor generation, Morgan molecular fingerprint generation, compound structure similarity calculation, 2D and 3D molecular display, etc.
[0066] S3. According to the feature map, perform adaptive feature extraction on the molecular graph structure through a molecular graph neural network based on a ternary message passing mechanism to obtain the molecular graph features of the drug, that is, the structural features. .
[0067] In this embodiment, a molecular graph neural network based on a ternary message passing mechanism (abbreviation: TrimNet) is used for the refined characterization learning of drug molecules, capturing atomic-level features that play a decisive role in risk level prediction for structure and relationship extraction.
[0068] Specifically, input the processed graph structure of the drug combination into the TrimNet model. The TrimNet model is constructed based on the ternary message passing mechanism and realizes the adaptive feature extraction of the molecular graph structure through the interaction modeling of atom-chemical bond-atom triples. In TrimNet, the atoms of a single drug are used as nodes, and the chemical bonds between atoms are used as edges. Perform characterization learning on the atoms and chemical bonds of a single drug, and finally obtain the molecular graph features of different drugs, that is, the structural features. .
[0069] As a variant of the message passing neural network, in the message passing stage, TrimNet dynamically weights the contribution degrees of adjacent atoms and chemical bonds through a multi-head attention mechanism. In the readout stage, a serialization aggregation strategy based on the long short-term memory network (LSTM) is adopted. This design enables the model to directly learn the feature expressions of key atomic groups from the original molecular graph and provide highly discriminative structural characterizations for downstream risk level prediction.
[0070] Based on the above embodiment, in an optional embodiment of the present invention, step S3 specifically includes steps S31 to S39.
[0071] S31. According to the feature map, project the node pair and its connecting edge into a unified feature space through a triple attention network, and calculate the interaction intensity through a non-linear transformation.
[0072] .
[0073] In the formula, is the node and the node interaction intensity, represents the LeakyReLU activation function, U is a learnable weight vector, represents the transpose, and are two different trainable parameter matrices, and respectively represent the hidden states of the node at time step and the node , is the edge feature between the node and the node , represents vector concatenation. The LeakyReLU activation function is a non-linear activation function that allows negative values to pass through with a small slope.
[0074] Specifically, in the message passing stage, TrimNet generates the attention weights between nodes by fusing the interaction features of atoms and bonds. As shown in the above calculation model of the interaction intensity, at time step , given the node feature and the edge feature , the triple attention network first projects the node pair and its connecting edge into a unified feature space, and calculates the interaction intensity through a non-linear transformation.
[0075] S32. According to the interaction intensity, generate an attention distribution through normalization.
[0076] .
[0077] In the formula, is the attention distribution, is the exponential function with the natural constant e as the base, is the interaction intensity between the node and the node , is the neighborhood set of the node , is the interaction intensity between the node and the node .
[0078] S33. Aggregate the messages of nodes through the hidden states of nodes and the features of edges by weighted summation. The aggregated message of the node is obtained.
[0079] .
[0080] In the formula, represents the aggregated message of node , is the attention coefficient, and are two different trainable weight matrices, is the hidden state of node , represents the Hadamard product, is the node and node The edge feature between.
[0081] S34. According to the aggregated message, use the multi-head attention mechanism to generate multiple groups of messages in parallel and splice them to obtain multi-dimensional features, enhancing the robustness of the model.
[0082] .
[0083] In the formula, is the multi-dimensional feature, represents concatenation, is the attention coefficient calculated by the th attention head, is the number of attention heads, is the first weight matrix of the input linear transformation, is the second weight matrix of the input linear transformation, is the node 's hidden state, represents the Hadamard product, is the node and node The edge feature between, is the node 's neighborhood set.
[0084] S35. Use the gated recurrent unit as the node update function to fuse the previously extracted messages with the current integrated features, realize the integration of temporal features, introduce layer normalization to alleviate gradient anomalies, and gradually optimize the node representation through T iterations to output high-order features.
[0085] .
[0086] In the formula, where HF stands for high - order feature, LN stands for layer normalization, and GRU stands for gated recurrent unit, CF is the currently integrated feature, and MF is the multi - dimensional feature.
[0087] Specifically, in the node update operation, the gated recurrent unit (GRU) is used as the node update function to fuse the previously extracted message with the current integrated feature, realizing the integration of temporal features.
[0088] In the read - out stage of TrimNet, based on the finally updated node features the Set2Set network is used as the read - out function to dynamically aggregate the global node features to generate the graph - level embedding representation. Specifically, Set2Set aggregates the node features according to different attention weights and concatenates the aggregated features with the previous messages.
[0089] S36. Update the hidden state using LSTM.
[0090] .
[0091] In the formula, h is the hidden state, LSTM represents the LSTM network, CFt - 1 represents the aggregated feature at the previous time step.
[0092] S37. Calculate the attention weight of the node according to the hidden state.
[0093] .
[0094] In the formula, αi is the attention weight, softmax is the softmax activation function, ht is the final hidden state of node i, h is the hidden state.
[0095] S38. Aggregate the node features according to the attention weight.
[0096] .
[0097] In the formula, CF is the aggregated node feature, N is the total number of nodes, αi is the attention weight, ht is the final hidden state of node i.
[0098] S39. After aggregating the node features, through T iterations, finally obtain the molecular graph feature of the drug, that is, the structural feature .
[0099] 。
[0100] Wherein, is the molecular graph feature of the drug, is the th hidden state of the iteration, is the aggregated node feature.
[0101] S4. Input the molecular graph feature generated by TrimNet into a multi-level cascaded graph convolutional network to extract high-order topological features in the drug network and obtain relationship features . Specifically, based on the molecular graph feature generated by TrimNet, input it into a multi-level cascaded graph convolutional network as the initial feature of the drug nodes in the network to further extract high-order topological features in the drug network.
[0102] In the graph convolutional network, the graph representation is defined as , where A represents the adjacency matrix of the drug nodes, X represents the set of molecular graph features generated by TrimNet as the initial node feature matrix, V represents the set of nodes in the network, and E represents the set of edges between the nodes. Through the multi-level cascaded graph convolutional network combined with the residual structure and the attention mechanism, while preventing overfitting, the feature expression of key atomic groups is enhanced.
[0103] In the embodiment of the present invention, two-level cascaded graph convolutional networks are used for extraction. The overall feature extraction process is as shown in part a of Figure 2 .
[0104] The extraction model of the relationship feature is: .
[0105] Wherein, is the output of the first-layer graph convolutional network, is the output of the second-layer graph convolutional network, is the relationship feature, represents the residual connection, represents the first-layer graph convolutional operation, represents the second-layer graph convolutional operation, is the initial graph representation output by TrimNet, AttnWeight represents the attention mechanism, is the set of edge indices, which is dynamically constructed through preprocessed edge indices (including multiple types of edges).
[0106] Specifically, for the input node feature matrix and the set of edge indices , the drug representation node, the risk level of the drug combination represents the edge, and the molecular graph features generated by TrimNet previously are used as the initial features of its nodes Perform graph convolution. Among them, the steps of single-layer graph convolution include step S41 to step S43.
[0107] S41. Use the molecular graph features of the drug as the initial features and perform graph convolution.
[0108] .
[0109] In the formula, is the feature after convolution, is the graph convolution, is the convolutional layer, is the initial feature, is the edge index set, is the degree matrix, is the adjacency matrix with self-loops, is the trainable parameter matrix.
[0110] S42. Process the feature after convolution through batch normalization, ReLU activation and residual connection to alleviate the problem of deep network degradation.
[0111] .
[0112] In the formula, is the processed feature, is dropout, is the ReLU activation function, is batch normalization, is the dimensional projection function, is the feature after convolution, is the initial feature.
[0113] S43. According to the processed feature, calculate the node attention weight through a shallow perception network to enhance the contribution of key features.
[0114] .
[0115] In the formula, is the node attention weight, is the tanh activation function, is the attention parameter matrix, is the softmax activation function, is the processed feature.
[0116] S5. According to the drug dataset, based on Morgan molecular fingerprints and Tanimoto coefficients, obtain interaction embedding features .
[0117] In the field of drug research, drugs with similar chemical structures usually exhibit similar activities. The underlying mechanism of this phenomenon is that similar chemical structures prompt drugs to bind to the same or similar biological targets in the body, thereby triggering similar physiological responses. When two or more drugs with similar structures are used simultaneously, they are very likely to compete for the same metabolic enzymes or transporters. In view of this, in a model for predicting drug interactions, by deeply analyzing a large amount of data on drug interactions of similar drugs, the model can learn the risk patterns presented by different drug combinations, so as to more accurately evaluate the risk level of new drug combinations.
[0118] Based on the above embodiments, in an optional embodiment of the present invention, the drug similarity calculation method is based on Morgan molecular fingerprints and Tanimoto coefficients. Then step S5 specifically includes steps S51 to S56.
[0119] S51. Obtain the Morgan molecular fingerprints of the drug combination according to the drug data set.
[0120] Specifically, convert the SMILES sequences of the drug combination into Morgan molecular fingerprints respectively using the RDKit tool. The Morgan molecular fingerprint is a fixed-length bit vector (composed of 0s and 1s), and each position corresponds to the presence or absence of a specific chemical structure feature.
[0121] S52. Calculate the similarity between two drug molecular fingerprints according to the Morgan molecular fingerprints of the drug combination, and obtain the similarity feature .
[0122] .
[0123] In the formula, is the similarity, and respectively represent the fingerprint vectors of two drug molecules. represents the number of bits that are 1 in both vectors, represents the number of bits that are at least 1 in both vectors.
[0124] Specifically, use Tanimoto to calculate the similarity between two molecular fingerprints. The value range of this similarity is between 0 and 1. The closer the value is to 1, the higher the similarity between the two molecules. By calculating the similarity of the molecular fingerprints of all drug molecules in pairs, a similarity feature is finally obtained. The rows and columns of this matrix correspond to different drug molecules respectively, and each element in the matrix represents the similarity score between the corresponding two molecules.
[0125] S53. Introduce the identity matrix as structured prior information, map the identity matrix to the target embedding space via a linear projection layer, and generate a set of orthogonal basis vectors.
[0126] After obtaining the similarity features of the drugs, to further improve the robustness of the model, the embodiments of the present invention introduce the identity matrix as structured prior information. Learning with the identity matrix can endow the model with a kind of prior knowledge, enabling the model to learn the potential semantics and relationships between drugs.
[0127] Specifically, each row of the identity matrix can be regarded as a one-hot encoded vector, representing an "idealized" independent representation of a drug in the chemical space. By mapping the identity matrix to the target embedding space via a linear projection layer, the model can learn the independent representation of each drug. This operation generates a set of orthogonal basis vectors, and each basis vector corresponds to the independent representation of a drug. This independent representation can be combined with the input drug similarity features, thereby enhancing the model's ability to capture drug features.
[0128] .
[0129] In the formula, is the orthogonal basis vector, is the identity matrix, represents the linear projection layer, represents a real number, is the total number of nodes (i.e., the number of drugs), is the dimension of the projected embedding space.
[0130] S54. Project the similarity features to the target embedding space, and then perform weighted mixing with the orthogonal basis vectors to obtain a mixed embedding matrix.
[0131] .
[0132] In the formula, is the new similarity feature projected into the embedding space, is the mixed embedding matrix, is a learnable weight parameter, is the similarity feature, represents the linear projection layer, is the orthogonal basis vector.
[0133] In this embodiment, project the input similarity feature to the same embedding space, and perform weighted mixing on the identity matrix basis vectors and the similarity features with the learnable weight parameter . By dynamically adjusting the contributions of the prior basis vectors and the similarity features, the generalization ability of the model is effectively enhanced. Obtained by adaptively weighted mixing of orthogonal basis vectors and similarity features after that.
[0134] S55. For the drug combination and , retrieve the mixed embedding vectors corresponding to the drug combination from the mixed embedding matrix and .
[0135] .
[0136] In the formula, and respectively represent the mixed embedding vectors of drug and drug , and are the mixed embedding matrices of the two drug sets respectively.
[0137] S56. According to the mixed embedding vectors and , capture the synergistic effect of the drug pair through the Hadamard product, and apply the L2 normalization operation to alleviate the influence of the feature scale difference on the downstream task, and obtain the interaction embedding feature .
[0138] .
[0139] In the formula, is the interaction embedding feature, represents the L2 normalization operation, represents the Hadamard product, represents the L2 norm. and respectively represent the mixed embedding vectors of drug and drug . The L2 norm calculates the square root of the sum of the squares of the vector elements and is used to measure the magnitude of the vector. is used to represent the interaction embedding feature of drug and drug .
[0140] The embodiment of the present invention realizes robust representation learning of drug interaction embedding by fusing the prior structure and the similarity features of drugs. The hybrid weighting mechanism effectively improves the model's ability to capture sparse interaction patterns while retaining the topological properties of the chemical space.
[0141] The overall drug interaction embedding extraction is as shown in part b of Figure 2 . Its model is: .
[0142] In the formula, is the finally obtained interactive embedding feature, represents the interactive embedding function, and represents the drug and the drug , and respectively represent the similarity features of the drug and the drug , is dropout, represents the L2 normalization operation, is the mixed embedding matrix of the drug , is the mixed embedding matrix of the drug , is the mixed embedding matrix of the drug or the drug , is the node attention weight, is the identity matrix, represents the linear projection layer of the drug or the drug , is the similarity feature of the drug or the drug , is a replacement symbol, replaced by or .
[0143] is a learnable weight parameter for dynamically adjusting the contributions of the prior basis vector and the similarity feature.
[0144] S6. According to the structural features and the similarity features, through a contrast learning mechanism, potential associations and differences between the features are mined to obtain the structure contrast learning feature and the similarity contrast learning feature .
[0145] After the drug interactive embedding extraction is completed, to further enhance the discriminative ability of the drug molecular representation, the embodiment of the present invention introduces a contrast learning mechanism. As a key self-supervised learning method, in traditional graph contrast learning, the way of pulling positive sample pairs closer and pushing negative sample pairs farther in the feature space is usually adopted to learn more effective feature expressions.
[0146] However, the initial features of the method nodes are highly dependent. When the discriminability of the initial features of the nodes is insufficient or there are missing cases, it is difficult to effectively mine key features, which has a negative impact on the quality of molecular representation. In addition, when dealing with heterogeneous graphs, due to the complex and diverse semantics of their nodes and edges, traditional methods are difficult to effectively integrate information.
[0147] To solve the above problems, the embodiment of the present invention adopts the graph contrast learning method without enhancement proposed by MRGCDDI. This method uses the SimGRACE framework and uses the Heterogeneous Graph Transformer (HGT) and its perturbed version as double encoders to extract two related views for comparison.
[0148] Based on the above embodiments, in an optional embodiment of the present invention, a perturbation method is further designed, and an Adaptive Gradient-Noise Hybrid Perturbation (AGNHP for short) is proposed. AGNHP can, while maintaining the semantic integrity of the original graph data, guide the perturbation direction through the gradient direction and dynamically adjust the noise weight according to the current training loss, thereby improving the effectiveness and robustness of contrast learning. AGNHP includes steps S61 to S66.
[0149] S61. Given the original model parameters of the heterogeneous graph transformer, calculate the gradient sign direction of its corresponding loss function. The calculation model of the gradient sign direction is: .
[0150] In the formula, is the gradient sign direction, is the sign function, represents the gradient direction, represents the gradient, is the original model parameter, is the corresponding loss function.
[0151] The gradient direction is ternarized to , effectively avoiding the influence brought by the gradient magnitude difference. To improve the perturbation adaptability, stronger perturbations are applied in the initial stage of training (when the loss is large) to explore the feature space, and the perturbations are weakened in the later stage (when the loss is small) to ensure stable convergence.
[0152] S62. Introduce a loss weight factor to dynamically adjust the noise intensity according to the current loss value.
[0153] .
[0154] In the formula, is the loss weight factor, is the current loss value, is the smoothing factor.
[0155] S63. Perturb the parameters of the GNN layer of the model according to the loss weight factor, and the perturbed parameters are: .
[0156] In the formula, represents the perturbed GNN layer parameters, is the original model parameter, is the loss weight factor, is a learnable weight parameter used to balance the contributions of the gradient direction and Gaussian noise, is the gradient sign direction, is the random noise obeying the Gaussian distribution.
[0157] S64. Use the heterogeneous graph transformer and its perturbed version as dual encoders for encoding to obtain the original view and the perturbed view.
[0158] .
[0159] In the formula, and respectively represent the original views generated by inputting the structural features and similarity features into the heterogeneous graph transformer, and respectively represent the perturbed views generated by the perturbed heterogeneous graph transformer, ( ) is the original encoder, ( ) is the encoder after applying the hybrid perturbation, is the graph representation, represents the perturbed GNN layer parameters, is the original model parameter.
[0160] S65. Map the representations of the original view and the perturbed view to the latent space through the non-linear projection head.
[0161] .
[0162] In the formula, and are respectively the structural features and similarity features of the original view, and the latent feature representations obtained through the non-linear transformation of the MLP (set as positive samples), and They are the structural features and similarity features of the perturbed view after perturbation, the potential feature representation obtained by the nonlinear transformation of MLP (set as negative samples), represents the nonlinear projection head (i.e., the MLP after two layers of nonlinear transformation), is the nonlinear projection head corresponding to the original encoder, is the nonlinear projection head corresponding to the perturbed encoder, For multi-layer perceptron, is the ReLU activation function, and They are two different learnable parameter matrices.
[0163] S66. According to the latent feature representation mapped to the latent space, obtain structure contrast learning features and similarity contrast learning features.
[0164] .
[0165] In the formula, Learning features for structural comparison, Learn features for similarity comparison, is average pooling.
[0166] S7. Perform feature fusion and feature alignment on the structural features, the relationship features, the structural contrast learning features, the similarity contrast learning features, and the interactive embedding features to obtain fused features.
[0167] In the construction of drug research models, how to effectively fuse features extracted from different modules is a key issue to improve model performance. To solve this problem, the embodiment of the present invention proposes a hierarchical feature fusion module (HFFM), which adopts a progressive fusion strategy and combines residual connection and attention mechanism to achieve efficient fusion of structural and relational features, interactive embedding features and contrastive learning features, so that the model can more comprehensively and accurately capture the feature information related to the risk level of drug combinations.
[0168] Based on the above embodiment, in an optional embodiment of the present invention, step S7 specifically includes steps S71 to S74.
[0169] S71. Combine structural features with relationship features to obtain new drug structural features .
[0170] .
[0171] In the formula, For the structural characteristics of drugs, For the concatenated features, indicating vector concatenation, For the features after residual projection, and are two different learnable parameters respectively, indicating the ReLU activation function, For the structural features, For the relational features, is the non - linear bias vector.
[0172] In the fusion of structural and relational features, the TrimNet structural features serve as the basic structural representation of drug molecules, containing rich atomic and chemical bond information. The deep relational features extracted by the graph convolutional network, on the other hand, focus on revealing the interaction relationships between drug molecules. Concatenating the structural features and relational features can provide a more comprehensive information basis for subsequent fusion operations, through a dual - path fusion with residual connections.
[0173] This fusion method can not only fully integrate the advantages of the two types of features, but also retain the key information of the original features through residual connections, effectively avoiding the problem of information loss that may occur during the fusion process, thus obtaining new drug structural features and providing a more representative input for subsequent feature processing.
[0174] S72: Adopt the method of dual - path fusion to fuse the structural contrast learning features and the similarity contrast learning features to obtain the contrast learning features.
[0175] Specifically, input the structural features of TrimNet and the similarity features into the contrast learning. Through the unique mechanism of contrast learning, the potential correlations and differences between the features can be mined, and then the structural contrast learning features and the similarity contrast learning features are obtained. Then, adopt the same dual - path fusion method as in step S71 to make these two types of contrast learning features complement each other, further enhancing the expression ability of the features and obtaining the contrast learning features This fusion method can make full use of the features learned by contrast learning, improve the discriminative ability of the model for drug molecule features, and provide a more discriminative feature representation for subsequent attention fusion.
[0176] S73: Align the structural features , the contrast learning features and the previous interaction embedding features to obtain the aligned features.
[0177] .
[0178] In the formula, represents the alignment feature, FFN is a two-layer feed-forward network, is the normalization operation, is the attention weight matrix, is the Softmax activation function, represents the transpose, is the query, is the key, is the value, is the hidden layer dimension, are respectively the projection matrices of
[0179] To achieve semantic alignment and associative fusion between multi-source features, the embodiments of the present invention design an improved multi-head attention mechanism for attention fusion, and enhance the ability of collaborative representation through multi-source feature interaction modeling. This mechanism performs feature alignment on the structural feature , the contrast learning feature and the previous interactive embedding feature . In this way, features from different sources can be aligned at the semantic level.
[0180] S74. Introduce a learnable affine transformation to dynamically adjust the fused features and obtain the finally output fused features .
[0181] .
[0182] In the formula, and are respectively different trainable recalibration parameters.
[0183] Specifically, introducing a learnable affine transformation to dynamically adjust the fused features can further enhance the expressive ability of the fused features. Retain the original feature information through residual connection, and use attention to achieve semantic alignment between features. The finally output fused features can effectively capture the associated features of the drug combination risk level, providing a more accurate feature representation for subsequent risk level prediction, thereby improving the performance of the model in the drug combination risk assessment task.
[0184] S8. According to the fused features, use a classification model to judge the drug combination risk level and obtain the drug combination risk.
[0185] After completing the hierarchical feature fusion, the model needs to convert the fused features into a prediction of the drug combination risk level. The embodiments of the present invention implement this key task through a multi-layer perceptron MLP. Given the fused features of the drug pair ( and First, feature concatenation is performed to form a new vector containing drug combination information. Subsequently, the concatenated feature vector is input into a three-layer fully connected network for prediction. Each layer in the fully connected network performs a linear transformation on the input through a weight matrix and applies an activation function to introduce non-linearity, thereby gradually extracting features and finally outputting the prediction result of the risk level of the drug combination.
[0186] 。
[0187] In the formula, is and the concatenated features, is the multi-layer perceptron, is the Softmax activation function, is the prediction result.
[0188] Through the above series of model optimization and training strategies, the embodiments of the present invention construct an efficient, accurate and well-generalized drug combination risk level prediction model, which can effectively handle complex situations in practical applications and provide reliable support for drug R & D and clinical medication.
[0189] The drug combination risk prediction method based on multi-source feature fusion and contrast learning in the embodiments of the present invention can accurately predict the risk level, effectively handle data imbalance, and deeply mine feature associations. And it has good generalization ability.
[0190] Accurately predict the risk level: Through the multi-source feature fusion architecture, multiple features are integrated to construct a multi-level feature representation space, improving the ability to predict the risk level. Experiments on multiple datasets show that MSFCL is significantly better than the baseline model in various evaluation metrics. For example, in the risk level prediction task of the DDInter dataset, the Accuracy is increased by an average of 9.84%, and the Macro-F1 is increased by an average of 14.97%.
[0191] Effectively handle data imbalance: The AGNHP strategy in the adaptive contrast learning mechanism enhances feature discriminability while retaining data semantics, improving the model's ability to distinguish minority class features and alleviating the problem of data distribution imbalance.
[0192] Deeply mine feature associations: The hierarchical feature fusion module uses residual connections and attention mechanisms to achieve progressive fusion of multi-source features, effectively mining the association rules of the risk level, enabling the model to more comprehensively and accurately capture the feature information related to the drug combination risk level.
[0193] Good generalization ability: In the multi-classification tasks of DrugBank and MDF-SA-DDI datasets, MSFCL also demonstrated excellent generalization ability, proving that this method can stably and accurately predict the risk level of drug combinations on different datasets.
[0194] Embodiment 2: The present invention provides a drug combination risk prediction device based on multi-source feature fusion and comparative learning, which includes a data set acquisition module, a feature graph construction module, a structural feature extraction module, a relationship feature extraction module, an interactive embedding feature extraction module, a comparative learning module, a feature fusion module and a risk prediction module.
[0195] The data set acquisition module is used to acquire the drug data set.
[0196] The feature map construction module is used to obtain a feature map according to the drug data set.
[0197] The structural feature extraction module is used to extract the molecular graph structure adaptively based on the feature graph through a molecular graph neural network based on a ternary message passing mechanism to obtain the molecular graph features of the drug. The molecular graph features are defined as structural features. .
[0198] The relational feature extraction module is used to input the molecular graph features generated by TrimNet into the multi-layer cascaded graph convolutional network to extract high-order topological features in the drug network and obtain relational features. .
[0199] An interactive embedding feature extraction module for calculating similarity features based on the drug dataset based on Morgan molecular fingerprints and Tanimoto coefficients , and extract interactive embedding features .
[0200] The contrastive learning module is used to mine the potential associations and differences between the features through the contrastive learning mechanism according to the structural features and the similarity features, and obtain the structural contrastive learning features. Comparative learning features with similarity .
[0201] A feature fusion module is used to fuse and align the structural features, the relationship features, the structural contrast learning features, the similarity contrast learning features, and the interactive embedding features to obtain a fused feature. .
[0202] The risk prediction module is used to determine the risk level of the drug combination based on the fusion features and use a classification model to obtain the drug combination risk.
[0203] Embodiment 3. The present invention provides a drug combination risk prediction device based on multi-source feature fusion and contrast learning, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a drug combination risk prediction method as described in any paragraph of Embodiment 1.
[0204] Embodiment 4. The present invention provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute a drug combination risk prediction method as described in any paragraph of Embodiment 1.
[0205] Obviously, the embodiments described above are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0206] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0207] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0208] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs. It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0209] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0210] It should be understood that the term "and / or" used herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the situation where A exists alone, the situation where A and B exist simultaneously, and the situation where B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0211] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0212] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0213] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A drug combination risk prediction method based on multi-source feature fusion and contrast learning, characterized in that, Include: Get the drug dataset; According to the drug data set, a feature map is obtained; According to the feature graph, adaptive feature extraction of the molecular graph structure is performed through a molecular graph neural network based on a ternary message passing mechanism to obtain the molecular graph features of the drug; Define the molecular graph feature as a structural feature ; Input the molecular graph features generated by TrimNet into a multi-level cascaded graph convolutional network to extract high-order topological features in the drug network and obtain relational features ; Calculate similarity features based on the Morgan molecular fingerprint and Tanimoto coefficient according to the drug data set and extract interactive embedding features ; Based on the described structural features and the similarity features, through a contrastive learning mechanism, potential associations and differences between the features are mined to obtain structural contrastive learning features and similarity contrastive learning features ; Perform feature fusion and feature alignment on the structural features, the relational features, the structural contrast learning features, the similarity contrast learning features, and the interactive embedding features to obtain fused features ; According to the fusion features, a classification model is used to judge the risk level of the drug combination to obtain the drug combination risk.
2. The method for predicting the risk of a drug combination based on multi-source feature fusion and contrastive learning according to claim 1, wherein According to the feature graph, the molecular graph structure is adaptively extracted through a molecular graph neural network based on a ternary message passing mechanism to obtain the molecular graph features of the drug, specifically including: The molecular graph neural network based on the ternary message passing mechanism performs the following steps in the message passing phase: According to the feature map, the node pairs and their connecting edges are projected into a unified feature space through a triple attention network, and the interaction intensity is calculated through a non-linear transformation; ; where is the interaction intensity, represents the LeakyReLU activation function, U is a learnable weight vector, represents the transpose, and are two different trainable parameter matrices, and respectively represent the hidden states of node at time step and node , is the edge feature between node and node , represents vector concatenation; Generate an attention distribution through normalization according to the interaction intensity; ; where is the attention distribution, is the exponential function with the natural constant e as the base, is the node 's neighborhood set, is the node and the node interaction intensity; Through nodes Hidden state And edge features The aggregated message of the node is obtained by weighted summation ; ; In the formula, Represents the aggregated message of node , Is the attention coefficient, And Are two different trainable weight matrices, Is the hidden state of node , Represents the Hadamard product; According to the aggregated message, multiple groups of messages are generated in parallel by using the multi-head attention mechanism and concatenated to obtain multi-dimensional features; ; where, is the multi-dimensional feature, represents concatenation, is the attention coefficient calculated by the th attention head, is the first weight matrix of the input linear transformation, is the second weight matrix of the input linear transformation; Using a gated recurrent unit as the node update function, fuse the previously extracted message with the current integrated feature to achieve temporal feature integration, introduce layer normalization to alleviate gradient anomalies, and gradually optimize the node representation through T iterations to output high-order features; ; where, is the high-order feature, LN represents layer normalization, GRU represents the gated recurrent unit, is the currently integrated feature; The molecular graph neural network based on the ternary message passing mechanism performs the following steps in the readout phase: Update the hidden state using LSTM; ; where, is the hidden state, represents the LSTM network, represents the aggregated feature of the previous time step; Calculate the attention weights of the nodes according to the hidden state; ; where is the attention weight, is the softmax activation function, is the final hidden state of node i; Aggregate node features according to the attention weights; ; where is the aggregated node feature, is the total number of nodes; After aggregating the node features, through T iterations, the molecular graph features of the drug are finally obtained; ; where, is the molecular graph feature of the drug, is the hidden state of the t-th iteration; The molecular graph features are defined as structural features.
3. A method for predicting the risk of a drug combination based on multi-source feature fusion and contrastive learning according to claim 1, characterized in that, The molecular graph features generated by TrimNet are input into a multi-layer cascaded graph convolutional network to extract high-order topological features in the drug network and obtain relational features, including: Input the molecular graph features generated by TrimNet into a two - layer cascaded graph convolutional network to extract relational features; the extraction model of relational features is: ; where is the output of the first - layer graph convolutional network, is the output of the second - layer graph convolutional network, is the relational feature, represents the residual connection, represents the first - layer graph convolutional operation, represents the second - layer graph convolutional operation, is the initial graph representation output by TrimNet, is the edge index set, and AttnWeight represents the attention mechanism; The steps of a single-layer graph convolution are as follows: Using the molecular graph features of the said drug as the initial features, perform graph convolution; ; where, is the feature after convolution, is the graph convolution, is the convolutional layer, is the initial feature, is the edge index set, is the degree matrix, is the adjacency matrix with self-loops, is the trainable parameter matrix; Process the convolved features through batch normalization, ReLU activation, and residual connection; ; where, is the output feature, is dropout, is the ReLU activation function, is batch normalization, is the dimensional projection function; Calculate the node attention weights through a shallow perception network based on the processed features; ; where is the node attention weight, is the tanh activation function, is the attention parameter matrix.
4. A method for predicting the risk of a drug combination based on multi-source feature fusion and contrastive learning according to claim 1, characterized in that, According to the drug dataset, similarity features are calculated based on Morgan molecular fingerprints and Tanimoto coefficients, and interactive embedding features are extracted, including: According to the drug data set, obtaining a Morgan molecular fingerprint of the drug combination; Calculate the similarity between two drug molecular fingerprints according to the Morgan molecular fingerprint of the drug combination, and obtain the similarity features; ; In the formula, is the similarity, and respectively represent the fingerprint vectors of two drug molecules; Introduce the identity matrix as structured prior information, map the identity matrix to the target embedding space via a linear projection layer, and generate a set of orthogonal basis vectors; ; where is the orthogonal basis vector, is the identity matrix, represents the linear projection layer, represents a real number, is the total number of nodes, is the dimension of the projected embedding space; Project the similarity features into the target embedding space, and then perform weighted mixing with the orthogonal basis vectors to obtain a mixed embedding matrix; ; where is the new similarity feature generated by projecting into the embedding space, is the mixed embedding matrix, is a learnable weight parameter; For a drug combination and , the mixed embedding vectors corresponding to the drug combination are retrieved from the mixed embedding matrix index and ; ; where and respectively represent the mixed embedding vectors of drugs and drug , and are the mixed embedding matrices of two drug sets respectively; According to the mixed embedding vector and , capture the synergistic effect of the drug pair through the Hadamard product, and apply the L2 normalization operation to obtain the interaction embedding features; ; where is the interaction embedding feature, represents the L2 normalization operation, represents the Hadamard product, represents the L2 norm.
5. A method for predicting the risk of drug combinations based on multi-source feature fusion and contrastive learning according to claim 1, characterized in that According to the structural features and the similarity features, potential associations and differences between the features are mined through a comparative learning mechanism to obtain structural comparative learning features and similarity comparative learning features, specifically including: Given the original model parameters of the heterogeneous graph transformer, calculate the gradient sign direction of its corresponding loss function; the calculation model of the gradient sign direction is: ; where is the gradient sign direction, is the sign function, represents the gradient direction, represents the gradient, are the original model parameters, is the corresponding loss function. Introduce a loss weight factor to dynamically adjust the noise intensity according to the current loss value; ; where is the loss weight factor, is the current loss value, is the smoothing factor; Perturb the parameters of the GNN layer of the model according to the loss weight factor, and the perturbed parameters are: ; where represents the perturbed GNN layer parameters, is a learnable weight parameter used to balance the contributions of the gradient direction and Gaussian noise, is the gradient sign direction, is random noise that follows a Gaussian distribution; Encode using a heterogeneous graph transformer and its perturbed version as dual encoders to obtain the original view and the perturbed view; ; where, and respectively represent the original views generated by inputting the structural features and similarity features into the heterogeneous graph transformer, and respectively represent the perturbed views generated by the perturbed heterogeneous graph transformer, ( ) is the original encoder, ( ) is the encoder after applying the mixed perturbation, is the graph representation; Map the representations of the original view and the perturbed view to the latent space through a non - linear projection head; ; where, and are the structural feature and the similarity feature of the original view respectively, the latent feature representations obtained by the non - linear transformation of the MLP, and are the structural feature and the similarity feature of the perturbed view after perturbation respectively, the latent feature representations obtained by the non - linear transformation of the MLP, represents the non - linear projection head, is the non - linear projection head corresponding to the original encoder, is the non - linear projection head corresponding to the encoder after perturbation, is the multi - layer perceptron, and are two different learnable parameter matrices respectively; Obtain the structure contrast learning feature and the similarity contrast learning feature according to the latent feature representation mapped to the latent space; ; where is the structure contrast learning feature, is the similarity contrast learning feature, is average pooling.
6. The method for predicting the risk of a drug combination based on multi-source feature fusion and contrastive learning according to claim 1, wherein The structural feature, the relationship feature, the structural contrast learning feature, the similarity contrast learning feature, and the interactive embedding feature are subjected to feature fusion and feature alignment to obtain a fusion feature, specifically including: Concatenate the structural features and the relational features to obtain new drug structural features ; ; In the formula, is the drug structural feature, is the concatenated feature, represents vector concatenation, is the feature after residual projection, and are two different learnable parameters respectively, represents the ReLU activation function, is the structural feature, is the relational feature, is the non-linear bias vector; Adopt a dual-channel fusion method to fuse the structural contrast learning features and the similarity contrast learning features, and obtain the contrast learning features ; Align the structural features contrastive learning features and the previous interaction embedding features to obtain aligned features; ; where represents the aligned features, FFN is a two-layer feed-forward network, is the normalization operation, is the attention weight matrix, is the Softmax activation function, represents the transpose, is the query, is the key, is the value, is the hidden layer dimension, are respectively the projection matrices of According to the alignment feature, a learnable affine transformation is introduced to dynamically adjust the fused feature, and the finally output fused feature is obtained. ; ; where and are respectively different trainable recalibration parameters.
7. A method for predicting the risk of drug combinations based on multi-source feature fusion and contrastive learning according to claim 1, characterized in that The drug dataset contains SMILES sequences of drug combinations.
8. A drug combination risk prediction device based on multi-source feature fusion and contrast learning, characterized in that, Include: A data set acquisition module, used to acquire drug data sets; A feature graph construction module, used for obtaining a feature graph according to the drug dataset; A structural feature extraction module, which is used to perform adaptive feature extraction of the molecular graph structure through a molecular graph neural network based on a ternary message passing mechanism according to the feature map, so as to obtain the molecular graph features of the drug; define the molecular graph features as structural features ; The relationship feature extraction module is used to input the molecular graph features generated by TrimNet into a multi-level cascaded graph convolutional network, extract high-order topological features in the drug network, and obtain relationship features ; An interactive embedding feature extraction module, configured to calculate similarity features based on the Morgan molecular fingerprint and the Tanimoto coefficient according to the drug data set and extract interactive embedding features ; A contrastive learning module, configured to, according to the structural features and the similarity features, through a contrastive learning mechanism, discover potential associations and differences between the features, and obtain structural contrastive learning features and similarity contrastive learning features ; A feature fusion module, configured to perform feature fusion and feature alignment on the structural feature, the relational feature, the structural contrast learning feature, the similarity contrast learning feature, and the interaction embedding feature, so as to obtain a fused feature ; The risk prediction module is used to determine the risk level of the drug combination based on the fusion features and use a classification model to obtain the drug combination risk.
9. A drug combination risk prediction device based on multi-source feature fusion and contrast learning, characterized in that, It includes a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement a drug combination risk prediction method based on multi-source feature fusion and comparative learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a drug combination risk prediction method based on multi-source feature fusion and comparative learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-relation comparative learning drug interaction prediction method, system, medium and equipment
CN118173198A
Drug target interaction prediction method based on multi-source feature interaction
CN118212974A
Pathological analysis method, device and equipment based on machine learning and storage medium
CN118280585A
Drug interaction prediction method based on multilayer attention and elastic message passing
CN119314698A
Multi-modal data-based disability intervention quality evaluation method and device and readable storage medium
CN119480036A
Cited By
Catalyst prediction method and system based on molecular characterization contrast learning
CN121034442A
Polymer thermal property prediction method based on gated multi-modal fusion and multi-task learning
CN121096478A
Drug full life cycle risk monitoring and early warning method based on multi-source data fusion
CN121215308A
A Drug Lifecycle Risk Monitoring and Early Warning Method Based on Multi-Source Data Fusion
CN121215308B
Dual-channel drug interaction prediction method of positive definite non-commutative polynomial filter
CN121260529A