Task processing method, training method and device based on intermolecular interaction

By introducing embedded feature representations of node connection edges between molecular graphs and a loss value training method, the problem of insufficient modeling of intermolecular interactions is solved, and the accuracy of task processing is improved, especially in solvent-solute and drug molecule interaction analysis, providing more accurate predictions.

CN121148518APending Publication Date: 2025-12-16SUZHOU INST FOR ADVANCED STUDY USTC +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511297050.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing technologies lack modeling of intermolecular interactions in molecular graph embedding feature representations, resulting in the embedding representations being isolated from the interaction environment and reducing the accuracy of task processing.

Method used

By introducing node connection edges between molecular graphs, embedding feature representations are performed to identify the features of the active nodes, determine the target substructure, and make predictions in regression or classification tasks. The model is trained by combining structural and outcome loss values.

Benefits of technology

It improves the accuracy of identifying target substructures and enhances the accuracy of task processing, especially providing more accurate predictive capabilities in the analysis of solvent-solute and drug molecule interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148518A_ABST
    Figure CN121148518A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method based on intermolecular interaction, a training method based on intermolecular interaction and a device thereof. Based on a connecting edge formed between a first node in the first molecular graph and a second node in the second molecular graph, respectively performing embedding feature representation on the first node and the second node to obtain a first node embedding feature and a second node embedding feature; identifying a first action node feature and a second action node feature; based on the first action node feature and the second action node feature, determining target substructures in chemical structures indicated by the first molecular graph and the second molecular graph respectively; the task matched with the first molecule and the second molecule is a regression task, and performing regression prediction on the target substructures corresponding to the first molecule graph and the second molecule graph to obtain a processing result; the task matched with the first molecule and the second molecule is a classification task, classification is carried out based on the target substructures corresponding to the first molecule graph and the second molecule graph, and a classification result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical fields of artificial intelligence, molecular biology and computer science, and particularly relates to a task processing method based on intermolecular interaction, a training method and a device thereof. BACKGROUND

[0002] Molecular relationship learning aims to represent the characteristics of intermolecular interaction, such as potential drug-drug interactions and chromophores in different solvents, and has attracted wide attention. The molecular core structure embodies the essence of the physical and chemical characteristics of the molecule in the molecular interaction. In the regression or classification task processing of the intermolecular interaction relationship by using the molecular graph, the key information is extracted from the embedded feature representation of the nodes in the single molecular graph, and then the task processing result is obtained based on the key information.

[0003] In the process of implementing the present application concept, it is found that although the task processing result can be obtained based on the key information, due to the lack of modeling of the interaction between the molecular graph and the information in the molecular graph in the embedded feature representation, the embedded representation is isolated from the interactive environment, which reduces the accuracy of extracting the key information, and further reduces the accuracy of the task processing. SUMMARY

[0004] Therefore, the present application provides a task processing method based on intermolecular interaction, a training method and a device thereof.

[0005] One aspect of the present application provides a task processing method based on intermolecular interaction, comprising: based on a connection edge formed between a first node in a first molecular graph and a second node in a second molecular graph, respectively embedding feature representation of the first node and the second node to obtain first node embedding feature and second node embedding feature, the first molecular graph indicating a chemical structure of a first molecule, and the second molecular graph indicating a chemical structure of a second molecule; based on the association relationship between the first node and the second node, respectively identifying first action node feature and second action node feature of the nodes used for interaction from the first node embedding feature and the second node embedding feature; based on the first action node feature and the second action node feature, determining a target substructure in the chemical structure indicated by the first molecular graph and the second molecular graph respectively, which occurs interaction; when it is determined that the task matched with the first molecule and the second molecule is a regression task, performing regression prediction on the target substructure corresponding to the first molecular graph and the second molecular graph respectively to obtain a processing result of an index used to represent the interaction between the first molecule and the second molecule; when it is determined that the task matched with the first molecule and the second molecule is a classification task, classifying whether the first molecule and the second molecule have interaction based on the target substructure corresponding to the first molecular graph and the second molecular graph respectively to obtain a classification result.

[0006] Another aspect of the present application provides a training method of a task processing model based on intermolecular interaction, comprising: based on a sample connection edge formed between a sample first node in a sample first molecular graph and a sample second node in a sample second molecular graph, embedding feature representations of the sample first node and the sample second node respectively to obtain a sample first node embedding feature and a sample second node embedding feature, the sample first molecular graph indicating a sample chemical structure of a sample first molecule, and the second molecular graph indicating a sample chemical structure of a sample second molecule; based on the association relationship between the sample first node and the sample second node, identifying sample first action node features and sample second action node features of sample nodes that interact from the sample first node embedding features and the sample second node embedding features respectively; based on the sample first action node features and the sample second action node features, determining sample target substructures in the sample chemical structures indicated by the sample first molecular graph and the sample second molecular graph respectively that interact; when it is determined that the task matched with the sample first molecule and the sample second molecule is a regression task, performing regression prediction on the target substructures corresponding to the sample first molecular graph and the sample second molecular graph respectively to obtain a sample processing result for an index representing the interaction between the sample first molecule and the sample second molecule; when it is determined that the task matched with the sample first molecule and the sample second molecule is a classification task, classifying whether the sample first molecule and the sample second molecule have an interaction based on the sample target substructures corresponding to the sample first molecular graph and the sample second molecular graph respectively to obtain a sample classification result; based on the sample first action node features and the sample second action node features, determining a structure loss value; based on the sample processing result and a real processing result, or based on the sample classification result and a real classification result, determining a result loss value; based on the structure loss value and the result loss value, training a task processing model matched with the task to obtain a trained task processing model.

[0007] Another aspect of the present application provides a task processing device based on intermolecular interaction, comprising: an embedding module configured to obtain first node embedding features and second node embedding features by embedding feature representations of a first node in a first molecular graph and a second node in a second molecular graph respectively based on a connection edge formed between the first node and the second node, the first molecular graph indicating a chemical structure of a first molecule, and the second molecular graph indicating a chemical structure of a second molecule; an identification module configured to identify first action node features and second action node features of nodes that interact with each other from the first node embedding features and the second node embedding features respectively based on a correlation between the first node and the second node; a determination module configured to determine target substructures in the chemical structures indicated by the first molecular graph and the second molecular graph respectively based on the first action node features and the second action node features; a first processing module configured to perform regression prediction on the target substructures corresponding to the first molecular graph and the second molecular graph respectively to obtain a processing result of an index representing an interaction between the first molecule and the second molecule when it is determined that a task matched with the first molecule and the second molecule is a regression task; and a second processing module configured to classify whether the first molecule and the second molecule have an interaction based on the target substructures corresponding to the first molecular graph and the second molecular graph respectively to obtain a classification result when it is determined that the task matched with the first molecule and the second molecule is a classification task.

[0008] Another aspect of the present application provides a training device of a task processing model based on intermolecular interaction, comprising: a sample embedding module configured to obtain sample first node embedding features and sample second node embedding features by embedding feature representation of sample first nodes and sample second nodes based on sample connection edges formed between the sample first nodes and the sample second nodes, wherein the sample first molecular graph indicates a sample chemical structure of a sample first molecule, and the sample second molecular graph indicates a sample chemical structure of a sample second molecule; a sample identification module configured to identify sample first action node features and sample second action node features of sample nodes used for interaction from the sample first node embedding features and the sample second node embedding features based on a correlation between the sample first nodes and the sample second nodes; a sample determination module configured to determine sample target substructures in the sample chemical structures indicated by the sample first molecular graph and the sample second molecular graph based on the sample first action node features and the sample second action node features; a sample first processing module configured to perform regression prediction on the target substructures corresponding to the sample first molecular graph and the sample second molecular graph to obtain a sample processing result used to represent an index of interaction between the sample first molecule and the sample second molecule when it is determined that a task matched with the sample first molecule and the sample second molecule is a regression task; a sample second processing module configured to classify whether the sample first molecule and the sample second molecule have interaction based on the sample target substructures corresponding to the sample first molecular graph and the sample second molecular graph to obtain a sample classification result when it is determined that the task matched with the sample first molecule and the sample second molecule is a classification task; a first loss determination module configured to determine a structure loss value based on the sample first action node features and the sample second action node features; a second loss determination module configured to determine a result loss value based on the sample processing result and a real processing result or based on the sample classification result and a real classification result; and a training module configured to train a task processing model matched with the task based on the structure loss value and the result loss value to obtain a trained task processing model.

[0009] According to the embodiments of the present application, since the connection edges formed between the nodes of different molecular graphs are introduced when embedding feature representation, the first node embedding features and the second node embedding features are not calculated based on a single molecule in isolation, but are fusion representations deeply embedding context information of interaction between two molecules, which enhances information interaction between molecules in the embedding representation stage. Furthermore, based on the correlation between the nodes across molecules, the action node features of the nodes used for interaction are identified, which can improve the accuracy of identifying target substructures and further improve the accuracy of task processing. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:

[0011] Figure 1 A flow chart of a task processing method based on intermolecular interaction according to an embodiment of the present application is shown;

[0012] Figure 2 A flow chart of a training method of a task processing model based on intermolecular interaction according to an embodiment of the present application is shown;

[0013] Figure 3 A target substructure visualization result chart under different iteration numbers is shown, taking a solubility reaction of amide in nitro solvent as an example, according to an embodiment of the present application, and the darker the color, the greater the weight;

[0014] Figure 4 A block diagram of a task processing device based on intermolecular interaction according to an embodiment of the present application is shown;

[0015] Figure 5 A block diagram of a training device of a task processing model based on intermolecular interaction according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0016] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that the description is merely exemplary of the present application, and is not intended to limit the scope of the present application. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without these specific details. In other instances, well-known structures and techniques have not been described in detail in order to avoid obscuring aspects of the present application.

[0017] The terms used herein are merely used to describe specific embodiments, and are not intended to limit the present application. The terms "include", "comprise" and the like used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0018] All terms used herein (including technical and scientific terms) have meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present specification, and should not be interpreted in an idealized or overly formal manner.

[0019] In the case of using expressions such as "at least one of A, B, and C", it will be understood that the meaning was intended to include any of the individual members when referring to the multiple member set {A, B, C} (e.g., referring to "a system having at least one of A, B, and C" is intended to cover a system having A alone as well as a system having B alone as well as a system having C alone).

[0020] In the process of implementing the present concept, it is found that although the selection of the substructure of a molecule in the related art can be significantly affected by another molecule, more accurate substructures can be extracted by iteratively extracting interactive substructures when modeling the embedded feature representation, however, due to the interaction between the information in the two molecular graphs depends on how the molecule indicated by the molecular graph interacts with another molecule, for example, a hydrophobic group is a key substructure when interacting with another hydrophobic molecule, but it can be irrelevant when interacting with a strongly polar molecule, therefore, due to the lack of interaction signals, the two molecular graphs will fail to embed the information of the interactive object which is crucial when initially embedding the representation, thus lacking the key context. Since the preliminary decision error will reduce the accuracy and reliability of substructure extraction, thus reducing the accuracy of using molecular graphs to process the regression or classification tasks of the interaction relationship between molecules.

[0021] Based on this, embodiments of the present application provide a task processing method based on intermolecular interaction, a training method and a device thereof. The method comprises: based on the connection edge formed between the first node in the first molecular graph and the second node in the second molecular graph, embedding feature representations of the first node and the second node respectively to obtain first node embedding features and second node embedding features, the first molecular graph indicating the chemical structure of the first molecule, and the second molecular graph indicating the chemical structure of the second molecule; based on the association relationship between the first node and the second node, identifying first action node features and second action node features of the nodes used to interact from the first node embedding features and the second node embedding features respectively; based on the first action node features and the second action node features, determining target substructures in the chemical structures indicated by the first molecular graph and the second molecular graph respectively which interact; when determining that the task matched with the first molecule and the second molecule is a regression task, performing regression prediction on the target substructures corresponding to the first molecular graph and the second molecular graph respectively to obtain a processing result of an index used to represent the interaction between the first molecule and the second molecule; when determining that the task matched with the first molecule and the second molecule is a classification task, classifying whether the first molecule and the second molecule have an interaction based on the target substructures corresponding to the first molecular graph and the second molecular graph respectively to obtain a classification result.

[0022] The following will be described through Figure 1The task processing method based on intermolecular interaction is described in detail.

[0023] Figure 1 A flowchart of the task processing method based on intermolecular interaction is shown.

[0024] As shown in Figure 1 The task processing method based on intermolecular interaction includes operations S110-S150.

[0025] In operation S110, based on the connection edge formed between the first node in the first molecular graph and the second node in the second molecular graph, the first node and the second node are respectively embedded to obtain the first node embedding feature and the second node embedding feature.

[0026] In operation S120, based on the association relationship between the first node and the second node, the first action node feature and the second action node feature of the node that interacts are respectively identified from the first node embedding feature and the second node embedding feature.

[0027] In operation S130, based on the first action node feature and the second action node feature, the target substructure that interacts in the chemical structure indicated by the first molecular graph and the second molecular graph is determined.

[0028] In operation S140, when it is determined that the task matched with the first molecule and the second molecule is a regression task, the target substructure corresponding to the first molecular graph and the second molecular graph is respectively subjected to regression prediction to obtain a processing result of an index for characterizing the interaction between the first molecule and the second molecule.

[0029] In operation S150, when it is determined that the task matched with the first molecule and the second molecule is a classification task, based on the target substructure corresponding to the first molecular graph and the second molecular graph, whether the first molecule and the second molecule have interaction is classified to obtain a classification result.

[0030] In the embodiment of the application, the first molecular graph can indicate the chemical structure of the first molecule. The second molecular graph can indicate the chemical structure of the second molecule. The first node characterizes any atom in the first molecule. The second node characterizes any atom in the second molecule. The connection edge is used to characterize the potential interaction strength between any atom in the first molecule and any atom in the second molecule.

[0031] The embedding feature representation can be obtained by message passing calculation through a graph neural network, for example. The first node embedding feature is used to represent the embedding representation of the first molecule that fuses the interaction information with the second molecule. The second node embedding feature is used to represent the embedding representation of the second molecule that fuses the interaction information with the first molecule.

[0032] For example, the feature of the first molecule can be updated based on the aggregated feature obtained by aggregating the features of the nodes adjacent to the first node in the first molecular graph and the features of the connection edges between the first node and the second node, to obtain the first node embedding feature. The number of first nodes in the first molecular graph can be determined by the first molecule, and the first node embedding feature of any first node can be obtained in the same way, which will not be described here.

[0033] The second node embedding feature can also be obtained by symmetric operation in the same way as obtaining the first node embedding feature, which will not be described here.

[0034] The first action node feature can represent the feature of the interactive substructure extracted from the first node embedding feature. The second action node feature can represent the feature of the interactive substructure extracted from the second node embedding feature.

[0035] The first action node feature can be mapped to the target substructure in the chemical structure indicated by the first molecular graph that interacts. The second action node feature can be mapped to the target substructure in the chemical structure indicated by the second molecular graph that interacts.

[0036] The index of the interaction between the first molecule and the second molecule can include at least one of the following: dissociation constant, solubility, free energy change, etc. For example, for the dissociation constant or solubility representing the interaction between the first molecule and the second molecule, the target substructure corresponding to the first molecular graph and the second molecular graph can be embedded and fused to obtain the fusion feature, which is then input into a regression predictor. The continuous value output by the regression predictor is in the logarithmic form of the dissociation constant or solubility. The negative logarithmic form of the dissociation constant or solubility is inversely transformed to obtain the processing result of the dissociation constant or solubility representing the interaction between the first molecule and the second molecule. For the free energy change representing the interaction between the first molecule and the second molecule, the fusion feature can be input into a regression predictor, and the continuous value output by the regression predictor is the free energy change. The regression predictor can be trained, and how to train the regression predictor is not limited in the embodiments of the present application.

[0037] For example, the fusion feature can be input into a binary classifier to output a classification result representing whether the first molecule and the second molecule have an interaction.

[0038] According to the embodiment of the present application, since the connection edges formed by the nodes between different molecular graphs are introduced when embedding the feature representation, the first node embedding feature and the second node embedding feature are not calculated based on a single molecular isolation, but are the fusion representation of the deep embedding of the context information of the interaction of the two molecules, which enhances the information interaction between molecules in the embedding representation stage. Further, based on the association relationship between the nodes across the molecules, the role node features of the nodes for interaction are identified, which can improve the accuracy of identifying the target substructure, and further improve the accuracy of task processing.

[0039] According to another embodiment of the present application, in addition to operations S110-S150 as shown in the above, the method for processing tasks based on the interaction between molecules can further include the following operations: in the case where the first molecule is determined to be a solvent molecule and the second molecule is determined to be a solute molecule, it is determined that the task matched with the first molecule and the second molecule is a regression task; in the case where the first molecule and the second molecule are both determined to be drug molecules, it is determined that the task matched with the first molecule and the second molecule is a classification task. Figure 1

[0040] The solute molecule can be dissolved in the solvent molecule. The drug molecule can include, but is not limited to, at least one of the following: a molecular compound, a biological macromolecule (such as a protein, a nucleic acid, etc.), or a derivative thereof, etc.

[0041] The index of the interaction between the first molecule and the second molecule can be used to represent the thermodynamic properties of the solute molecule in the solvent molecule, such as dissociation constant, solubility, free energy change, etc.

[0042] According to the embodiment of the present application, by performing the regression task processing for the solvent-solute molecule pair, the solubility reaction of the solvent and the solute can be accurately analyzed in an interpretable manner. By performing the classification task processing for the drug molecule pair, it can be determined whether the two drug molecules have interaction, thereby providing a reference for the research of drug molecules.

[0043] According to the embodiment of the present application, the connection edge can include a first directed connection edge from the first node to the second node and a second directed connection edge from the second node to the first node. For operation S110 as shown in the above, based on the connection edge formed between the first node in the first molecular graph and the second node in the second molecular graph, the embedding feature representation of the first node and the second node is obtained, which can include the following operations: determining the attribute features of the first directed connection edge, the second directed connection edge, the first node, and the second node respectively; obtaining the first node embedding feature based on the merging of the attribute features of the second directed connection edge and the first node; obtaining the second node embedding feature based on the merging of the attribute features of the first directed connection edge and the second node. Figure 1 ​​

[0044] In the embodiments of the present application, the attribute features of the first node and the second node can be used to represent the properties of the atoms, which can include, but are not limited to, electronegativity, radius, charge, hybridization state, bond energy, oxidation state, electrophilicity, nucleophilicity, hydrophobicity, etc. The attribute features of the first directed connection edge and the second directed connection edge can represent the properties of the bond formed between the molecules, which can include, but are not limited to, the type of the bond, the strength of the bond, the length of the bond, the polarity of the bond, the electron cloud density of the bond, the vibration frequency of the bond, the reactivity of the bond, the stereochemistry of the bond, the resonance structure of the bond, the electron transfer of the bond, the thermodynamic and kinetic properties of the bond, the optical properties of the bond, the magnetic properties of the bond, and the charge distribution of the bond, etc. The directionality of the first directed connection edge and the second directed connection edge is used to distinguish the interaction behaviors of the atoms in the first molecular graph and the second molecular graph.

[0045] The features of all the directed connection edges connected to the first node can be aggregated to obtain aggregated edge features, and the aggregated edge features and the attribute features of the first node are combined to obtain the first node embedding features. Similarly, the second node embedding features can be obtained by using a method similar to that of the first node embedding features, which will not be described here.

[0046] According to the embodiments of the present application, by combining the attribute features of the first node or the second node itself and the attribute features of the first directed connection edge or the second directed connection edge, the node embedding features of each node are determined, which can more comprehensively capture the structure and attribute information of the node, thereby facilitating the extraction of the substructure of the interaction between the information in the molecular graph, and providing rich feature representations for the property prediction between the solvent and solute molecules and the classification between the drug molecules.

[0047] According to the embodiments of the present application, based on the combination of the attribute features of the second directed connection edge and the first node, the first node embedding features can include operations of: fusing the attribute features of the second node with similar features to obtain first fusion features; processing the first fusion features through a fully connected layer and an activation layer in sequence to obtain first node conversion features; based on a first parameter, combining the attribute features of the second directed connection edge and the first node conversion features to obtain first combined features; processing the first combined features through a fully connected layer and an activation layer in sequence to obtain first edge conversion features; based on a second parameter, combining the first edge conversion features and the attribute features of the first node to obtain the first node embedding features.

[0048] In the embodiments of the present application, the similar features can be used to represent the similarities between the attribute features of the first node and the second node. The determination method of the similar features is not limited in the present application.

[0049] The first parameter can be used to control the merging ratio of the attribute features of the second directed connection edge. The second parameter can be used to control the fusion ratio of the attribute features of the first node.

[0050] For example, for a first molecular graph and a second molecular graph , the preliminary interaction of and can be simulated by means of the first directed connection edge and the second directed connection edge. The first node i1 and the second node j2 correspond to a pair of directed connection edges, i.e. the first directed connection edge from i1 to j2 and the second directed connection edge from j2 to i1.

[0051] Message passing can be performed in the molecular interior indicated by and respectively. The feature of the edge between the node i and the node j indicated by or , the delivery process can be shown in the following formula (1)~formula (3):

[0052] (1)

[0053] (2)

[0054] (3)

[0055] wherein, and are activation functions, is a fully connected layer, is the initial attribute feature of the node i in or , is the initial attribute feature of the node j in or , is the graph-level feature of or , is an element-wise multiplication, is a smoothing constant, thereby the attribute feature of the node i in or , .

[0056] Taking the attribute feature of the first node i1 in , and the attribute feature of the second node j2 in , as an example, the first merged feature can be obtained by the following formula (4):

[0057]

[0057] (4)

[0058] wherein, is a first parameter, is an attribute feature of a second directed connection edge pointing from j2 to i1.

[0059] In order to enhance the interaction of information across molecules, a large number of connection edges are introduced between the two molecules. However, the introduction of a large number of connection edges may lead to over-smoothing problems. Based on this, when determining the first node embedding feature, a mask vector , is introduced to randomly mask part of the second directed connection edge. Therefore, the first node embedding feature can be obtained by the following formula (5) :

[0060] (5)

[0061] wherein, is a second parameter.

[0062] According to the embodiments of the present application, by introducing the attribute feature of the second node, the similarity feature and the attribute feature of the second directed connection edge, and processing through the full connection layer and the activation layer, the similarity and interaction between nodes can be captured. After merging with the attribute feature of the first node, the first node embedding feature can more comprehensively reflect the role and function of the first node in the interaction of the two molecules, thereby enhancing the interaction of information across molecules and improving the understanding and prediction ability of the interaction between molecules.

[0063] According to the embodiments of the present application, based on the merging of the attribute features of the first directed connection edge and the second node respectively, the second node embedding feature can include operations: fusing the attribute feature of the first node with the similarity feature to obtain a second fusion feature; processing the second fusion feature through the full connection layer and the activation layer in turn to obtain a second node conversion feature; based on a third parameter, merging the attribute feature of the first directed connection edge and the second node conversion feature to obtain a second merging feature; processing the second merging feature through the full connection layer and the activation layer in turn to obtain a second edge conversion feature; based on a fourth parameter, merging the first edge conversion feature and the attribute feature of the second node to obtain the second node embedding feature.

[0064] In the embodiments of the present application, the third parameter can be used to control the merging ratio of the attribute feature of the first directed connection edge. The fourth parameter can be used to control the fusion ratio of the attribute feature of the second node.

[0065] For example, the second merging feature can be obtained by the following formula (6) :

[0066] (6)

[0067] wherein, is a third parameter, is an attribute feature of the first directed connection edge pointing from i1 to j2. Similar to introducing the mask vector when determining the first node embedding feature, a mask vector can also be introduced when determining the second node embedding feature for randomly masking part of the first directed connection edges.

[0068] The second node embedding feature can be obtained by the following formula (7) :

[0069] (7)

[0070] wherein, is a fourth parameter, is the number of atoms in the first molecule.

[0071] According to the embodiments of the present application, by introducing the attribute feature of the first node, the similarity feature and the attribute feature of the first directed connection edge, and processing through the full connection layer and the activation layer, the similarity and interaction between nodes can be captured. After merging with the attribute feature of the second node, the second node embedding feature can more comprehensively reflect the role and function of the second node in the interaction between the two molecules, thereby enhancing the information interaction between molecules and improving the understanding and prediction ability of the interaction between molecules.

[0072] According to the embodiments of the present application, for the operation S120 as shown in Figure 1 , based on the association relationship between the first node and the second node, the first action node feature and the second action node feature of the nodes for interaction are identified from the first node embedding feature and the second node embedding feature respectively, which can include the operation: performing T rounds of identification processing on the first node embedding feature and the second node embedding feature to obtain the Tth round of identification result, T represents a predetermined iteration round, and T is a positive integer; determining the first action node feature and the second action node feature according to the Tth round of identification result.

[0073] Performing T rounds of identification processing on the first node embedding feature and the second node embedding feature to obtain the Tth round of identification result can include the operation: identifying the second action node feature from the second node embedding feature based on the tth round parameter value of the predetermined parameter and the importance of the tth round second node to obtain the tth round first sub-identification result; updating the (t-1)th round second sub-identification result based on the tth round first sub-identification result to obtain the tth round second sub-identification result; and obtaining the tth round of identification result according to the tth round first sub-identification result and the tth round second sub-identification result.

[0074] In the embodiments of the present application, t is an integer greater than or equal to 1 and less than or equal to T. The t-th round parameter value is less than or equal to the termination parameter value, or the t-th round parameter value is greater than or equal to the starting parameter value. The predetermined parameter can be used to control the recognition accuracy of the second action node feature in the process of changing from the starting parameter value to the termination parameter value.

[0075] The predetermined parameter can be used to adjust the sensitivity to noise randomness. For example, the t-th round parameter value of the predetermined parameter can be obtained based on the following formula (8):

[0076] (8)

[0077] wherein, represents an initial value of the predetermined parameter in the initial stage, such as = 2.0. represents a termination value of the predetermined parameter in the end stage, such as = 0.1.

[0078] In order to ensure the differentiability in the sampling process, a technique for differentiable sampling from a discrete distribution can be used to calculate the t-th round discrete random variable for identifying the second action node feature according to the t-th round parameter value of the predetermined parameter and the importance of the t-th round second node , as shown in the following formula (9):

[0079] (9)

[0080] wherein, the importance of the t-th round second node i, from a uniform distribution, , represents an activation function.

[0081] The t-th round first sub-recognition result can be obtained according to the following formula (10):

[0082] (10)

[0083] wherein, represents a second node embedding feature, obeys a distribution, represents a plurality of first node embedding feature matrices.

[0084] The lower bound of evidence can be calculated according to the first sub-identification result of the tth round and the second sub-identification result of the (t-1)th round. Due to the symmetry of molecular interaction, the second sub-identification result of the tth round that maximizes the lower bound of evidence can be identified from the first node embedding feature based on a method similar to the above formula (8) to formula (10).

[0085] According to an embodiment of the present application, when t is much smaller than T, sampling is allowed to be performed in a softer way, so as to encourage the model to explore more potential node combinations in the early iterations. As t approaches T, the lower bound of evidence gradually approaches , and in this phase, the change of becomes steeper, making tend to a hard binary choice {0, 1}. This change helps to stabilize the convergence in the later stage, select a fixed substructure, and strengthen the information bottleneck effect, reduce the number of iterations, and save computing resources.

[0086] According to an embodiment of the present application, the task processing method based on intermolecular interaction can include operations S110 to S150 as shown in the figure, and can further include operations: obtaining tth round similar features based on the similarity between the second sub-identification result of the (t-1)th round and the second node embedding feature; performing weighted summation on the tth round similar features and the first node embedding feature to obtain tth round weighted features; and performing feature conversion and linear mapping on the tth round weighted features to obtain the importance of the tth round second node. Figure 1 The similarity between the second sub-identification result of the (t-1)th round

[0087] and the second node embedding feature can be obtained according to the following formula (11):

[0088] (11)

[0089] wherein, similarity matrix obtained by the similarity between each of the tth round all second node embedding features and each of the (t-1)th round second sub-identification result corresponding to all first nodes.

[0090] The importance of the tth round second node can be obtained according to the following formula (12):

[0091]

[0092] wherein, importance matrix obtained by the importance of the tth round all second nodes, and MLP represents a two-layer perceptron.

[0093] According to an embodiment of the present application, by determining the importance of the nodes, it is possible to facilitate the introduction of random noise in the nodes to promote the identification of the target substructures.

[0094] According to an embodiment of the present application, after the iteration is completed, the set2set network can be used to pool the target substructures corresponding to the first molecular graph and the second molecular graph respectively, to obtain a target substructure representation vector. The target substructure representation vector can be used for regression prediction or classification.

[0095] The following will be described in detail through Figure 2 The training method of the intermolecular interaction-based task processing model according to an embodiment of the present application will be described in detail.

[0096] Figure 2 A flowchart of the training method of the intermolecular interaction-based task processing model according to an embodiment of the present application is shown.

[0097] As Figure 2 shown, the training method of the intermolecular interaction-based task processing model includes operations S210-S280.

[0098] In operation S210, based on the sample connection edges formed between the sample first nodes in the sample first molecular graph and the sample second nodes in the sample second molecular graph, the sample first node embedding features and the sample second node embedding features are obtained by embedding the sample first nodes and the sample second nodes respectively.

[0099] In operation S220, based on the association relationship between the sample first nodes and the sample second nodes, the sample first action node features and the sample second action node features of the sample nodes that interact are identified from the sample first node embedding features and the sample second node embedding features respectively.

[0100] In operation S230, based on the sample first action node features and the sample second action node features, the sample target substructures that interact in the sample chemical structures indicated by the sample first molecular graph and the sample second molecular graph are determined.

[0101] In operation S240, when it is determined that the task matching the sample first molecule and the sample second molecule is a regression task, the target substructures corresponding to the sample first molecular graph and the sample second molecular graph are subjected to regression prediction to obtain a sample processing result for characterizing the index of the interaction between the sample first molecule and the sample second molecule.

[0102] In operation S250, when it is determined that the task matched with the sample first molecule and the sample second molecule is a classification task, a classification is performed on whether there is an interaction between the sample first molecule and the sample second molecule based on the respective sample target substructure corresponding to the sample first molecule graph and the sample second molecule graph, to obtain a sample classification result.

[0103] In operation S260, a structure loss value is determined based on the sample first action node feature and the sample second action node feature.

[0104] In operation S270, a result loss value is determined based on the sample processing result and the real processing result, or based on the sample classification result and the real classification result.

[0105] In operation S280, a task processing model matched with the task is trained based on the structure loss value and the result loss value, to obtain a trained task processing model.

[0106] In an embodiment of the present application, the sample first molecule graph can indicate a sample chemical structure of the sample first molecule, and the second molecule graph can indicate a sample chemical structure of the sample second molecule.

[0107] It should be noted that the specific implementation method of operations S210 to S250 can refer to the method of operations S110 to S150 described above, which will not be repeated here. Figure 1

[0108] According to the embodiment of the present application, since the structure loss value and the result loss value are combined to jointly train the task processing model matched with the task, the iterative identification of the target substructure can be further strengthened to identify the functional group with actual chemical semantics and function, thereby enhancing the accuracy of the task processing model.

[0109] In another embodiment of the present application, for operation S260, based on the sample first action node feature and the sample second action node feature, the structure loss value can include the following operations: based on the sample first action node feature and the sample first molecule graph, a first substructure loss value is determined; based on the sample second action node feature and the sample second molecule graph, a second substructure loss value is determined; based on the first substructure loss value and the second substructure loss value, the structure loss value is obtained.

[0110] ​For example, determining the first substructure loss value based on the features of the first active node of the sample and the first molecular graph of the sample may include the following operations: determining the importance of the first node of the sample based on the features of the first active node of the sample; applying sparsity constraints to the first node of the sample based on the first molecular graph of the sample and the importance of the first node of the sample to obtain a first constraint value; applying peak constraints to the first node of the sample based on the importance of the first node of the sample to obtain a second constraint value; applying clustering constraints to the first node of the sample based on the importance of the first node of the sample and the sample edges in the first molecular graph of the sample to obtain a third constraint value; and obtaining the first substructure loss value based on the first constraint value, the second constraint value, and the third constraint value.

[0111] Based on the first molecular map of the sample The first node of each sample The probability of being selected determines the importance of the first node in the sample. . This represents the set of the first nodes of the sample in the first molecular graph. This represents the set of sample edges in the first molecular graph of the sample.

[0112] First constraint value It can be obtained from the following formula (13):

[0113] (13).

[0114] Second constraint value It can be obtained from the following formula (14):

[0115] (14)

[0116] Third constraint value We can obtain the following from equations (15) to (17):

[0117] (15)

[0118] (16)

[0119] (17)

[0120] Where Z represents the median, This represents the first node of the sample in the first molecular graph. The converted value is approximately 1; k represents the slope control parameter. This indicates the first node of the sample in the first molecular graph. Other sample nodes with connection relationships The transformed value can approach 0; each sample edge in the first molecular graph is represented as... , is a weight of the sample edge , is a bias of the sample edge , represents an importance degree of other sample nodes having a connection relationship with the sample first node .

[0121] The first substructure loss value can be obtained by weighted sum of the first constraint value, the second constraint value and the third constraint value.

[0122] It should be noted that the method for determining the second substructure loss value can be similar to the method for determining the first substructure loss value, which will not be described here.

[0123] The first substructure loss value and the second substructure loss value can be summed to obtain the structure loss value.

[0124] The following takes i=1 as the sample first molecular graph and i=2 as the sample second molecular graph to obtain the structure loss value which can be shown as the following formula (18):

[0125] (18)

[0126] , , respectively represent the sparsity constraint weight, the peak constraint weight and the cluster constraint weight.

[0127] According to the embodiments of the present application, the sparsity constraint can encourage to retain only a few nodes. The peak constraint can ensure that there is at least one high-confidence node. Through the cluster constraint, the high-weight nodes can be focused on a certain connected substructure. Thus, the combination of the three to determine the structure loss can identify the functional groups with actual chemical semantics and functions.

[0128] The task processing model obtained according to the embodiment of the present application: for a regression task, for example, an evaluation experiment can be performed using a data set containing organic molecules and their optical properties, such as the Chromophore data set, using the root mean square error evaluation index to measure the difference between the processing result of the task processing model and the actual value, for example, the root mean square error of the three indicators of absorption, emission and excited state lifetime can be as low as 16.06, 22.88, 0.702. The data set describing the solvation free energy of the solute-solvent pair, such as MNSol, FreeSolv, CompSol, Abraham and Combi-Solv, can also be used to perform evaluation experiments, and the root mean square error of the free energy indicators of these data sets can be as low as 0.545, 0.663, 0.250, 0.321, 0.373. Since the smaller the root mean square error value, the higher the processing accuracy of the task processing model, the smaller the error. Therefore, the task processing model obtained according to the embodiment of the present application has higher accuracy than the IGIB-ISE and ISE models in the prior art.

[0129] For a classification task, such as a drug-drug interaction classification task, three drug-drug interaction data sets recording adverse reactions between drug-drug pairs, such as ZhangDDI, ChChMiner and DeepDDI, can be used to perform evaluation experiments, using the classification accuracy evaluation index to measure the classification accuracy of the task processing model, such as classifying on the classes seen in the training set, the classification accuracy of the three data sets is 89.88%, 96.13%, and 96.98%, respectively. On the classes not seen during training, the first ensures that at least one drug molecule in the test data set is not seen in the training data set, and the classification accuracy of the three data sets is 70.88%, 81.40%, and 76.81%, respectively. And the first ensures that two drug molecules in the test data set are not present in the training data set, and the classification accuracy of the three data sets is 60.86%, 70.76%, and 71.88%, respectively. Compared with the IGIB-ISE and ISE models in the prior art, the task processing model has higher classification accuracy and strong generalization ability. It has potential development prospects in processing emerging drug molecules.

[0130] Since the task processing model in the embodiment of the present application iteratively identifies more accurate target substructures, it can ensure that more accurate information is provided in task processing, thereby improving the performance of the model.

[0131] Figure 3 The target substructure visualization results under different iterations (Iterations) are shown according to the embodiment of the present application, taking the solubility reaction of amide in nitro solvent as an example, and the deeper the color, the greater the weight.

[0132] For example, taking the solubilization reaction of amide in nitro solvent as an example, in the reaction of solubilizing amide ‘CNC=O’ in nitrosomethane ‘C[N+]([O-])=O’, the target substructure is the N-H atom in the amide (providing a hydrogen bond) and the two oxygen atoms in the nitro solvent (accepting a hydrogen bond, especially the negatively charged O - ), which together constitute an intermolecular hydrogen bond network, which is the key to the solubilization reaction. As shown in Figure 3 , using the task processing method based on intermolecular interaction of the embodiment of the present application, the target substructure can be accurately located using 5 iterations, so the dynamic recognition ability in the solubilization reaction of solute and solvent is strong.

[0133] Figure 4 A block diagram of a task processing apparatus based on intermolecular interaction according to an embodiment of the present application is shown.

[0134] As shown in Figure 4 , the task processing apparatus 400 based on intermolecular interaction includes an embedding module 410, an identification module 420, a determination module 430, a first processing module 440, and a second processing module 450.

[0135] The embedding module 410 is configured to perform embedding feature representation on the first node and the second node based on the connection edge formed between the first node in the first molecular graph and the second node in the second molecular graph, to obtain first node embedding features and second node embedding features, the first molecular graph indicating the chemical structure of the first molecule, and the second molecular graph indicating the chemical structure of the second molecule. The identification module 420 is configured to identify first action node features and second action node features of the nodes used to interact from the first node embedding features and the second node embedding features based on the association relationship between the first node and the second node. The determination module 430 is configured to determine target substructures in the chemical structures indicated by the first molecular graph and the second molecular graph based on the first action node features and the second action node features. The first processing module 440 is configured to perform regression prediction on the target substructures corresponding to the first molecular graph and the second molecular graph respectively when it is determined that the task matched with the first molecule and the second molecule is a regression task, to obtain a processing result of an index used to represent the interaction between the first molecule and the second molecule. The second processing module 450 is configured to classify whether the first molecule and the second molecule have an interaction based on the target substructures corresponding to the first molecular graph and the second molecular graph respectively when it is determined that the task matched with the first molecule and the second molecule is a classification task, to obtain a classification result.

[0136] According to an embodiment of the present disclosure, any of the modules of the embedding module 410, the identifying module 420, the determining module 430, the first processing module 440 and the second processing module 450 can be combined in one module, or any of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of the modules can be combined with at least part of the functions of other modules, and implemented in one module. According to an embodiment of the present disclosure, at least one of the embedding module 410, the identifying module 420, the determining module 430, the first processing module 440 and the second processing module 450 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging a circuit, etc. hardware or firmware, or implemented in any one of software, hardware and firmware or in a proper combination of any of them. Alternatively, at least one of the embedding module 410, the identifying module 420, the determining module 430, the first processing module 440 and the second processing module 450 can be at least partially implemented as a computer program module that can perform corresponding functions when the computer program module is run.

[0137] It should be noted that the task processing device part based on inter-molecular interaction in the embodiments of the present application corresponds to the task processing method part based on inter-molecular interaction in the embodiments of the present application, and the description of the task processing device part based on inter-molecular interaction is specifically referred to the task processing method part based on inter-molecular interaction, which will not be repeated here.

[0138] Figure 5 A block diagram of a training device of a task processing model based on inter-molecular interaction according to an embodiment of the present application is shown.

[0139] As shown in Figure 5 The training device 500 of the task processing model based on inter-molecular interaction includes a sample embedding module 510, a sample identifying module 520, a sample determining module 530, a sample first processing module 540, a sample second processing module 550, a first loss determining module 560, a second loss determining module 570 and a training module 580.

[0140] The sample embedding module 510 is configured to embed the sample first nodes and the sample second nodes based on the sample connection edges formed between the sample first nodes and the sample second nodes, to obtain sample first node embedding features and sample second node embedding features. The sample identification module 520 is configured to identify sample first action node features and sample second action node features of sample nodes that interact with each other from the sample first node embedding features and the sample second node embedding features based on the association relationship between the sample first nodes and the sample second nodes. The sample determination module 530 is configured to determine sample target substructures in the sample chemical structures indicated by the sample first molecular graph and the sample second molecular graph based on the sample first action node features and the sample second action node features. The sample first processing module 540 is configured to perform regression prediction on the target substructures corresponding to the sample first molecular graph and the sample second molecular graph to obtain a sample processing result for characterizing an index of interaction between the sample first molecule and the sample second molecule when it is determined that the task matched with the sample first molecule and the sample second molecule is a regression task. The sample second processing module 550 is configured to classify whether the sample first molecule and the sample second molecule have interaction based on the sample target substructures corresponding to the sample first molecular graph and the sample second molecular graph to obtain a sample classification result when it is determined that the task matched with the sample first molecule and the sample second molecule is a classification task. The first loss determination module 560 is configured to determine a structure loss value based on the sample first action node features and the sample second action node features. The second loss determination module 570 is configured to determine a result loss value based on the sample processing result and a real processing result, or based on the sample classification result and a real classification result. The training module 580 is configured to train the task processing model matched with the task based on the structure loss value and the result loss value to obtain a trained task processing model.

[0141] It should be noted that the training device part of the task processing model based on intermolecular interaction in the embodiments of the present application corresponds to the training method part of the task processing model based on intermolecular interaction in the embodiments of the present application. The description of the training device part of the task processing model based on intermolecular interaction is specifically referred to the training method part of the task processing model based on intermolecular interaction, which will not be repeated here.

[0142] Those skilled in the art will appreciate that features recited in the various embodiments of the present application can be combined and / or interchanged, even if this is not explicitly stated in the present application. In particular, features recited in the various embodiments of the present application can be combined and / or interchanged, without departing from the spirit and teachings of the present application. All such combinations and / or interchanges are within the scope of the present application.

[0143] The above describes embodiments of the present application. However, these embodiments are merely for illustrative purposes and are not intended to limit the scope of the present application. Although each embodiment is described above separately, this does not mean that the measures in the various embodiments cannot be advantageously used in combination. The scope of the present application is defined by the appended claims and their equivalents. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present application, and all such substitutions and modifications are intended to fall within the scope of the present application.

Claims

1. A task processing method based on intermolecular interactions, characterized in that, The method includes: Based on the connection edges formed between the first node in the first molecular graph and the second node in the second molecular graph, embedding feature representations are performed on the first node and the second node respectively to obtain the first node embedding feature and the second node embedding feature. The first molecular graph indicates the chemical structure of the first molecule, and the second molecular graph indicates the chemical structure of the second molecule. Based on the association between the first node and the second node, first and second action node features for nodes that interact are identified from the first node embedding features and the second node embedding features, respectively. Based on the first interaction node features and the second interaction node features, the target substructures that interact in the chemical structures indicated by the first molecular map and the second molecular map are determined. When the task matching the first molecule and the second molecule is determined to be a regression task, regression prediction is performed on the target substructures corresponding to the first molecule map and the second molecule map respectively to obtain the processing results of the index used to characterize the interaction between the first molecule and the second molecule. When determining that the task matching the first molecule and the second molecule is a classification task, the interaction between the first molecule and the second molecule is classified based on the target substructure corresponding to the first molecule map and the second molecule map, and the classification result is obtained.

2. The method according to claim 1, characterized in that, The method further includes: If the first molecule is determined to be a solvent molecule and the second molecule is determined to be a solute molecule, the task that matches the first molecule and the second molecule is determined to be the regression task. If it is determined that both the first molecule and the second molecule are drug molecules, the task that matches the first molecule and the second molecule is the classification task. The index of the interaction between the first molecule and the second molecule is used to characterize the thermodynamic properties of the solute molecule in the solvent molecule.

3. The method according to claim 1, characterized in that, The connecting edge includes a first directed connecting edge from the first node to the second node and a second directed connecting edge from the second node to the first node; The method involves embedding feature representations of the first node and the second node based on the connection edges formed between the first node in the first molecular graph and the second node in the second molecular graph, respectively, to obtain the first node embedding feature and the second node embedding feature, including: Determine the attribute characteristics of the first directed connection edge, the second directed connection edge, the first node, and the second node respectively; The embedding features of the first node are obtained by merging the attribute features of the second directed connection edge and the first node. The embedding feature of the second node is obtained by merging the attribute features of the first directed connection edge and the second node.

4. The method according to claim 3, characterized in that, The method of obtaining the first node embedding feature based on merging the attribute features of the second directed connection edge and the first node includes: The attribute features of the second node are fused with similarity features to obtain a first fused feature, wherein the similarity features are used to characterize the similarity features between the attribute features of the first node and the second node respectively. The first fused feature is processed sequentially through a fully connected layer and an activation layer to obtain the first node transformation feature; Based on the first parameter, the attribute features of the second directed connection edge and the first node transformation feature are merged to obtain the first merged feature. The first parameter is used to control the merging ratio of the attribute features of the second directed connection edge. The first merged feature is processed sequentially through a fully connected layer and an activation layer to obtain the first edge transformation feature; Based on the second parameter, the first edge transformation feature and the attribute feature of the first node are merged to obtain the first node embedding feature. The second parameter is used to control the fusion ratio of the attribute features of the first node.

5. The method according to claim 4, characterized in that, The step of merging the attribute features of the first directed connection edge and the second node to obtain the embedding feature of the second node includes: The attribute features of the first node are fused with the similarity features to obtain the second fused feature; The second fusion feature is processed sequentially through a fully connected layer and an activation layer to obtain the second node transformation feature; Based on the third parameter, the attribute features of the first directed connection edge and the transformation features of the second node are merged to obtain the second merged feature. The third parameter is used to control the merging ratio of the attribute features of the first directed connection edge. The second merged feature is processed sequentially through a fully connected layer and an activation layer to obtain the second edge transformation feature; Based on the fourth parameter, the second edge transformation feature and the attribute feature of the second node are merged to obtain the second node embedding feature. The fourth parameter is used to control the fusion ratio of the attribute features of the second node.

6. The method according to claim 1, characterized in that, The step of identifying first and second acting node features for nodes that interact, based on the association relationship between the first and second nodes, from the first node embedding features and the second node embedding features, respectively, includes: The first node embedding feature and the second node embedding feature are subjected to T rounds of recognition processing to obtain the Tth round of recognition results, where T represents the predetermined iteration round and T is a positive integer; Based on the Tth round of identification results, the first active node features and the second active node features are determined; The step of performing T rounds of recognition processing on the first node embedding features and the second node embedding features to obtain the Tth round of recognition results includes: Based on the parameter value of the t-th round and the importance of the second node in the t-th round, the second active node feature is identified from the embedded feature of the second node to obtain the first sub-identification result of the t-th round. The parameter value of the t-th round is less than or equal to the termination parameter value, or the parameter value of the t-th round is greater than or equal to the starting parameter value. The predetermined parameter is used to control the identification accuracy of the second active node feature during the change process from the starting parameter value to the termination parameter value. Based on the first sub-identification result of round t, the second sub-identification result of round t-1 is updated to obtain the second sub-identification result of round t; Based on the first sub-identification result of round t and the second sub-identification result of round t, the identification result of round t is obtained, where t is an integer greater than or equal to 1 and less than or equal to T.

7. The method according to claim 6, characterized in that, The method further includes: Based on the similarity between the second sub-identification result of the (t-1)th round and the embedded features of the second node, the similar features of the tth round are obtained; The weighted sum of the similarity features in round t and the embedding features of the first node is obtained to get the weighted features in round t. The weighted features of round t are transformed and linearly mapped to obtain the importance of the second node in round t.

8. A training method for a task processing model based on intermolecular interactions, characterized in that, The training method includes: Based on the sample connection edges formed between the first sample node in the first sample molecular graph and the second sample node in the second sample molecular graph, embedding feature representations are performed on the first sample node and the second sample node respectively to obtain the first sample node embedding feature and the second sample node embedding feature. The first sample molecular graph indicates the sample chemical structure of the first sample molecule, and the second sample molecular graph indicates the sample chemical structure of the second sample molecule. Based on the association between the first sample node and the second sample node, the first sample node feature and the second sample node feature for interacting sample nodes are identified from the first sample node embedding feature and the second sample node embedding feature, respectively. Based on the first and second interaction node features of the sample, the target substructures of the sample that interact in the sample chemical structure indicated by the first and second molecular maps of the sample are determined. When the task matching the first sample molecule and the second sample molecule is determined to be a regression task, regression prediction is performed on the target substructures corresponding to the first sample molecule map and the second sample molecule map respectively to obtain the sample processing results of the index used to characterize the interaction between the first sample molecule and the second sample molecule. When determining that the task matching the first sample molecule and the second sample molecule is a classification task, the interaction between the first sample molecule and the second sample molecule is classified based on the sample target substructure corresponding to the first sample molecule map and the second sample molecule map respectively, and the sample classification result is obtained. Based on the features of the first and second action nodes of the sample, the structural loss value is determined. Based on the sample processing results and the actual processing results, or based on the sample classification results and the actual classification results, determine the result loss value; Based on the structural loss value and the result loss value, a task processing model matching the task is trained to obtain the trained task processing model.

9. The training method according to claim 8, characterized in that, The determination of the structural loss value based on the features of the first and second action nodes of the sample includes: Based on the features of the first active node of the sample and the first molecular graph of the sample, the loss value of the first substructure is determined; Based on the features of the second action node of the sample and the second molecular graph of the sample, the loss value of the second substructure is determined; The structural loss value is obtained based on the first substructure loss value and the second substructure loss value; The step of determining the first substructure loss value based on the first action node features of the sample and the first molecular graph of the sample includes: Based on the characteristics of the first functional node of the sample, the importance of the first node of the sample is determined; Based on the importance of the first molecular graph of the sample and the first node of the sample, a sparsity constraint is applied to the first node of the sample to obtain a first constraint value; Based on the importance of the first node of the sample, a peak value constraint is applied to the first node of the sample to obtain a second constraint value; Based on the importance of the first node of the sample and the sample edges in the first molecular graph of the sample, a cluster constraint is applied to the first node of the sample to obtain a third constraint value. The first substructure loss value is obtained based on the first constraint value, the second constraint value, and the third constraint value.

10. A task processing device based on intermolecular interactions, characterized in that, The device includes: An embedding module is used to perform embedding feature representation on the first node and the second node respectively based on the connection edge formed between the first node in the first molecular graph and the second node in the second molecular graph, to obtain the first node embedding feature and the second node embedding feature, wherein the first molecular graph indicates the chemical structure of the first molecule and the second molecular graph indicates the chemical structure of the second molecule; The identification module is used to identify, based on the association relationship between the first node and the second node, a first acting node feature and a second acting node feature for nodes that interact, respectively, from the first node embedding feature and the second node embedding feature. The determination module is used to determine the target substructures that interact in the chemical structures indicated by the first molecular map and the second molecular map, based on the first interaction node features and the second interaction node features. The first processing module is used to perform regression prediction on the target substructures corresponding to the first molecular map and the second molecular map respectively when the task matching the first molecule and the second molecule is determined to be a regression task, so as to obtain the processing result of the index used to characterize the interaction between the first molecule and the second molecule. The second processing module is used to classify whether there is an interaction between the first molecule and the second molecule when the task matching the first molecule and the second molecule is determined to be a classification task, based on the target substructures corresponding to the first molecule map and the second molecule map respectively, and to obtain the classification result.

Citation Information

Cited By

  • Solute-solvent interaction prediction method based on conditional factor subgraph recognition

    CN122177270A