Multi-expert molecular attribute prediction method based on substructure information bottleneck

By splitting the molecular graph into skeleton substructure and functional group substructure, and using the substructure information bottleneck module and gating mechanism to dynamically allocate expert models, the problems of redundant information retention and lack of task-related information in molecular property prediction in existing technologies are solved, and more efficient molecular property prediction is achieved.

CN120708757APending Publication Date: 2025-09-26ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510815020.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively utilizing molecular substructure information, are unable to effectively remove redundant information and retain task-related information, resulting in limited accuracy and efficiency in molecular property prediction.

Method used

By splitting the molecular graph into skeleton substructure and functional group substructure, encoding it using graph neural network, constructing a substructure information bottleneck module to optimize the substructure representation, and dynamically assigning it to the expert model in combination with the gating mechanism, a two-layer optimization strategy is adopted to train the model parameters.

Benefits of technology

The accuracy and robustness of molecular property predictions have been significantly improved, and the model's adaptability to complex molecules has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708757A_ABST
    Figure CN120708757A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-expert molecular attribute prediction method based on substructure information bottleneck, and aims to solve the problems that an existing method cannot fully excavate molecular substructure information and lacks an effective feature filtering mechanism. Substructure representation is coded through a graph neural network; constructing a substructure information bottleneck module, and filtering redundant features by taking maximization of mutual information of a substructure and a target attribute and minimization of mutual information of the substructure and a complete molecular graph as an optimization target; dynamically distributing the substructures to an expert model based on a skeleton and a functional group in combination with a gating mechanism, and calculating a distribution probability by utilizing a trainable clustering center; updating the model by adopting a double-layer strategy of optimizing discriminator parameters through internal circulation and optimizing main model parameters through external circulation; and finally, weighting and fusing expert output to realize attribute prediction. The method significantly improves the accuracy and robustness of molecular attribute prediction, and is suitable for the fields of drug research and development and material design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of molecular property prediction and graph neural network technology, and in particular relates to a multi-expert molecular property prediction method based on substructure information bottleneck. Background Art

[0002] Molecular property prediction is a fundamental task in cheminformatics, widely used in fields such as drug discovery, materials science, and environmental chemistry. Accurately predicting properties such as solubility, reactivity, and bioavailability is crucial for identifying promising compounds and optimizing their design. With the rise of deep learning, data-driven representation learning methods have brought significant progress to molecular property prediction. Early sequence-based methods, such as simplified molecular linear input systems, provided linear symbolic representations of molecules. While these methods achieved initial success, they struggled to capture the necessary spatial and structural information. In contrast, graph-based methods represent molecules as graphs, where atoms and chemical bonds correspond to nodes and edges, respectively, enabling more accurate capture of complex molecular patterns. Graph neural networks use message passing mechanisms to iteratively update node representations based on neighboring node features, significantly improving prediction accuracy. However, due to their data-driven nature, these methods still have limitations when dealing with heterogeneous molecular structures, as a single graph neural network model often struggles to effectively capture complex molecular information.

[0003] To address the heterogeneity of molecular data, multi-expert architectures have attracted widespread attention. They employ a divide-and-conquer strategy, utilizing multiple specialized models or "experts," each responsible for a different subdomain of the input space. This improves learning efficiency and model specialization, and by assigning data with similar topological or structural features to different experts, multi-expert architectures can effectively handle diverse molecular structures. Existing multi-expert molecular property prediction methods first construct a hybrid expert model, encode the task information text and text context through a cross-modal projector, and input the result into the routing of the hybrid expert model. Molecular encoders are then trained to handle different tasks, and the trained molecular encoders are integrated into a unified 3D molecular encoder. Multiple molecular encoders are then combined to improve comprehensive understanding of molecular properties. However, existing technologies fail to address the impact of finer-grained information, such as molecular substructure, and irrelevant information in the molecule on prediction results. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a multi-expert molecular property prediction method based on substructure information bottleneck to solve the problems existing in the above-mentioned prior art.

[0005] In a first aspect, to achieve the above-mentioned object, the present invention provides a multi-expert molecular property prediction method based on substructure information bottleneck, comprising the following steps: Representing molecules as a molecular graph including an adjacency matrix, a node feature matrix, and an edge feature matrix; Decomposing the molecular graph to obtain a skeleton substructure and a functional group substructure; Using a graph neural network to encode the skeleton substructure and the functional group substructure respectively to obtain a substructure representation; Construct a substructure information bottleneck module to optimize substructure representation by maximizing the mutual information between substructures and target attributes and minimizing the mutual information between substructures and the complete molecular graph; Constructing a multi-expert module to dynamically assign substructures to backbone- and functional group-based expert models in conjunction with a gating mechanism that calculates assignment probabilities via trainable cluster centers; The model is updated using a two-layer optimization strategy with an inner loop optimizing the discriminator parameters and an outer loop optimizing the main model parameters. Molecular property predictions are calculated based on expert output and assignment probabilities.

[0006] Optionally, the process of splitting the molecular graph includes: Decompose the molecular graph into backbone substructure and functional group substructure; The message passing mechanism of graph neural networks is used to iteratively update node embeddings, and substructure representations are obtained through aggregation functions.

[0007] Optionally, the process of constructing the substructure information bottleneck module includes: The optimization goal is to maximize the mutual information between the substructure and the target attribute and minimize the mutual information between the substructure and the complete molecular graph; Use a learnable discriminator to estimate the mutual information loss between the substructure and the complete molecular graph; Combine the downstream task losses to construct the information bottleneck total loss.

[0008] Optionally, the process of the gating mechanism includes: Linearly map the substructure representation to a topological representation; Calculate the probability of molecules being assigned to each cluster based on the trainable cluster centers; Gumbel-Softmax distribution is used to generate the assignment probability with temperature parameter.

[0009] Optionally, the process of constructing the multi-expert module includes: Set up two types of expert models based on skeleton and functional groups; Optimize clustering loss by minimizing the difference between the current cluster distribution and the target distribution through KL divergence; The final prediction value is obtained by weighted fusion of all expert outputs.

[0010] Optionally, the process of the two-layer optimization strategy includes: The inner loop updates the discriminator parameters to optimize the mutual information loss; The outer loop updates the main model parameters to optimize the sum of the information bottleneck loss and the clustering loss.

[0011] In a second aspect, the present invention further provides a multi-expert molecular property prediction system based on substructure information bottleneck, which is used to implement a multi-expert molecular property prediction method based on substructure information bottleneck, and the system comprises: Molecular structure decomposition module, used to represent molecules as molecular graphs including adjacency matrix, node feature matrix and edge feature matrix, and split them into skeleton substructure and functional group substructure; The substructure encoding module is used to encode the skeleton substructure and functional group substructure using graph neural networks to generate substructure representations; An information bottleneck optimization module for optimizing substructure representations by maximizing the mutual information between substructures and target properties and minimizing the mutual information between substructures and the complete molecular graph; a multi-expert routing module for dynamically assigning substructures to backbone- and functional group-based expert models in conjunction with a gating mechanism that calculates assignment probabilities via trainable cluster centers; Parameter optimization module, which is used to implement a two-level optimization strategy of optimizing the discriminator parameters in an inner loop and the main model parameters in an outer loop; The property prediction module is used to calculate the molecular property prediction results based on the expert model output and assignment probability.

[0012] Optionally, the molecular structure decomposition module includes: Subgraph decomposition unit, used to decompose the molecular graph into skeleton substructure and functional group substructure; The embedding generation unit is used to iteratively update node embeddings through the message passing mechanism of the graph neural network and generate substructure representations using an aggregation function.

[0013] In a third aspect, the present invention further provides a computer terminal device, comprising: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a multi-expert molecular property prediction method based on a substructure information bottleneck.

[0014] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a multi-expert molecular property prediction method based on a substructure information bottleneck.

[0015] Compared with the prior art, the present invention has the following advantages and technical effects: The present invention provides a multi-expert molecular property prediction method based on substructure information bottleneck. The present invention fully exploits molecular hierarchical features through decomposition characterization of skeleton substructure and functional group substructure; combines the information bottleneck principle to filter redundant information and force the model to retain key substructure features related to the target properties; utilizes a gating mechanism to achieve dynamic routing of substructures to dedicated experts, ensuring that heterogeneous molecules are processed by the most suitable model; and uses a two-layer optimization strategy to collaboratively train the discriminator and main model parameters, significantly improving the model's adaptability to complex molecular topologies, ultimately achieving overall optimization of the accuracy and robustness of molecular property prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings: Figure 1 This is an overall flow chart of an embodiment of the present invention.

[0017] Figure 2 This is an overall framework diagram of an embodiment of the present invention.

[0018] Figure 3 Schematic diagram of a substructure information bottleneck module according to an embodiment of the present invention.

[0019] Figure 4 This is a schematic diagram of a multi-expert module combined with a gating mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION

[0020] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0021] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0022] Example 1 like Figure 1 As shown, this embodiment provides a multi-expert molecular property prediction method based on substructure information bottleneck, including: Representing molecules as a molecular graph including an adjacency matrix, a node feature matrix, and an edge feature matrix; Decomposing the molecular graph to obtain a skeleton substructure and a functional group substructure; Using a graph neural network to encode the skeleton substructure and the functional group substructure respectively to obtain a substructure representation; Construct a substructure information bottleneck module to optimize substructure representation by maximizing the mutual information between substructures and target attributes and minimizing the mutual information between substructures and the complete molecular graph; Constructing a multi-expert module to dynamically assign substructures to backbone- and functional group-based expert models in conjunction with a gating mechanism that calculates assignment probabilities via trainable cluster centers; The model is updated using a two-layer optimization strategy with an inner loop optimizing the discriminator parameters and an outer loop optimizing the main model parameters. Molecular property predictions are calculated based on expert output and assignment probabilities.

[0023] Specifically, S1: perform substructure decomposition on the molecular graph to obtain the skeleton substructure and functional group substructure, and use graph neural network to obtain the substructure representation; S2: Build a substructure information bottleneck module to optimize the molecular substructure representation; S3: Construct a multi-expert module and combine it with a gating mechanism to enhance substructure representation; S4: Optimize the model using a two-level optimization approach and update the model parameters to obtain the optimal model for molecular property prediction; S5: Based on the optimal model obtained in step S4, the model performance is tested using the evaluation indicators corresponding to the downstream tasks.

[0024] As an implementation method in this embodiment, the process of splitting the molecular graph includes: Decompose the molecular graph into backbone substructure and functional group substructure; The message passing mechanism of graph neural networks is used to iteratively update node embeddings, and substructure representations are obtained through aggregation functions.

[0025] Specifically, in step S1, the molecular graph is split to obtain the skeleton substructure and the functional group substructure, specifically as follows: atoms are used as nodes and chemical bonds as edges to form a molecular graph ,in Represented as an adjacency matrix, represents the node feature matrix, Represents the edge feature matrix; each molecular graph Split into skeleton substructure and functional group substructures ;Use graph neural network to encode the obtained substructure and obtain the substructure representation and .

[0026] More specifically, a molecule is represented as a molecular graph. ,in Represented as an adjacency matrix, represents the node feature matrix, Represents the edge feature matrix; then, based on chemical rules, each molecular graph Decomposition into skeleton substructures and functional group substructures ;Use graph neural network to obtain initial node embedding , through the message passing mechanism, further representation learning is performed to update node embedding: in and is a node and In the The embedding of the iterations, is a node Neighbor node set, is the edge weight, Represents a function that combines its own node information and neighbor aggregation information, The representation propagation function is used to obtain the embedded representation from the neighbor nodes; After layer iteration, we get , and then obtain the graph-level representation by node aggregation : The read function You can use pooling functions, such as sum pooling, average pooling, etc.; then perform the following operations on the skeleton substructures: and functional group substructures Encode and get the substructure representation and .

[0027] like Figure 3 As shown in FIG. 1 , as an implementation method of this embodiment, the process of constructing the substructure information bottleneck module includes: The optimization goal is to maximize the mutual information between the substructure and the target attribute and minimize the mutual information between the substructure and the complete molecular graph; Use a learnable discriminator to estimate the mutual information loss between the substructure and the complete molecular graph; Combine the downstream task losses to construct the information bottleneck total loss.

[0028] Specifically, in step S2, a substructure information bottleneck module is constructed to optimize the molecular substructure representation, as follows: S2.1: To reduce the impact of noise and irrelevant or redundant molecular features on model performance, the information bottleneck theory is introduced. It aims to maximize the information in the substructure that is relevant to the target task while minimizing the information of irrelevant or redundant molecular features. It enables the model to focus on the core substructure that is critical for property prediction. The optimization objectives of the substructure representation are as follows: in It is a substructure and target attributes The mutual information between is the complete molecular graph and its substructure Mutual information between S2.2: Based on step S1, obtain the skeleton substructure and functional group substructures , the optimization objective of the substructure representation can be redefined as: in and Skeleton substructure and functional group substructures With target attributes The mutual information between and The complete molecular graph With its skeleton substructure and functional group substructures Mutual information between S2.3: To reduce the complete molecular graph The irrelevant information between its substructures, firstly, the present invention uses the random variable mutual information evaluation method to estimate ,Right now: in represents a learnable discriminator; then, define Loss to approximate , ensuring that the substructure representation contains only the most important information for attribute prediction without redundant information; at the same time, by maximizing , so that the substructure representation retains the most task-related information, namely: in is the true posterior probability Variational approximation can further obtain downstream task losses , to guide the model to focus on Task-related information in ; For classification tasks, is the binary cross entropy loss, for regression tasks, is the mean square error loss; the total loss of the final substructure information bottleneck module is: in A hyperparameter representing the weight of the mutual information loss.

[0029] As an implementation method in this embodiment, the process of the gating mechanism includes: Linearly map the substructure representation to a topological representation; Calculate the probability of molecules being assigned to each cluster based on the trainable cluster centers; Gumbel-Softmax distribution is used to generate the assignment probability with temperature parameter.

[0030] As an implementation method in this embodiment, the process of constructing the multi-expert module includes: Set up two types of expert models based on skeleton and functional groups; Optimize clustering loss by minimizing the difference between the current cluster distribution and the target distribution through KL divergence; The final prediction value is obtained by weighted fusion of all expert outputs.

[0031] Specifically, such as Figure 4 As shown, in step S3, a multi-expert module is constructed and combined with a gating mechanism to enhance substructure representation, as follows: S3.1: To coordinate multiple experts, a gating mechanism is introduced to cluster molecules and assign each molecule to each expert based on its similarity to the corresponding molecular cluster; the gating mechanism uses a linear mapping layer to embed the substructure graph into a representation and Convert to topological representation and ,Right now: in Represents a linear mapping layer; S3.2: Construct skeleton-based and functional group-based experts respectively; for each expert there is experts, including Indicates the number of tasks, Indicates the number of cluster categories; Experts in topological clustering The generated graph embedding is: in express Substructure graph embedding (skeleton substructure graph embedding Or functional group substructure diagram embedding ), It is The learnable weight matrix of the experts; S3.3: Introduction trainable cluster centers, calculate the probability that each molecule belongs to each cluster, so the The molecules are assigned to The probability of a topological cluster is: in, and The two categories of experts are Cluster centers, Indicates the To help the expert model learn topologically relevant patterns in the early stages of training, the present invention adopts a random selection strategy based on the Gumbel-Softmax distribution. While maintaining differentiability, the present invention approximately samples from the cluster assignment distribution and assigns probabilities. Can be calculated as in and is the Gumbel noise term, is the temperature parameter, which decreases from the initial value to the set minimum value during the training phase; S3.4: In order to optimize the clustering process, a target distribution is introduced , to guide the current cluster assignment probability , the target distribution is obtained by the current probability Squaring is performed to enhance high confidence assignments; by minimizing the current distribution With target distribution The KL divergence between is used to define the clustering loss: in is the batch size, the total clustering loss It can be obtained by the following formula: S3.5: Based on the output of multiple experts and the assignment probability from the gating mechanism, the final prediction value of each molecule is calculated: in is the normalization function, Represents a splicing operation.

[0032] As an implementation method in this embodiment, the process of the two-layer optimization strategy includes: The inner loop updates the discriminator parameters to optimize the mutual information loss; The outer loop updates the main model parameters to optimize the sum of the information bottleneck loss and the clustering loss.

[0033] Specifically, in step S4, a two-layer optimization method is used to optimize the model and update the model parameters to obtain the optimal model for molecular property prediction, as follows: S4.1: Based on the losses in steps S2 and S3, the final loss of the model is obtained, namely: in is a hyperparameter of the mutual information loss weight, is a hyperparameter used to control the quality of topological clustering; S4.2: Design inner loop optimization and outer loop optimization. The present invention first performs inner loop optimization, and its optimization goal is to optimize the model parameters. ,Right now: in is the learning rate, It represents the gradient of the model parameters; then the model is optimized in an outer loop, the goal of which is to optimize the model parameters ,Right now: in It represents the gradient of the model parameters; the model parameters are updated to obtain the optimal model for molecular property prediction.

[0034] In step S5, the optimal model obtained in step S4 is tested for model performance using the evaluation index corresponding to the downstream task, specifically as follows: Based on the optimal model obtained in step S4, for each molecule, the first The output of an expert and its corresponding distribution probability : in is the normalization function, Represents the splicing operation; further obtain the final prediction score : Then, the evaluation indicators corresponding to the downstream tasks are used. For downstream classification tasks, the ROC-AUC performance evaluation indicator is used to evaluate the model performance; for downstream regression tasks, the root mean square error performance evaluation indicator is used to evaluate the model performance.

[0035] The present invention can be applied to molecular property prediction tasks in real scenarios and can greatly improve the accuracy of prediction results.

[0036] In summary, the present invention proposes a multi-expert molecular property prediction method based on the substructure information bottleneck, which makes up for the defect that the existing methods fail to effectively utilize molecular substructure information and overcomes the problem that the existing methods cannot remove redundant information and retain task-related information. This method significantly improves the accuracy and efficiency of the molecular property prediction model.

[0037] Based on this, an embodiment of the present invention provides a multi-expert molecular property prediction method based on substructure information bottleneck: (1) Existing technical solutions are usually based on the overall molecular graph or atomic level information for representation, and fail to fully explore the structural information in the molecular substructure (such as molecular functional groups, etc.) level, resulting in limited model representation capabilities. The present invention proposes a multi-expert mechanism based on the substructure information bottleneck. By splitting the molecular graph, the molecular substructure representation is obtained, and the substructure representation is optimized by combining the substructure information bottleneck module. Then, the multi-expert network is dynamically allocated at the substructure level for representation learning, thereby enhancing the model's ability to represent molecules.

[0038] (2) The present invention introduces the substructure information bottleneck theory and constructs a substructure information bottleneck module based on the theory, which retains the molecular features related to the task and filters out redundant or irrelevant information. This solves the problem that the existing technical solutions lack an effective mechanism to filter out irrelevant information in the molecular graph and retain information related to the task, and effectively improves the performance of the model in the molecular property prediction task.

[0039] Example 2 In this embodiment, a computer terminal device is provided, including: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the methods in the above embodiments.

[0040] In this embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the method in the above embodiment is implemented.

[0041] In this embodiment, an electronic device is further provided, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the method in the above embodiment.

[0042] The above program can be executed in a processor or stored in a memory (or computer-readable medium). Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0043] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more blocks can be implemented by different modules corresponding to different steps.

[0044] like Figure 2 As shown, this embodiment provides such a device or system. The system is called a multi-expert molecular property prediction system based on substructure information bottleneck, including: Molecular structure decomposition module, used to represent molecules as molecular graphs including adjacency matrix, node feature matrix and edge feature matrix, and split them into skeleton substructure and functional group substructure; The substructure encoding module is used to encode the skeleton substructure and functional group substructure using graph neural networks to generate substructure representations; An information bottleneck optimization module for optimizing substructure representations by maximizing the mutual information between substructures and target properties and minimizing the mutual information between substructures and the complete molecular graph; a multi-expert routing module for dynamically assigning substructures to backbone- and functional group-based expert models in conjunction with a gating mechanism that calculates assignment probabilities via trainable cluster centers; Parameter optimization module, which is used to implement a two-level optimization strategy of optimizing the discriminator parameters in an inner loop and the main model parameters in an outer loop; The property prediction module is used to calculate the molecular property prediction results based on the expert model output and assignment probability.

[0045] As an implementation in this embodiment, the molecular structure decomposition module includes: Subgraph decomposition unit, used to decompose the molecular graph into skeleton substructure and functional group substructure; The embedding generation unit is used to iteratively update node embeddings through the message passing mechanism of the graph neural network and generate substructure representations using an aggregation function.

[0046] As an implementation method of this embodiment, the information bottleneck optimization module includes: A target optimization unit is used to maximize the mutual information between the substructure and the target attribute and minimize the mutual information between the substructure and the complete molecular graph; The discriminator unit is used to estimate the mutual information loss between the substructure and the complete molecular graph through a learnable discriminator; Loss integration unit, used to combine downstream task losses to construct the information bottleneck total loss.

[0047] As an implementation method of this embodiment, the multi-expert routing module includes: A topological mapping unit, used for linearly mapping the substructure representation into a topological representation; A cluster assignment unit, configured to calculate the probability of assigning a molecule to each cluster based on the trainable cluster centers; The probability generation unit is used to generate the allocation probability with the temperature parameter using the Gumbel-Softmax distribution.

[0048] As an implementation manner in this embodiment, the multi-expert routing module further includes: Expert configuration unit, used to set two types of expert models based on skeleton and functional groups; Cluster optimization unit, used to optimize clustering loss by minimizing the difference between the current cluster distribution and the target distribution through KL divergence; The result fusion unit is used to weight the outputs of all expert models to generate the final prediction value.

[0049] As an implementation method in this embodiment, the parameter optimization module includes: The inner loop unit is used to update the discriminator parameters to optimize the mutual information loss; The outer loop unit is used to update the main model parameters to optimize the sum of the information bottleneck loss and the clustering loss.

[0050] The system or device is used to implement the functions of the method in the above-mentioned embodiment. Each module in the system or device corresponds to each step in the method, which has been explained in the method and will not be repeated here.

[0051] Through the above implementation, the problem of multi-expert molecular property prediction based on the substructure information bottleneck in the related art is solved, thereby ensuring that the problems existing in the prior art are solved.

[0052] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multi-expert molecular property prediction method based on substructure information bottleneck, characterized by: The following steps are involved: Representing molecules as a molecular graph including an adjacency matrix, a node feature matrix, and an edge feature matrix; Decomposing the molecular graph to obtain a skeleton substructure and a functional group substructure; Using a graph neural network to encode the skeleton substructure and the functional group substructure respectively to obtain a substructure representation; Construct a substructure information bottleneck module to optimize substructure representation by maximizing the mutual information between substructures and target attributes and minimizing the mutual information between substructures and the complete molecular graph; Constructing a multi-expert module to dynamically assign substructures to backbone- and functional group-based expert models in conjunction with a gating mechanism that calculates assignment probabilities via trainable cluster centers; The model is updated using a two-layer optimization strategy with an inner loop optimizing the discriminator parameters and an outer loop optimizing the main model parameters. Molecular property predictions are calculated based on expert output and assignment probabilities.

2. The method according to claim 1, characterized in that The process of splitting the molecular graph includes: Decompose the molecular graph into backbone substructure and functional group substructure; The message passing mechanism of graph neural networks is used to iteratively update node embeddings, and substructure representations are obtained through aggregation functions.

3. The method according to claim 1, characterized in that The process of constructing the substructure information bottleneck module includes: The optimization goal is to maximize the mutual information between the substructure and the target attribute and minimize the mutual information between the substructure and the complete molecular graph; Use a learnable discriminator to estimate the mutual information loss between the substructure and the complete molecular graph; The information bottleneck total loss is constructed by combining the downstream task losses.

4. The method according to claim 1, wherein The process of the gating mechanism includes: Linearly map the substructure representation to a topological representation; Calculate the probability of molecules being assigned to each cluster based on the trainable cluster centers; Gumbel-Softmax distribution is used to generate the assignment probability with temperature parameter.

5. The method according to claim 1, wherein The process of constructing the multi-expert module includes: Set up two types of expert models based on skeleton and functional groups; Optimize clustering loss by minimizing the difference between the current cluster distribution and the target distribution through KL divergence; The final prediction value is obtained by weighted fusion of all expert outputs.

6. The method according to claim 1, characterized in that The process of the two-layer optimization strategy includes: The inner loop updates the discriminator parameters to optimize the mutual information loss; The outer loop updates the main model parameters to optimize the sum of the information bottleneck loss and the clustering loss.

7. A multi-expert molecular property prediction system based on substructure information bottleneck, characterized by: The system comprises: Molecular structure decomposition module, used to represent molecules as molecular graphs including adjacency matrix, node feature matrix and edge feature matrix, and split them into skeleton substructure and functional group substructure; The substructure encoding module is used to encode the skeleton substructure and functional group substructure using graph neural networks to generate substructure representations; An information bottleneck optimization module for optimizing substructure representations by maximizing the mutual information between substructures and target properties and minimizing the mutual information between substructures and the complete molecular graph; a multi-expert routing module for dynamically assigning substructures to backbone- and functional group-based expert models in conjunction with a gating mechanism that calculates assignment probabilities via trainable cluster centers; Parameter optimization module, which is used to implement a two-level optimization strategy of optimizing the discriminator parameters in an inner loop and the main model parameters in an outer loop; The property prediction module is used to calculate the molecular property prediction results based on the expert model output and assignment probability.

8. The system according to claim 7, characterized in that The molecular structure decomposition module includes: Subgraph decomposition unit, used to decompose the molecular graph into skeleton substructure and functional group substructure; The embedding generation unit is used to iteratively update node embeddings through the message passing mechanism of the graph neural network and generate substructure representations using an aggregation function.

9. A computer terminal device, characterized in that: include: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-expert molecular property prediction method based on substructure information bottleneck according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-expert molecular property prediction method based on substructure information bottleneck according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Node separation learning method and system based on discrimination difficulty perception

    CN120911511A

  • Multi-feature fusion oral bioavailability prediction method, device, equipment and medium

    CN121885236A

  • Multi-feature fusion oral bioavailability prediction method, device, equipment and medium

    CN121885236B