Multi-expert molecular attribute prediction method and system based on adaptive substructure perception
By decomposing molecules into substructures and introducing marginal ternary loss and dynamic routing mechanisms, we adaptively handle multi-expert molecular property predictions, solving the problem of insufficient quantification of substructure contribution differences in existing technologies and improving the prediction accuracy and interpretability of the model.
Patent Information
- Application Number
- CN202510815021.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing multi-expert molecular property prediction methods fail to effectively quantify the differences in the contributions of different substructures to the prediction results, resulting in confusion between key features and noise features, affecting the model's generalization ability and practical application value.
By decomposing molecules into a set of substructures, the BRICS splitting method and graph neural network are used to extract node representations, attribution analysis is combined to identify positive and negative substructures, and marginal ternary loss and downstream task loss are introduced to optimize the substructure representation. A multi-expert architecture is constructed, and a dynamic routing mechanism is designed to adaptively assign expert fusion weights to finally generate molecular property prediction values.
It significantly improves the accuracy and generalization ability of molecular property prediction, reduces the impact of noise characteristics, enhances the model's ability to analyze complex molecular topologies, alleviates data imbalance problems, and ensures the stability and interpretability of the model in low-quality data scenarios.
Smart Images

Figure CN120708758A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of molecular property prediction and graph neural network, and in particular relates to a multi-expert molecular property prediction method and system based on adaptive substructure perception. Background Art
[0002] In recent years, with the rapid development of artificial intelligence (AI) technology, molecular property prediction has gradually become a research hotspot in fields such as drug discovery. This technology can significantly reduce the cost and risk of traditional wet-lab drug development. Among various molecular property prediction technologies, deep learning methods based on graph neural networks have attracted much attention due to their ability to effectively perceive molecular topology. These methods primarily model molecules as molecular graphs, transforming the property prediction task into a graph classification or regression problem. However, due to their data-driven nature, their reliance on high-quality labeled data and data imbalance significantly restrict the model's generalization ability, resulting in a lack of practical application value.
[0003] To address the above challenges, multi-expert model technology has been introduced into the field of molecular property prediction. Its core idea is to improve overall performance by integrating multiple specialized sub-models. Existing multi-expert molecular property prediction methods, such as the multimodal biomedical data processing methods disclosed in the prior art, first obtain multimodal data including corresponding molecular structures, knowledge graphs, and texts, and then use pre-built multi-dimensional molecular encoding models, knowledge graph encoding models, and text encoding models to encode the data to obtain feature representations. Finally, a multi-expert model is used to mix different molecular feature representations to adapt to different downstream tasks such as molecular property prediction. However, this existing technology fails to effectively quantify the differences in the contributions of different substructures to the prediction results, resulting in the problem of confusion between key features and noise features. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a multi-expert molecular property prediction method and system based on adaptive substructure perception to solve the problems existing in the above-mentioned prior art.
[0005] In a first aspect, to achieve the above-mentioned object, the present invention provides a multi-expert molecular property prediction method based on adaptive substructure perception, comprising the following steps: decomposing an input molecule into a set of substructures, and identifying positive and negative substructures of the molecule through attribution analysis; For the positive substructures and negative substructures, marginal ternary loss is introduced together with downstream task loss to jointly optimize the substructure recognition quality and obtain substructure representation; Based on the substructure representation, a multi-expert architecture is constructed to generate substructure embedding representations of different experts; Embedding the substructure, designing a dynamic routing mechanism, adaptively assigning expert fusion weights, and generating molecular logits based on different experts; Integrate the logit values of the different experts to obtain the predicted values of molecular properties, and train the optimization model by calculating the loss; Based on the optimized model, downstream task evaluation indicators are used to test the model performance.
[0006] Optionally, the process of decomposing the input molecule into a set of substructures includes: The BRICS splitting method is used to split the input molecule into a set of substructures; Extract molecular node representations based on graph neural networks, selectively mask molecular substructures through substructure mask vectors, and generate substructure embeddings; Attribution analysis was used to compare the prediction differences before and after masking, quantify the contribution of substructures to molecular properties, and select the top ψ substructures with the highest contribution values as positive substructures, and the top ψ substructures with the lowest contribution values as negative substructures.
[0007] Optionally, the process of jointly optimizing the substructure recognition quality includes: Taking the original molecular graph, the positive substructure, and the negative substructure as three views, a marginal triplet loss is defined to force the feature similarity between the positive substructure and the molecular graph to be higher than the feature similarity between the negative substructure and the molecular graph; The marginal triplet loss is weightedly summed with the downstream task loss to optimize substructure identification and representation.
[0008] Optionally, the process of constructing a multi-expert architecture includes: Multiple expert models are assigned to positive and negative substructures respectively. Each expert performs a linear transformation on the substructure representation through a weight matrix to generate an expert-specific substructure embedding representation. The number of experts is determined by the number of tasks and preset parameters.
[0009] Optionally, the process of the dynamic routing mechanism includes: For the embedded representation of the active substructure, the routing score is calculated through the parameter matrix and the expert weight is assigned; For the embedded representation of the negative substructure, the positive substructure and the negative substructure are concatenated and input into the router to calculate the routing score; The expert embedding representations are weighted and fused according to the routing scores to generate positive logits and negative logits.
[0010] Optionally, the process of calculating the loss training optimization model includes: The importance of experts in a batch is calculated based on the expert routing score, and the balance of expert division of labor is measured by the coefficient of variation to generate the importance loss. The importance loss is weightedly summed with the downstream task loss and the marginal triplet loss to optimize the model parameters.
[0011] In a second aspect, the present invention further provides a multi-expert molecular property prediction system based on adaptive substructure perception, which is used to implement a multi-expert molecular property prediction method based on adaptive substructure perception, and the system comprises: a substructure decomposition and attribution module for decomposing an input molecule into a set of substructures and identifying positive and negative substructures of the molecule through attribution analysis; A joint optimization module is used to introduce a marginal ternary loss and a downstream task loss to the positive substructure and the negative substructure to jointly optimize the substructure recognition quality and obtain a substructure representation; A multi-expert architecture building module is used to build a multi-expert architecture based on the substructure representation and generate substructure embedding representations of different experts; A dynamic routing assignment module is used to design a dynamic routing mechanism for the substructure embedding representation, adaptively assign expert fusion weights, and generate molecular logits based on different experts; A prediction integration and optimization module is used to integrate the logit values of the different experts to obtain the molecular property prediction value and train the optimization model by calculating the loss; The performance testing module is used to test the model performance based on the optimized model using downstream task evaluation indicators.
[0012] Optionally, the substructure decomposition and attribution module includes: BRICS splitting unit, used to split the input molecule into a set of substructures using the BRICS method; The substructure embedding generation unit is used to extract molecular node representations based on graph neural networks and selectively mask molecular substructures through substructure mask vectors to generate substructure embeddings; The substructure contribution quantification unit is used to compare the prediction differences before and after masking through attribution analysis, quantify the contribution of substructures to molecular properties, and select the top ψ substructures with the highest contribution values as positive substructures and the top ψ substructures with the lowest contribution values as negative substructures.
[0013] In a third aspect, the present invention further provides a computer terminal device, comprising: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a multi-expert molecular property prediction method based on adaptive substructure perception.
[0014] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a multi-expert molecular property prediction method based on adaptive substructure perception.
[0015] Compared with the prior art, the present invention has the following advantages and technical effects: (1) Existing technical solutions often ignore the different contributions of different substructures to molecular properties and treat them uniformly. The present invention can dynamically adapt to different molecular structures, avoiding confusion when a single model is used to process diverse and complex molecular features.
[0016] (2) The present invention is universal and can be used for all deep neural network models.
[0017] (3) The present invention is based on BRICS decomposition and substructure perception technology to identify positive and negative substructures, and uses a dynamic routing mechanism based on a multi-expert architecture to assign specialized experts to different molecular substructures, thereby effectively processing and integrating positive and negative substructure information. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings: Figure 1 This is an overall flow chart of an embodiment of the present invention.
[0019] Figure 2 This is an overall framework diagram of an embodiment of the present invention.
[0020] Figure 3 A schematic diagram of substructure identification and optimization according to an embodiment of the present invention; Figure 4 Schematic diagram of multi-expert routing mechanism training and optimization according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0022] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0023] This paper proposes a multi-expert molecular property prediction method and system based on adaptive substructure perception. Through splitting and attribution analysis, molecular substructures are divided into positive and negative substructures. A dynamic routing mechanism, based on a multi-expert architecture, is designed to assign different molecular substructures to expert models, enabling efficient processing and integration of positive and negative substructures.
[0024] Example 1 like Figure 1 As shown, this embodiment provides a multi-expert molecular property prediction method based on adaptive substructure perception, including: decomposing an input molecule into a set of substructures, and identifying positive and negative substructures of the molecule through attribution analysis; For the positive substructures and negative substructures, marginal ternary loss is introduced together with downstream task loss to jointly optimize the substructure recognition quality and obtain substructure representation; Based on the substructure representation, a multi-expert architecture is constructed to generate substructure embedding representations of different experts; Embedding the substructure, designing a dynamic routing mechanism, adaptively assigning expert fusion weights, and generating molecular logits based on different experts; Integrate the logit values of the different experts to obtain the predicted values of molecular properties, and train the optimization model by calculating the loss; Based on the optimized model, downstream task evaluation indicators are used to test the model performance.
[0025] Specifically, such as Figure 2 As shown, S1: Substructure decomposition of molecules in the dataset based on BRICS, and identification of positive and negative substructures of each molecule through attribution analysis; S2: For the positive and negative substructures obtained in S1, the marginal triple loss is introduced together with the downstream task loss to jointly optimize the substructure recognition quality and obtain the substructure representation; S3: Design a multi-expert architecture for the substructure representation obtained in S2 to obtain substructure embedding representations based on different experts; S4: Based on the substructure embedding representation obtained in S3, a dynamic routing mechanism is designed to adaptively assign expert fusion weights to obtain molecular logits based on different experts; S5: Integrate the logit values of different experts to obtain the molecular property prediction value, and train the optimization model by calculating the loss to obtain the optimal model for molecular property prediction; S6: For the optimal model obtained in S5, the model performance is tested using the evaluation indicators corresponding to the downstream tasks.
[0026] As an implementation method in this embodiment, the process of decomposing the input molecule into a substructure set includes: The BRICS splitting method is used to split the input molecule into a set of substructures; Extract molecular node representations based on graph neural networks, selectively mask molecular substructures through substructure mask vectors, and generate substructure embeddings; Attribution analysis was used to compare the prediction differences before and after masking, quantify the contribution of substructures to molecular properties, and select the top ψ substructures with the highest contribution values as positive substructures, and the top ψ substructures with the lowest contribution values as negative substructures.
[0027] Specifically, in step S1, the molecules in the dataset are decomposed into substructures based on BRICS, and the positive and negative substructures of each molecule are identified through attribution analysis, as follows: S1.1: Input molecules Use BRICS splitting method to split into substructure sets ,in is the number of substructures per molecule; S1.2: Using the graph neural network model to obtain the node representation set of molecules ,in For each molecule's node number, define a substructure mask vector to selectively mask the molecule's substructure: ; Therefore, the corresponding s-th substructure embedding for: ; S1.3: Quantify substructure contributions by comparing the differences in predictions before and after masking through attribution analysis: ; in is the molecular graph representation obtained by averaging the pooled node representations, is a multi-layer perceptron, is the label of the molecule; finally, the top one with the highest attribution value is selected The substructure is the “positive substructure”, and the lowest attribution value is The substructure is called a "negative substructure" and the corresponding embedding representation is obtained and .
[0028] As an implementation method in this embodiment, the process of jointly optimizing the substructure recognition quality includes: Taking the original molecular graph, the positive substructure, and the negative substructure as three views, a marginal triplet loss is defined to force the feature similarity between the positive substructure and the molecular graph to be higher than the feature similarity between the negative substructure and the molecular graph; The marginal triplet loss is weightedly summed with the downstream task loss to optimize substructure identification and representation.
[0029] Specifically, such as Figure 3 As shown, in step S2, based on the positive and negative substructures obtained in step S1, the marginal triplet loss is introduced to jointly optimize the substructure recognition quality and obtain the substructure representation, specifically as follows: To optimize the recognition of substructures, the present invention introduces the marginal triplet loss and the corresponding downstream task loss The original molecular graph, "positive substructure" and "negative substructure" are regarded as three different views. The graph representation obtained from each view should ensure high quality while making the "positive substructure" and "negative substructure" distinct to a certain extent. Specifically, the marginal triplet loss is defined as: in is the Sigmoid function, is the marginal loss threshold, represents the i-th molecular graph representation, represents the i-th active substructure graph representation, represents the i-th negative substructure graph representation; Next, and Combined, the overall goal of substructure identification optimization is obtained for: ; in is the loss balance parameter, which is used to adjust the weight of the classification task loss and the triplet loss.
[0030] As an implementation method in this embodiment, the process of building a multi-expert architecture includes: Multiple expert models are assigned to positive and negative substructures respectively. Each expert performs a linear transformation on the substructure representation through a weight matrix to generate an expert-specific substructure embedding representation. The number of experts is determined by the number of tasks and preset parameters.
[0031] Specifically, in step S3, based on the substructure representation obtained in S2, a multi-expert architecture is designed to obtain substructure embedding representations based on different experts, specifically as follows: the positive and negative substructures obtained are assigned to different expert models; tasks, each substructure category is assigned experts, including is the number of experts; each expert is represented by a weight matrix The substructure embedding representation based on the expert model is calculated: ; in is the weight matrix, is the dimension of the embedding representation; for each type of expert, each substructure is embedded Take separately or .
[0032] As an implementation method of this embodiment, the process of the dynamic routing mechanism includes: For the embedded representation of the active substructure, the routing score is calculated through the parameter matrix and the expert weight is assigned; For the embedded representation of the negative substructure, the positive substructure and the negative substructure are concatenated and input into the router to calculate the routing score; The expert embedding representations are weighted and fused according to the routing scores to generate positive logits and negative logits.
[0033] Specifically, such as Figure 4 As shown, in step S4, based on the substructure embedding representation obtained in S3, a dynamic routing mechanism is designed to adaptively assign expert fusion weights to obtain molecular logits based on different experts, as follows: S4.1: For “active substructures”, use the parameter matrix As a router; the present invention uses the Linear mapping layer to obtain the embedded representation of the input router , and then perform a dot product operation with the weight matrix, and the routing score is calculated as: in is the temperature hyperparameter, is the sampling noise; S4.2: For “negative substructure”, it represents the substructure that is negatively correlated with the target attribute, the purpose of which is to reduce its adverse effect on expert optimization; the present invention adopts the parameter matrix As a router; in order to make the model learn more predictive features from the “negative substructure”, the present invention combines the embedding of positive and negative substructures as the joint input of the router ,in It is also obtained by a Linear mapping layer; then the routing score is calculated as: S4.3: Based on routing score and Combine the embedding representations of all experts to obtain positive logits and negative logit : in represents the positive substructure expert embedding, represents the negative substructure expert embedding.
[0034] As an implementation method in this embodiment, the process of calculating the loss training optimization model includes: The importance of experts in a batch is calculated based on the expert routing score, and the balance of expert division of labor is measured by the coefficient of variation to generate the importance loss. The importance loss is weightedly summed with the downstream task loss and the marginal triplet loss to optimize the model parameters.
[0035] Specifically, in step S5, the logit values of different experts are integrated to obtain the molecular property prediction value, and the model is optimized by calculating the loss training to obtain the optimal model for molecular property prediction, as follows: S5.1: The present invention obtains the predicted value of each molecule by integrating the positive logit value and the negative logit value: in, Represents a splicing operation, is the predicted value of the numerator; S5.2: To optimize the expert learning process, this paper introduces the classification task loss. At the same time, it is observed during the training process that a small number of experts tend to dominate all samples. To encourage better division of labor among experts, an importance loss function is added to penalize experts who are too dominant. Given a batch of graphs , the importance of each expert is defined as the sum of the routing scores of all graphs in the batch ,Right now: Therefore, the importance loss is calculated as the mean of the coefficient of variation of the importance values of all experts : in To block the function of gradient propagation, when the coefficient of variation is lower than the preset threshold When , this loss term does not conduct error; the model is then optimized by minimizing the following loss function: in, To balance the two loss parameters, Represents the total loss; finally, the optimal model is used to predict the molecular properties to obtain the prediction results of each molecule .
[0036] As an implementation method of this embodiment, in step S6, based on the optimal model obtained in S5, the model performance is tested using the evaluation index corresponding to the downstream task, specifically as follows: the predicted value of each molecule is obtained using the optimal model obtained after training optimization , and obtain the final prediction score : in is the Sigmoid function; then the evaluation indicators corresponding to the downstream tasks are used, that is, the ROC-AUC indicator is used for classification tasks, and the RMSE indicator is used for regression tasks to evaluate the performance of the model.
[0037] Based on this, an embodiment of the present invention provides a multi-expert molecular property prediction method based on adaptive substructure perception. The present invention significantly improves the accuracy and generalization ability of molecular property prediction through adaptive substructure perception and multi-expert dynamic routing mechanism. Based on BRICS decomposition and attribution analysis, the positive contribution and negative interference of the local structure of the molecule are accurately distinguished, and the impact of noise characteristics on the prediction is reduced; through multi-expert architecture and dynamic routing allocation, the specialized characteristics of different substructures are adaptively integrated to enhance the model's ability to analyze complex molecular topologies; the marginal triple loss and importance loss are jointly optimized to effectively alleviate the data imbalance problem and suppress the expert-dominated phenomenon, ensuring the stability of the model in low-quality data scenarios. In addition, attribution analysis clarifies the contribution of key substructures, improves model interpretability, and provides a reliable decision-making basis for drug research and development.
[0038] Example 2 In this embodiment, a computer terminal device is provided, including: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the methods in the above embodiments.
[0039] In this embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the method in the above embodiment is implemented.
[0040] In this embodiment, an electronic device is further provided, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the method in the above embodiment.
[0041] The above program can be executed in a processor or stored in a memory (or computer-readable medium). Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0042] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more blocks can be implemented by different modules corresponding to different steps.
[0043] This embodiment provides such a device or system. The system is called a multi-expert molecular property prediction system based on adaptive substructure perception, and includes: a substructure decomposition and attribution module for decomposing an input molecule into a set of substructures and identifying positive and negative substructures of the molecule through attribution analysis; A joint optimization module is used to introduce a marginal ternary loss and a downstream task loss to the positive substructure and the negative substructure to jointly optimize the substructure recognition quality and obtain a substructure representation; A multi-expert architecture building module is used to build a multi-expert architecture based on the substructure representation and generate substructure embedding representations of different experts; A dynamic routing assignment module is used to design a dynamic routing mechanism for the substructure embedding representation, adaptively assign expert fusion weights, and generate molecular logits based on different experts; A prediction integration and optimization module is used to integrate the logit values of the different experts to obtain the molecular property prediction value and train the optimization model by calculating the loss; The performance testing module is used to test the model performance based on the optimized model using downstream task evaluation indicators.
[0044] As an implementation method of this embodiment, the substructure decomposition and attribution module includes: BRICS splitting unit, used to split the input molecule into a set of substructures using the BRICS method; The substructure embedding generation unit is used to extract molecular node representations based on graph neural networks and selectively mask molecular substructures through substructure mask vectors to generate substructure embeddings; The substructure contribution quantification unit is used to compare the prediction differences before and after masking through attribution analysis, quantify the contribution of substructures to molecular properties, and select the top ψ substructures with the highest contribution values as positive substructures and the top ψ substructures with the lowest contribution values as negative substructures.
[0045] As an implementation method of this embodiment, the joint optimization module includes: a triplet loss definition unit, configured to define a marginal triplet loss by taking the original molecular graph, the positive substructure, and the negative substructure as three views, and forcing the feature similarity between the positive substructure and the molecular graph to be higher than the feature similarity between the negative substructure and the molecular graph; A loss weighted optimization unit is used to weight the sum of the marginal triplet loss and the downstream task loss to optimize substructure identification and representation.
[0046] As an implementation method in this embodiment, the multi-expert architecture building module includes: an expert assignment unit, for assigning multiple groups of expert models to positive substructures and negative substructures respectively; A linear transformation unit is used to perform a linear transformation on the substructure representation through a weight matrix to generate an expert-specific substructure embedding representation; the number of experts is determined by the number of tasks and preset parameters.
[0047] As an implementation method of this embodiment, the dynamic routing allocation module includes: Active subrouting unit, used to embed the active substructure, calculate the routing score through the parameter matrix, and assign expert weights; The negative sub-routing unit is used to embed the negative sub-structure, concatenate the embedding representations of the positive sub-structure and the negative sub-structure, and input them into the router to calculate the routing score; The expert fusion unit is used to fuse the expert embedding representation according to the weighted routing score to generate positive logit and negative logit.
[0048] As an implementation method of this embodiment, the prediction integration and optimization module includes: Importance loss calculation unit, used to calculate the importance of experts in a batch based on the expert routing score, measure the balance of expert division of labor through the coefficient of variation, and generate importance loss; The multi-loss optimization unit is used to perform a weighted summation of the importance loss, the downstream task loss, and the marginal triplet loss to optimize the model parameters.
[0049] The system or device is used to implement the functions of the method in the above-mentioned embodiment. Each module in the system or device corresponds to each step in the method, which has been explained in the method and will not be repeated here.
[0050] Through the above implementation, the problem of multi-expert molecular property prediction based on adaptive substructure perception in the related art is solved, thereby ensuring that the problem of confusion that occurs when using a single model in existing methods to process diverse and complex molecular features is solved. This method significantly improves the accuracy and interpretability of the molecular property prediction model.
[0051] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A multi-expert molecular property prediction method based on adaptive substructure perception, characterized in that: The following steps are involved: decomposing an input molecule into a set of substructures, and identifying positive and negative substructures of the molecule through attribution analysis; For the positive substructures and negative substructures, marginal ternary loss is introduced together with downstream task loss to jointly optimize the substructure recognition quality and obtain substructure representation; Based on the substructure representation, a multi-expert architecture is constructed to generate substructure embedding representations of different experts; Embedding the substructure, designing a dynamic routing mechanism, adaptively assigning expert fusion weights, and generating molecular logits based on different experts; Integrate the logit values of the different experts to obtain the predicted values of molecular properties, and train the optimization model by calculating the loss; Based on the optimized model, downstream task evaluation indicators are used to test the model performance.
2. The method according to claim 1, characterized in that The process of decomposing an input molecule into a set of substructures includes: The BRICS splitting method is used to split the input molecule into a set of substructures; Extract molecular node representations based on graph neural networks, selectively mask molecular substructures through substructure mask vectors, and generate substructure embeddings; Attribution analysis was used to compare the prediction differences before and after masking, quantify the contribution of substructures to molecular properties, and select the top ψ substructures with the highest contribution values as positive substructures, and the top ψ substructures with the lowest contribution values as negative substructures.
3. The method according to claim 1, characterized in that The process of jointly optimizing the substructure recognition quality includes: Taking the original molecular graph, the positive substructure, and the negative substructure as three views, a marginal triplet loss is defined to force the feature similarity between the positive substructure and the molecular graph to be higher than the feature similarity between the negative substructure and the molecular graph; The marginal triplet loss is weightedly summed with the downstream task loss to optimize substructure identification and representation.
4. The method according to claim 1, wherein The process of building a multi-expert architecture includes: Multiple expert models are assigned to positive and negative substructures respectively. Each expert performs a linear transformation on the substructure representation through a weight matrix to generate an expert-specific substructure embedding representation. The number of experts is determined by the number of tasks and preset parameters.
5. The method according to claim 1, wherein The process of the dynamic routing mechanism includes: For the embedded representation of the active substructure, the routing score is calculated through the parameter matrix and the expert weight is assigned; For the embedded representation of the negative substructure, the positive substructure and the negative substructure are concatenated and input into the router to calculate the routing score; The expert embedding representations are weighted and fused according to the routing scores to generate positive logits and negative logits.
6. The method according to claim 1, characterized in that The process of calculating the loss training optimization model includes: The importance of experts in a batch is calculated based on the expert routing score, and the balance of expert division of labor is measured by the coefficient of variation to generate the importance loss. The importance loss is weightedly summed with the downstream task loss and the marginal triplet loss to optimize the model parameters.
7. A multi-expert molecular property prediction system based on adaptive substructure perception, characterized by: The system comprises: a substructure decomposition and attribution module for decomposing an input molecule into a set of substructures and identifying positive and negative substructures of the molecule through attribution analysis; A joint optimization module is used to introduce a marginal ternary loss and a downstream task loss to the positive substructure and the negative substructure to jointly optimize the substructure recognition quality and obtain a substructure representation; A multi-expert architecture building module is used to build a multi-expert architecture based on the substructure representation and generate substructure embedding representations of different experts; A dynamic routing assignment module is used to design a dynamic routing mechanism for the substructure embedding representation, adaptively assign expert fusion weights, and generate molecular logits based on different experts; A prediction integration and optimization module is used to integrate the logit values of the different experts to obtain the molecular property prediction value and train the optimization model by calculating the loss; The performance testing module is used to test the model performance based on the optimized model using downstream task evaluation indicators.
8. The system according to claim 7, characterized in that The substructure decomposition and attribution module includes: BRICS splitting unit, used to split the input molecule into a set of substructures using the BRICS method; The substructure embedding generation unit is used to extract molecular node representations based on graph neural networks and selectively mask molecular substructures through substructure mask vectors to generate substructure embeddings; The substructure contribution quantification unit is used to compare the prediction differences before and after masking through attribution analysis, quantify the contribution of substructures to molecular properties, and select the top ψ substructures with the highest contribution values as positive substructures and the top ψ substructures with the lowest contribution values as negative substructures.
9. A computer terminal device, characterized in that: include: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-expert molecular property prediction method based on adaptive substructure perception according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting molecular properties based on adaptive substructure perception by multiple experts is implemented as claimed in any one of claims 1 to 6.