Drug molecule generation method capable of explaining federal large model
By adopting an interpretable federal big model method in drug molecule generation, the problem of scarcity of data and insufficient transparency of the generation results is solved, and efficient and reliable drug molecule generation is achieved, which significantly improves the performance and robustness of the model.
Patent Information
- Application Number
- CN202411993968.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
AI Technical Summary
The existing drug molecule generation technology lacks sufficient data for effective training of unilateral models, and the generative logic and features are insufficiently explained in the process of generating molecules, resulting in a lack of transparency and verifiability in the generation results.
The drug molecule generation method that can explain the federal big model is adopted to construct and initialize the drug molecule generation model through data feature extraction and model initialization; the federal learning framework is used to optimize the parameter to generate drug molecule structure information, and improve the interpretability of the model through enhanced mechanisms.
Effectively ensure data privacy, improve model performance and data utilization efficiency, significantly improve model interpretability and credibility of results, reduce drug design cycles, and improve the generation ability and diversity of new molecules.
Smart Images

Figure CN119920356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for generating drug molecules that can interpret a large federated model. Background Art
[0002] In the field of drug research and development, the method of generating drug molecules plays an important role in the discovery and optimization of new drugs. With the increasing complexity of drug research and development and the continuous rise in research and development costs, how to develop efficient and innovative drug molecule generation technology has become the focus of current scientific research. The biological activity and properties of drug molecules are affected by multiple factors, including complex features such as molecular structure, drug similarity, and biological targets. The effective integration of these features can significantly improve the effectiveness of generated molecules.
[0003] In recent years, with the rapid development of technologies such as interpretable models, large language models, and federated learning, the integration of various technologies has made data privacy protection and efficient model training possible. Specifically, federated learning allows institutions to use decentralized drug data for local model training by building a distributed collaboration framework, avoiding the direct transmission of sensitive data, thereby ensuring data privacy while improving model performance. In addition, in order to improve the interpretability of the model, some technologies have introduced interpretability mechanisms that can analyze the structural features and generation process of generated molecules. This mechanism not only helps drug developers understand the logic of molecular design, but also identifies key features in the generation process, thereby enhancing trust in the generated results.
[0004] For example, some optimization tools such as Sophia optimizer have been proposed to improve the efficiency of model training and resource utilization through efficient second-order optimization methods. In addition, models such as Huawei Cloud Pangu Drug Molecular Model can perform AI-assisted design of small molecule drugs throughout the entire process, improving drug design efficiency by 33% and molecular binding energy optimization by more than 40%. Large language models such as Y-Mol, which combine multi-faceted knowledge guidance, have shown high efficiency and potential in lead compound discovery, molecular property prediction, and drug interaction identification. The development of these technologies has promoted the drug molecule generation technology to move towards high quality, diversity, and explainability.
[0005] However, existing drug molecule generation technologies still have some limitations. First, since drug development data are scattered among different institutions, there is a lack of sufficient data for effective training of unilateral models. Although data sharing can reduce R&D costs and improve the quality of molecule generation, direct data sharing poses a risk of privacy leakage due to data protection regulations. In addition, existing methods do not adequately explain the generation logic and features during the process of generating molecules, resulting in a lack of transparency and verifiability in the generation results, which affects the trust of R&D personnel in the quality of the generated molecules. Summary of the invention
[0006] The purpose of the present invention is to provide a method for generating drug molecules with an interpretable federated large model, so as to solve the problems in the prior art of lack of sufficient data for effective training of unilateral models and insufficient explanation of generation logic and features in the process of generating molecules, resulting in lack of transparency and verifiability of the generated results.
[0007] To achieve the above object, the present invention provides the following technical solution: a method for generating drug molecules that can explain a federated large model, the method comprising: S1. Through data feature extraction and model initialization steps, multimodal features are extracted from distributed data, and a drug molecule generation model is constructed and initialized; S2. Optimize the parameters of the federated learning client through the large model, including fine-tuning local training based on the distributed client, and integrate the fine-tuned parameters into the global model of the server to improve model performance and data utilization efficiency; S3. Generate drug molecular structure information based on the optimized model, including molecular description sequence and molecular graph; S4. Improve the interpretability of the model through enhancement mechanisms to reveal the feature association and weight distribution in the process of drug molecule generation; S5. Analyze and evaluate the generated drug molecules to verify their chemical rationality, drug potential and generate visualization results.
[0008] Preferably, the S1 further comprises: Systematically collect multimodal drug molecule data from multiple drug development institutions, including SMILES sequences and molecular graphs, use denoising algorithms to remove redundant data, and standardize features through Z-score to ensure data consistency and reliability; The embedding layer is used to map discrete molecular features to a continuous vector space to reveal the potential relationship between molecules. Subsequently, the graph neural network structure is used to propagate information through the adjacency matrix and node features to extract high-dimensional feature representations. The specific calculation formula is as follows: ; in, represents the node features of the i-th layer, N(i) is the neighbor set of node i; On the central server, based on the requirements of the drug molecule generation task, an adapted generation model architecture is selected, and the parameters of the selected model architecture are randomly initialized to provide initial weights.
[0009] Preferably, S2 further includes: Under the federated learning framework, each client performs independent training based on local drug data. Each client extracts features from local drug data, updates model parameters, calculates gradients, and uploads the calculated gradient information to the central server. This work uses the following formula to fine-tune the model parameters: ; Among them, B and A are training parameters, , During the entire training period, pre-training Keep it fixed and do not update the gradient. is an intermediate learning variable. After completing the global training round, each client will Upload to the server for parameter update; The central server receives and aggregates the gradient information from each client, and updates the parameters of the global model by weighted average. The specific calculation formula is as follows: ; in, is the weight of client i, which is determined by its sample number n i Decide, is the global model parameter of the tth round, η is the learning rate, is the gradient calculated by client i in round t; After each round of training, the client makes local adjustments based on the updated global model to optimize model performance.
[0010] Preferably, S3 further includes: Based on the trained large model, the target drug molecule is used as input, and the target mapping model from feature to structure is constructed by analyzing its atomic features and molecular structure relationship; The large model dynamically adjusts the mapping of features through the generation process; The generated drug molecules are presented in the form of SMILES sequences and molecular graphs.
[0011] Preferably, the S4 further includes: Build a distributed causal reasoning framework and use graph relationship networks to analyze the impact paths of drug characteristics; The self-attention mechanism is introduced to automatically adjust the feature weight by calculating the contribution of each feature to the model decision. The specific calculation formula is as follows: ; Among them, Q, K, V are queries, is the dimension scaling factor.
[0012] Preferably, S5 includes evaluating the drug development potential of the generated molecules through molecular docking simulation and chemical indexes, and visually displaying the evaluation results of the molecular structure and key performance indicators.
[0013] Preferably, S1 also includes using a denoising algorithm to remove redundant information of data collected in distributed clients, and performing consistency processing on the data through a standardized method to improve data reliability.
[0014] Preferably, S3 also includes optimizing the structural rationality and chemical diversity of the generated molecules by adjusting the feature mapping method.
[0015] Preferably, S5 also includes using a method combining binding energy analysis and molecular docking simulation to evaluate the binding stability of the molecule with the target protein and its drug development potential.
[0016] Preferably, S5 also includes using a machine learning model to predict the ADMET properties of the generated molecules.
[0017] It can be seen from the above technical solution that the present invention has the following beneficial effects: This method for generating drug molecules based on an interpretable federated large model, through federated distributed learning, enables all participants to achieve collaborative training without sharing original data, effectively protects data privacy, generates drug molecules using a large model, and efficiently parses input features through deep learning to generate molecular structures that meet target properties, reducing the design cycle of traditional drugs. It uses causal analysis and attention mechanisms to intuitively reveal the relationship between input features and generated results, and clarifies the contribution of features to the generation path and the basis for decision-making, thereby significantly improving the interpretability of the model and the credibility of the results. It increases the training sample size of the large model through federated learning, effectively alleviating the problem of poor model training caused by data scarcity, and significantly improving the generalization ability, accuracy and robustness of the model. On the basis of protecting data privacy, this method significantly improves the generation ability and diversity of new molecules, while using causal analysis and attention mechanisms to enhance the transparency and credibility of the model, improve the performance and robustness of the model, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flow chart of the method of the present invention; Figure 2 The overall structural flow chart of a drug molecule generation method that can interpret the federated large model proposed by the present invention; Figure 3 A schematic diagram of the interaction between a molecule and a target crystal structure generated by the molecular docking simulation of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] like Figure 1-Figure 3 As shown, the present invention provides a technical solution: a method for generating drug molecules that can explain a federated large model, the method comprising: S1. Through data feature extraction and model initialization steps, multimodal features are extracted from distributed data, and a drug molecule generation model is constructed and initialized; S2. Optimize the parameters of the federated learning client through the large model, including fine-tuning local training based on the distributed client, and integrate the fine-tuned parameters into the global model of the server to improve model performance and data utilization efficiency; S3. Generate drug molecular structure information based on the optimized model, including molecular description sequence and molecular graph; S4. Improve the interpretability of the model through enhancement mechanisms to reveal the feature association and weight distribution in the process of drug molecule generation; S5. Analyze and evaluate the generated drug molecules to verify their chemical rationality, drug potential and generate visualization results.
[0021] In the above method, first, through distributed data feature extraction and model initialization (S1), molecular feature information from multiple data sources, such as molecular structure, pharmacological properties, and molecular activity, is integrated based on a distributed computing framework. The model is initialized through a specific encoder-decoder architecture to capture multimodal input features. Then, through a federated learning mechanism (S2), local data training is completed on the client to optimize the processing efficiency of distributed data; the server collects model parameters from each client for global update to avoid direct access to distributed data, thereby protecting data privacy. After the model optimization is completed, drug molecular structure information is generated based on this model (S3), specifically including generating description sequences (such as SMILES sequences) and molecular graph structures of drug molecules through decoders. These outputs can fully characterize the physicochemical properties and spatial topological structures of the generated drug molecules. Then, by introducing an interpretability enhancement mechanism (S4), the attention mechanism or feature contribution analysis method is used to reveal the features that the model focuses on during the generation process. Specifically, the weight of the influence of feature input on molecule generation can be analyzed to reveal the key decision-making process of the model, which is convenient for further optimization. Finally, the generated molecules are submitted to the analysis and evaluation module (S5). This module includes chemical rationality verification (such as bond length and valence state analysis), drug potential assessment (such as ADMET prediction), and visualization of generated results (such as three-dimensional molecular graph display) to ensure the quality and practicality of the generated molecules. This method optimizes model performance through a federated learning mechanism, fully utilizes the diversity of distributed data, and protects data privacy, overcoming the limitations of traditional centralized data collection methods. The generated drug molecule model can effectively generate molecules with chemical rationality and drug potential, and achieves a high level of interpretability through enhanced mechanisms, helping researchers understand the key features and model decision logic in the generation process, and improving the transparency and reliability of the method. This method can improve the efficiency and accuracy of drug molecule design while reducing the trial and error costs in the process of new drug development.
[0022] S1 also includes systematic collection of multimodal drug molecule data from multiple drug development institutions, including SMILES sequences and molecular graphs, using denoising algorithms to remove redundant data, and standardizing features through Z-score to ensure data consistency and reliability; The embedding layer is used to map discrete molecular features to a continuous vector space to reveal the potential relationship between molecules. Subsequently, the graph neural network structure is used to propagate information through the adjacency matrix and node features to extract high-dimensional feature representations. The specific calculation formula is as follows: ; in, represents the node features of the i-th layer, N(i) is the neighbor set of node i; On the central server, based on the requirements of the drug molecule generation task, an adapted generation model architecture is selected, and the parameters of the selected model architecture are randomly initialized to provide initial weights.
[0023] In the above methods, the systematic collection and processing of drug molecule data (such as SMILES sequences and molecular graphs) aims to provide a multimodal input data source. The denoising algorithm is used to remove data redundancy and outliers, and the Z-score standardization is used to ensure the uniformity and reliability of feature distribution, providing a high-quality data foundation for subsequent model training. Through the embedding layer, the molecular features are projected into a high-dimensional continuous vector space. This operation retains the molecular characteristics and similarity information. As the core model, the graph neural network (GNN) uses the adjacency matrix and node features for information propagation, and integrates the neighborhood information layer by layer to capture the potential associations between molecules. W in the formula (l) is the weight matrix, σ is the activation function, which ensures the nonlinear processing capability of the data and gradually improves the model's ability to model complex molecular relationships. On the central server, an adapted generative model architecture is selected according to the task requirements, such as a variational autoencoder (VAE) or a generative adversarial network (GAN). The randomly initialized parameters are gradually updated through the optimization algorithm, so that the model has a strong generalization ability. This implementation method can systematically collect and process multimodal drug molecule data, improve data consistency and feature quality, and effectively improve model training efficiency and performance. The introduction of graph neural networks enables the model to capture complex relationships between molecules in high-dimensional space and enhance the accuracy and diversity of molecule generation. By randomly initializing the model parameters, the sufficient randomness of the generative model in the initial stage is ensured, which helps to learn the global optimal weights. This method can better adapt to different task requirements and provide a flexible and efficient solution for drug molecule generation.
[0024] S2 also includes that under the federated learning framework, each client performs independent training based on local drug data, each client extracts features from local drug data, updates model parameters and calculates gradients, and uploads the calculated gradient information to the central server; This work uses the following formula to fine-tune the model parameters: ; Among them, B and A are training parameters, , , during the entire training period, pre-training Keep it fixed and do not update the gradient. is an intermediate learning variable. After completing the global training round, each client will Upload to the server for parameter update; The central server receives and aggregates the gradient information from each client, and updates the parameters of the global model by weighted average. The specific calculation formula is as follows: ; in, is the weight of client i, which is determined by its sample number n i Decide, is the global model parameter of the tth round, η is the learning rate, is the gradient calculated by client i in round t; After each round of training, the client makes local adjustments based on the updated global model to optimize model performance.
[0025] In the above method, the federated learning framework consists of a client and a central server. Each client independently performs local model training to protect data privacy. Through local training, each client uses local drug molecule data to calculate the gradient information of the model parameters and uploads these gradients to the central server. After receiving the gradient information from each client, the central server calculates the global gradient based on the weighted sample size of each client data and updates the global model parameters through the formula. The weight value w i Determined by the amount of data on the client, it ensures that clients with larger amounts of data have a higher contribution ratio in the model update. The learning rate η controls the step size of the parameter update to ensure the stability of the model optimization. Through multiple rounds of iterations, each client obtains the global model updated by the central server, and fine-tunes it locally to optimize the model performance, thereby achieving collaborative optimization of the global model and local data. This process effectively improves the adaptability of the federated learning model to distributed data. Under the premise of protecting data privacy, this implementation method realizes the efficient use of distributed data through the federated learning framework, effectively avoiding the risk of privacy leakage caused by data centralization. At the same time, the weighted update of the model parameters ensures the contribution of different data distributions to the global model, and improves the model's adaptability to heterogeneous data. In addition, the performance of the model on client data is further optimized through local fine-tuning, thereby improving the accuracy and stability of drug molecule generation.
[0026] S3 also includes a target mapping model from features to structures based on the trained large model, taking the target drug molecule as input, and analyzing its atomic features and molecular structure relationships; the large model dynamically adjusts the mapping mode of features through the generation process; the generated drug molecules are presented in the form of SMILES sequences and molecular graphs. In the above method, the target drug molecule is first taken as input, and the large model is used to analyze the atomic features (such as atom type, electronegativity, valence, etc.) and molecular structure relationships (such as bond type, topological structure) of the drug molecule. These parsing processes construct a target mapping model from features to molecular structures and define the mapping rules between drug molecule features and generated structures. During the generation process, the large model optimizes the generation path by dynamically adjusting the mapping mode and combining the weight changes of different features in the molecular generation process. For example, in the early stage of molecular generation, the model may focus on capturing overall structural features, while in the later stage, the model pays more attention to detailed features (such as the accurate connection of functional groups). This dynamic adjustment process is implemented by an attention mechanism or a generative adversarial algorithm to ensure that the generated molecules achieve a balance between global structure and local characteristics. Finally, the generated drug molecules are presented in the form of SMILES sequences and molecular graphs. The SMILES sequence provides a compact way to describe molecules, while the molecular graph intuitively displays the topological structure of the molecule in a graph structure, which is convenient for subsequent analysis and verification. This implementation method ensures the pertinence and accuracy of drug molecule generation through the construction of feature analysis and mapping models of target drug molecules. The mechanism of dynamically adjusting the feature mapping method makes the generation process more flexible, thereby improving the structural integrity and rationality of the generated molecules at different stages. The generated drug molecules are presented in the form of SMILES sequences and molecular graphs at the same time, which is convenient for researchers to quickly analyze and verify. This method not only improves the quality of drug molecule generation, but also enhances the interpretability and practicality of the model.
[0027] S4 also includes: building a distributed causal reasoning framework to analyze the impact path of drug characteristics using graph relationship networks; The self-attention mechanism is introduced to automatically adjust the feature weight by calculating the contribution of each feature to the model decision. The specific calculation formula is as follows: ; Among them, Q, K, V are queries, is the dimension scaling factor.
[0028] In the above method, the distributed causal reasoning framework analyzes the influence paths between the features of drug molecules through a graph relationship network (such as a causal graph or a Bayesian network). This framework is based on the connectivity of feature relationships, explores the role of key features in drug molecule generation decisions, and reveals the implicit causal relationships in the drug molecule generation process. As a key component for interpretability enhancement, the self-attention mechanism dynamically adjusts the feature weights by calculating the influence of each feature on the model decision. Specifically, the query (Q) represents the target feature vector of the model, and the key (K) and value (V) represent the importance of the feature and its contribution to the decision, respectively. Through the attention score calculated in the formula, the importance of the features can be quantified and ranked, helping the model to focus on the most decision-making features when generating drug molecules. Combining causal reasoning with the attention mechanism, the large model can dynamically adjust the weights of the feature influence paths, thereby significantly improving the transparency and interpretability of the generation process, and facilitating the analysis of the logical basis of the generation decision. By constructing a distributed causal reasoning framework, this embodiment can systematically reveal the feature associations and causal paths in the drug molecule generation process, and improve the model's ability to interpret complex generation processes. The introduction of the self-attention mechanism enables dynamic adjustment of feature weights, allowing the model to focus on key features and optimize the generation of decision logic. This combination enhances the transparency of the model and provides a clear explanation for subsequent drug molecule optimization. In addition, this method improves the scientificity and repeatability of drug molecule generation, which helps support decision-making in the drug development process.
[0029] S5 includes evaluating the drug development potential of generated molecules through molecular docking simulation and chemical indicators, and visually displaying the evaluation results of molecular structure and key performance indicators. In the above method, the generated drug molecules are verified for their drug development potential through molecular docking simulation and chemical indicator evaluation. Molecular docking simulation uses computational chemistry methods to dock the generated molecules with the target protein binding site to evaluate the binding energy, binding stability and potential ligand-receptor interaction patterns. These results help determine the drug potential of the generated molecules on specific targets. At the same time, the chemical indicator evaluation quantitatively analyzes the physicochemical properties and drug development-related characteristics of the generated molecules, including molecular weight, LogP value (hydrophobicity), number of hydrogen bond donors and receptors, and ADMET (absorption, distribution, metabolism, excretion and toxicity) characteristics. These indicators provide key judgment basis for the chemical rationality and drug development feasibility of the generated molecules. The evaluation results are presented in a visual form, including the three-dimensional structure of the molecule, the ligand-receptor interaction diagram of the binding site, and the quantitative chart of the chemical indicators. These visualization results help researchers quickly understand the characteristics and drug potential of the generated molecules, and assist in drug screening and optimization decisions. This implementation provides a comprehensive analysis of the drug development potential of the generated molecules through molecular docking simulation and chemical index evaluation, ensuring the chemical rationality and biological activity of the generated molecules. Visual display of evaluation results not only enhances the interpretability of model generation results, but also improves researchers' intuitive understanding of the characteristics of generated molecules, facilitating further screening and optimization. This method can effectively reduce the trial and error cost of drug development, shorten the R&D cycle, and improve the practicality and accuracy of molecular generation results.
[0030] S1 also includes using a denoising algorithm to remove redundant information from data collected in distributed clients, and using a standardization method to process the data for consistency to improve the reliability of the data. In the above method, the denoising algorithm is used to remove redundant information from data collected in distributed clients. The denoising algorithm can improve the quality of data by identifying and removing outliers, repeated information, or low-correlation features based on statistical methods (such as principal component analysis PCA) or machine learning-based autoencoder models. At the same time, the standardization method processes the data for consistency to ensure the uniformity of the distribution of data features of each distributed client. For example, the data can be converted into a distribution with a mean of 0 and a standard deviation of 1 by using the Z-score standardization method, or the data can be normalized to a fixed range (such as [0, 1]) by using the Min-Max normalization method. This process can eliminate the scale differences between data from different clients, improve the overall reliability of the data and the training effect of the model. Through the dual processing of denoising and standardization, the distributed data is optimized into high-quality and highly consistent inputs, providing a reliable data basis for model training in the federated learning framework. This implementation method effectively removes redundant or useless information in distributed clients by using a denoising algorithm, thereby improving the purity and quality of the data. Standardization further unifies the data distribution of different clients and avoids model training deviations caused by differences in data feature scales. This method significantly improves the reliability of distributed data and provides high-quality data support for subsequent drug molecule generation model training, thereby improving model performance and the reliability of generation results.
[0031] S3 also includes optimizing the structural rationality and chemical diversity of generated molecules by adjusting the feature mapping method. In the above method, in order to optimize the structural rationality and chemical diversity of generated molecules, the large model achieves dynamic optimization by adjusting the feature mapping method. The adjustment of the feature mapping method is based on the calculation and feedback mechanism of the contribution of molecular features in the generation process. For example, in the early stage of molecular generation, the model focuses on the mapping of global structural features to ensure that the generated molecules have the basic rationality of the chemical structure (such as atomic valence, bonding rules, etc.); in the later stage of generation, the model adjusts the mapping weights to increase the focus on specific chemical features (such as functional groups, chiral centers) to enhance chemical diversity. This dynamic adjustment process can be achieved through attention mechanisms, graph neural networks or other reinforcement learning methods. Among them, the attention mechanism calculates the contribution value of each feature in the current generation step and adjusts the feature weights to optimize the generation path; the graph neural network updates the node embedding to achieve a fine description of the molecular structure relationship; the reinforcement learning method combines the reward mechanism to optimize the chemical rationality and diversity of the generated molecules and continuously adjusts the feature mapping strategy. Ultimately, the molecules generated by the model have the rationality of the chemical structure (such as meeting the chemical bonding rules) and diversity (such as the richness of molecular structure and functional group types), thereby meeting the needs of different drug research and development. This implementation method effectively optimizes the structural rationality of the generated molecules by dynamically adjusting the feature mapping method, ensuring that the generated molecules meet the basic requirements of chemical rules and drug development. At the same time, through the enhanced focus on specific chemical features, the large model significantly improves the chemical diversity of the generated molecules, enriches the possibility of molecular generation, and provides more innovative options for drug molecule design. This method can balance the rationality and diversity of molecule generation, and improve the applicability and practical value of drug molecule generation.
[0032] S5 also includes a method that combines binding energy analysis and molecular docking simulation for the generated molecules to evaluate the binding stability and drug development potential of the molecules and target proteins. In the above method, binding energy analysis is combined with molecular docking simulation to comprehensively evaluate the interaction stability and potential drug development value of the generated molecules and target proteins. Binding energy analysis: Based on quantum chemical calculations or molecular mechanics methods, the free energy changes of the binding between the generated molecules and the target proteins are evaluated, including key energy terms such as electrostatic interactions, van der Waals forces and solvent effects. The lower the binding energy, the higher the stability of the binding between the generated molecules and the target protein, which provides an important reference indicator for molecular screening. Molecular docking simulation: Using docking algorithms (such as AutoDock or DOCK), the binding sites of the generated molecules and the target proteins are subjected to ligand docking simulation to analyze the binding mode, interaction sites and key hydrogen bonds or hydrophobic interactions between the molecules and the target. The binding stability is further verified by simulating the conformation of the generated molecules at the binding site and the adaptability of the binding site. The combination of the two methods provides a multi-dimensional evaluation of the drug development potential of the generated molecules by quantifying the binding energy and visualizing the molecular docking results. For example, binding energy analysis provides global quantitative indicators, and molecular docking simulation displays detailed structural information of the interaction between molecules and proteins, helping researchers evaluate the developability and potential application scenarios of generated molecules. This embodiment comprehensively evaluates the binding stability and drug development potential of generated molecules and target proteins through the combination of binding energy analysis and molecular docking simulation. Binding energy analysis provides accurate quantitative indicators, while molecular docking simulation reveals the interaction mechanism between molecules and targets through the visualization of binding sites. The synergistic application of the two methods significantly improves the efficiency of screening and optimization of generated molecules, reduces the risk of blind experiments in the drug development process, and provides high-quality scientific support for new drug development.
[0033] S5 also includes the use of machine learning models to predict the ADMET properties of generated molecules. In the above method, the ADMET properties (absorption, distribution, metabolism, excretion and toxicity) of the generated molecules are predicted by machine learning models to comprehensively evaluate the drug development potential of the generated molecules. Feature extraction: structural features are extracted from the generated molecules, including molecular descriptors (such as molecular weight, LogP value, number of hydrogen bond donors and acceptors, etc.) and molecular topological features based on graph representation. A method based on molecular graph neural network (GNN) can also be introduced to encode the molecular structure into a high-dimensional vector representation as model input. Machine learning model construction: supervised learning models (such as random forests, support vector machines, deep neural networks, etc.) or dedicated models based on graph neural networks are used to predict the ADMET properties of generated molecules. The model is trained with experimental data of existing drug molecules to learn the mapping relationship between different molecular features and ADMET properties. For example, the toxicity category can be predicted by a classification model, and continuous value properties such as solubility and distribution coefficient can be predicted by a regression model. Result analysis: The prediction results are presented in numerical form or classification labels, such as absorption rate, hepatotoxicity prediction score, metabolic stability score, etc. These results are compared with drug development standards to provide guidance for the screening and optimization of generated molecules. Through the above process, the machine learning model can quickly and efficiently evaluate the ADMET properties of generated molecules, provide a scientific basis for the drug development process, and reduce the cost and time of experimental verification. This embodiment predicts the ADMET properties of generated molecules through a machine learning model, significantly improving the efficiency of drug development and the accuracy of initial screening. Compared with traditional experimental evaluations, machine learning models can quickly process large-scale generated molecules and reduce research and development costs. Through high-precision ADMET property predictions, researchers can exclude molecules with high toxicity and metabolic instability in the early screening stage, and concentrate resources on developing high-potential molecules, thereby shortening the drug development cycle and optimizing the research and development process.
[0034] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for generating drug molecules that can explain a federated large model, characterized in that: The method comprises: S1. Through data feature extraction and model initialization steps, multimodal features are extracted from distributed data, and a drug molecule generation model is constructed and initialized; S2. Optimize the parameters of the federated learning client through the large model, including fine-tuning local training based on the distributed client, and integrate the fine-tuned parameters into the global model of the server to improve model performance and data utilization efficiency; S3. Generate drug molecular structure information based on the optimized model, including molecular description sequence and molecular graph; S4. Improve the interpretability of the model through enhancement mechanisms to reveal the feature association and weight distribution in the process of drug molecule generation; S5. Analyze and evaluate the generated drug molecules to verify their chemical rationality, drug potential and generate visualization results.
2. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S1 further comprises: Systematically collect multimodal drug molecule data from multiple drug development institutions, including SMILES sequences and molecular graphs, use denoising algorithms to remove redundant data, and standardize features through Z-score to ensure data consistency and reliability; The embedding layer is used to map discrete molecular features to a continuous vector space to reveal the potential relationship between molecules. Subsequently, the graph neural network structure is used to propagate information through the adjacency matrix and node features to extract high-dimensional feature representations. The specific calculation formula is as follows: ; in, represents the node features of the i-th layer, N(i) is the neighbor set of node i; On the central server, based on the requirements of the drug molecule generation task, an adapted generation model architecture is selected, and the parameters of the selected model architecture are randomly initialized to provide initial weights.
3. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S2 further includes: Under the federated learning framework, each client performs independent training based on local drug data. Each client extracts features from local drug data, updates model parameters, calculates gradients, and uploads the calculated gradient information to the central server. This work uses the following formula to fine-tune the model parameters: ; Among them, B and A are training parameters, , , during the entire training period, pre-training Keep it fixed and do not update the gradient. is an intermediate learning variable. After completing the global training round, each client will Upload to the server for parameter update; The central server receives and aggregates the gradient information from each client, and updates the parameters of the global model by weighted average. The specific calculation formula is as follows: ; in, is the weight of client i, which is determined by its sample number n i Decide, is the global model parameter of the tth round, η is the learning rate, is the gradient calculated by client i in round t; After each round of training, the client makes local adjustments based on the updated global model to optimize model performance.
4. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S3 further includes: Based on the trained large model, the target drug molecule is used as input, and the target mapping model from feature to structure is constructed by analyzing its atomic features and molecular structure relationship; The large model dynamically adjusts the mapping of features through the generation process; The generated drug molecules are presented in the form of SMILES sequences and molecular graphs.
5. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S4 further comprises: Build a distributed causal reasoning framework and use graph relationship networks to analyze the impact paths of drug characteristics; The self-attention mechanism is introduced to automatically adjust the feature weight by calculating the contribution of each feature to the model decision. The specific calculation formula is as follows: ; Among them, Q, K, V are queries, is the dimension scaling factor.
6. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S5 includes evaluating the drug development potential of the generated molecules through molecular docking simulation and chemical indexes, and visually displaying the evaluation results of the molecular structure and key performance indicators.
7. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S1 also includes using a denoising algorithm to remove redundant information in the data collected in the distributed client, and performing consistency processing on the data through a standardized method to improve the reliability of the data.
8. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S3 also includes optimizing the structural rationality and chemical diversity of the generated molecules by adjusting the feature mapping method.
9. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S5 also includes a method of combining binding energy analysis with molecular docking simulation for the generated molecules to evaluate the binding stability between the molecules and the target protein and the drug development potential.
10. The method for generating drug molecules capable of interpreting a federated large model according to claim 1, characterized in that: The S5 also includes using a machine learning model to predict the ADMET properties of the generated molecules.
Citation Information
Cited By
Drug intermediate database construction and AI intelligent retrieval method
CN120199374A
Drug stability analysis system
CN121011259A
Anti-carbapenem drug-resistant enterobacter active molecule screening method based on LightGBM algorithm, medium and equipment
CN122157784A