An interpretable artificial intelligence-based molecular design constraint condition generation method, system, device and medium
By employing a molecular design method based on interpretable artificial intelligence, and utilizing machine learning and SHAP interpretive analysis to generate molecular design constraints, the problem of traditional molecular design relying on expert experience is solved, and an efficient and explicit molecular design process is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIV INST OF ADVANCED AGRI SCI
- Filing Date
- 2026-03-28
- Publication Date
- 2026-07-07
Smart Images

Figure CN121922245B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a method, system, device, and medium for generating molecular design constraints based on interpretable artificial intelligence. Background Technology
[0002] In the fields of medicinal chemistry, pesticide chemistry, and food flavoring, the efficient discovery of novel molecular structures with specific biological activities or functions has always been a core challenge. Traditional molecular design processes rely heavily on the experience and intuition of experts in the field, rationally modifying and optimizing structures by analyzing the structural differences between known active and inactive molecules. This approach is not only time-consuming and costly, but also severely limited by individual knowledge background, making it difficult to systematically explore the vast chemical space.
[0003] In recent years, with the development of artificial intelligence technology, various predictive models have provided new possibilities for molecular design. However, these models still have significant limitations in practical applications—most of them can only provide a judgment of "whether it is active" but cannot further answer two deep questions that are crucial to design work: first, "why is this molecule predicted to be active," and second, "based on this, how should new active molecules be designed?" The lack of these two questions makes it difficult for the models to be truly integrated into the core decision-making process of molecular design. Summary of the Invention
[0004] This invention provides a method, system, device, and medium for generating molecular design constraints based on interpretable artificial intelligence, in order to overcome the shortcomings of existing technologies.
[0005] This invention provides a method for generating molecular design constraints based on interpretable artificial intelligence, comprising:
[0006] Obtain molecular structure data for the target molecule type, wherein the molecular structure data is tagged with active / inactive tags;
[0007] Based on molecular structure data, molecular characteristics are obtained, including molecular property characteristics and / or molecular fingerprint structural characteristics.
[0008] Based on the activity / inactivity labels of molecular characteristics and molecular structure data, machine learning algorithms are used to train a model to learn the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model.
[0009] The molecular activity prediction model is applied to predict the molecular activity of the target molecule type. The molecular activity prediction results of the molecular activity prediction model are analyzed for interpretability. Based on the results of the interpretability analysis, molecular design constraints for the target molecule are generated. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction.
[0010] According to the present invention, a method for generating molecular design constraints based on interpretable artificial intelligence is provided. The molecular property characteristics include any one or any combination of the following: molecular weight, lipid-water partition coefficient, number of hydrogen bond donors, number of hydrogen bond acceptors, topological polar surface area, number of rotatable bonds, ring system characteristics, drug similarity score, synthetic accessibility score, natural product similarity score, total number of heavy atoms, and proportion of CSP3 hybrid carbon atoms.
[0011] According to the present invention, a method for generating molecular design constraints based on interpretable artificial intelligence is provided, wherein the molecular fingerprint structural feature is a 166-bit MACCS molecular fingerprint.
[0012] According to the present invention, a method for generating molecular design constraints based on interpretable artificial intelligence is provided, wherein the method uses a machine learning algorithm to train a model to learn the molecular characteristics of active and inactive molecules based on the activity / inactivity labels of molecular features and molecular structure data, thereby obtaining a molecular activity prediction model, comprising:
[0013] When molecular features include molecular property features and molecular fingerprint structural features, the molecular property features and molecular fingerprint structural features are fused to obtain molecular fusion features;
[0014] Based on the active / inactive labels of molecular fusion characteristics and molecular structure data, a machine learning algorithm is used to train a model to learn the molecular fusion characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model.
[0015] According to the present invention, a method for generating molecular design constraints based on interpretable artificial intelligence is provided, wherein the method uses a machine learning algorithm to train a model to learn the molecular characteristics of active and inactive molecules based on the activity / inactivity labels of molecular features and molecular structure data, thereby obtaining a molecular activity prediction model, comprising:
[0016] Active / inactive tags based on molecular characteristics and molecular structure data
[0017] By using the random forest algorithm and combining multiple random seeds, multiple models are trained to learn the molecular characteristics of active and inactive molecules, resulting in multiple molecular activity prediction models.
[0018] The generalization performance of multiple molecular activity prediction models was evaluated.
[0019] According to the present invention, a method for generating molecular design constraints based on interpretable artificial intelligence is provided, wherein the method involves applying a molecular activity prediction model to predict the molecular activity of a target molecule type, performing interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generating molecular design constraints for the target molecule based on the interpretability analysis results, including:
[0020] The interpretability analysis of the molecular activity prediction results of the molecular activity prediction model was performed using the Shapley additive interpretation method (SHAP) to obtain the global importance score of each molecular feature for predicting the activity of the molecule to be tested.
[0021] Key molecular features are selected based on the global importance score of each molecular feature;
[0022] For the molecular property features among the key molecular features, the kernel density estimation method is used to obtain the numerical optimization range of the molecular property features;
[0023] For the molecular fingerprint structure features in the key molecular features, the contribution direction of the fingerprint site to the prediction of the activity of the target molecule is obtained based on the average SHAP value of the fingerprint site with "1" and the fingerprint site and its contribution direction are generated into a natural language description using a text dictionary with preset constraints.
[0024] According to the present invention, a method for generating molecular design constraints based on interpretable artificial intelligence is provided, wherein the method involves applying a molecular activity prediction model to predict the molecular activity of a target molecule type, performing interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generating molecular design constraints for the target molecule based on the interpretability analysis results, including:
[0025] Based on the interpretability analysis results, a pre-set template is used to generate structured molecular design constraints for the molecule to be tested. The pre-set template includes the following parts:
[0026] Task definition: Clearly specify the type of molecules to be designed;
[0027] Property constraints: List each molecular property feature and its numerical optimization range in the key molecular features in the form of an item list;
[0028] Structural recommendations: List the chemical structural modifications or avoidance items extracted from the molecular fingerprint structural features of each key molecular feature in the form of an item list;
[0029] Output format specification: The generated molecules are required to be output as a standard SMILES string.
[0030] This invention also provides a molecular design constraint generation system based on interpretable artificial intelligence, comprising:
[0031] The data acquisition module is used to: acquire molecular structure data of the target molecule type, wherein the molecular structure data is tagged with active / inactive tags;
[0032] The data processing module is used to: obtain molecular features based on molecular structure data, wherein the molecular features include molecular property features and / or molecular fingerprint structural features;
[0033] The model training module is used to: train a model to learn the molecular characteristics of active and inactive molecules based on the activity / inactivity labels of molecular features and molecular structure data, using machine learning algorithms, and obtain a molecular activity prediction model.
[0034] The molecular design constraint generation module is used to: apply the molecular activity prediction model to the molecular activity prediction of the target molecular type of the molecule, perform interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generate molecular design constraints for the molecule based on the interpretability analysis results. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction.
[0035] The present invention also provides an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the computer program to implement any of the above-described methods for generating molecular design constraints based on interpretable artificial intelligence.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for generating molecular design constraints based on interpretable artificial intelligence.
[0037] The present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute any of the above-described methods for generating molecular design constraints based on interpretable artificial intelligence.
[0038] The present invention provides a method, system, device, and medium for generating molecular design constraints based on interpretable artificial intelligence, which can bring at least the following beneficial effects:
[0039] This invention introduces interpretability analysis to deeply analyze the output of molecular activity prediction models. It automates and systematically transforms the black-box predictions of machine learning models into specific, actionable molecular design constraints. This clearly informs designers which molecular properties (such as molecular weight, lipid-water partition coefficient, etc.) or molecular fingerprint structural features play a key role in activity, and how these features should be optimized (such as numerical range and contribution direction). This truly integrates artificial intelligence into the core decision-making process of molecular design, significantly reducing reliance on the personal experience and intuition of domain experts.
[0040] This invention, based on interpretable analysis results, can generate structured molecular design constraints. Through pre-set templates, it can transform complex model analysis results into clear, itemized design instructions containing "task definitions," "property constraints" (such as molecular weight optimization ranges), and "structural suggestions" (such as specific chemical modification suggestions). This explicit and specific output format can directly provide an actionable optimization path for molecular design work in fields such as medicinal chemistry and pesticide chemistry, avoiding the inefficient process of blind trial and error in traditional design, thereby effectively improving the efficiency and success rate of molecular design.
[0041] This invention's method can flexibly adapt to different molecular feature input modes, including using only molecular property features, only molecular fingerprint structure features, or a dual-modal feature input that fuses both. Under different feature input modes, this invention can stably generate clear and specific molecular design hints based on SHAP interpretability analysis. In particular, when using a dual-modal feature input that fuses molecular property features and molecular fingerprint structure features, the molecular activity prediction model of this invention exhibits superior prediction performance, and the resulting molecular design constraints are more comprehensive, providing design guidance from both macroscopic physicochemical properties and microscopic structural fragments, thus providing powerful automated tool support for efficient and rational molecular design. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0043] Figure 1 This is one of the flowcharts illustrating a method for generating molecular design constraints based on interpretable artificial intelligence, as provided by the present invention.
[0044] Figure 2This is the second flowchart illustrating a method for generating molecular design constraints based on interpretable artificial intelligence, as provided by the present invention.
[0045] Figure 3 This is the performance distribution of the model trained in Embodiment 1 of the present invention.
[0046] Figure 4 This is the performance distribution of the model trained in Embodiment 2 of the present invention.
[0047] Figure 5 This is the performance distribution of the model trained in Embodiment 3 of the present invention.
[0048] Figure 6 Example of molecular generation constraint text (prompt words) generated for embodiments of the present invention.
[0049] Figure 7 Example of using the molecular generation constraint text (prompt words) generated for embodiments of the present invention.
[0050] Figure 8 The chart shows the efficiency and success rate of molecule generation under constraints provided in the embodiments of the present invention. AI model 1 is Deepseek-R1, AI model 2 is GPT-4o, and AI model 3 is GPT-5. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0052] Figure 1 and Figure 2 This is a flowchart illustrating a method for generating molecular design constraints based on interpretable artificial intelligence, as provided by this invention. The execution entity of this method can be any applicable terminal-side device or network-side device, such as a molecular design constraint generation device based on interpretable artificial intelligence.
[0053] See Figure 1 and Figure 2 The present invention provides a method for generating molecular design constraints based on interpretable artificial intelligence, which may include:
[0054] S110. Obtain molecular structure data of the target molecule type (which can be in the form of SMILES string), wherein the molecular structure data is marked with an active / inactive tag.
[0055] In one embodiment, S110 can perform efficiency verification on molecular structure data, specifically by performing efficiency checks on the format and chemical rules of the SMILES expression for each molecule, and screening and retaining the molecular data that pass the verification.
[0056] S120. Based on the molecular structure data, obtain the molecular characteristics, which include molecular property characteristics and / or molecular fingerprint structural characteristics.
[0057] In one embodiment, the molecular property characteristics include any one or any combination of the following: molecular weight, lipid-water partition coefficient, number of hydrogen bond donors, number of hydrogen bond acceptors, topological polar surface area, number of rotatable bonds, ring system characteristics, drug similarity score, synthetic accessibility score, natural product similarity score, total number of heavy atoms, and proportion of CSP3 hybridized carbon atoms. The molecular property characteristics can be calculated using existing methods.
[0058] In one embodiment, the molecular fingerprint structure is characterized as a 166-bit MACCS molecular fingerprint.
[0059] S130. Based on the activity / inactivity labels of molecular characteristics and molecular structure data, a machine learning algorithm is used to train the model to learn the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model.
[0060] In one embodiment, when the molecular features include molecular property features and molecular fingerprint structure features, the molecular property features and molecular fingerprint structure features are fused (the fusion method can be vector concatenation) to obtain molecular fusion features; then, based on the molecular fusion features and the activity / inactivity labels of the molecular structure data, a machine learning algorithm is used to train a model to learn the molecular fusion features of active and inactive molecules to obtain a molecular activity prediction model.
[0061] In one embodiment, S130 can divide the activity / inactivity labels of molecular feature and molecular structure data into a training set, a validation set, and a test set. The sample size ratio among the training set, validation set, and test set can be customized through a configuration interface. The dataset partitioning supports two optional execution modes: a first mode is a random partitioning mode, and a second mode is a hierarchical partitioning mode, the latter ensuring that the proportion of different categories of samples in each subset remains consistent with the original dataset.
[0062] In one embodiment, S130 employs a random forest classification algorithm for model training. During training, the hyperparameter space of the model is traversed and optimized using a grid search method to obtain the optimal combination of hyperparameters. The hyperparameters include at least: the number of decision trees, the minimum number of samples required for further partitioning of internal nodes, and the maximum depth of the decision trees. The search range for each hyperparameter can be customized according to the specific task.
[0063] Specifically, S130 may include:
[0064] Based on the activity / inactivity labels of molecular features and molecular structure data, the random forest algorithm, combined with multiple random seeds, is used to train N models to learn the molecular features of active and inactive molecules, resulting in N molecular activity prediction models. N can be customized according to the actual situation. In this embodiment, 1≤N≤10. In this embodiment, the same data partitioning rules and hyperparameter search strategy will be used to repeatedly perform dataset partitioning, hyperparameter optimization and model training N times with different random seeds, thereby generating N independent molecular activity prediction models.
[0065] The generalization performance of multiple molecular activity prediction models was evaluated.
[0066] Multiple independent prediction models are trained to assess the robustness of feature importance and avoid overfitting or bias in a single model. The S130 performs rigorous performance evaluation on each trained model, providing a series of metrics such as accuracy, precision, recall, F1 score, and ROC-AUC value to ensure the reliability of model predictions.
[0067] S140. Apply the molecular activity prediction model to predict the molecular activity of the target molecule type, perform interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generate molecular design constraints for the target molecule based on the interpretability analysis results. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction.
[0068] In one embodiment, S140 may include:
[0069] The Shapley additive interpretation method (SHAP) was used to analyze the interpretability of the molecular activity prediction results of the molecular activity prediction model, and the global importance score of each molecular feature for predicting the activity of the test molecule was obtained. Specifically, the SHAP value of each molecular feature of all test molecules based on the molecular activity prediction results of the molecular activity prediction model can be calculated first. Then, the absolute average or norm of the SHAP values of the molecular feature of all test molecules can be taken to obtain the global importance score of each molecular feature (that is, the process of obtaining the global importance score can be: for each molecular activity prediction model, the SHAP method is first used to analyze and obtain the SHAP value of each molecular feature of each sample, and then the global importance score of each molecular feature is obtained: global SHAP feature importance = average absolute SHAP value = first take the absolute value of each SHAP value of each sample, and then take the average of the absolute values of the SHAP values of the feature of all samples).
[0070] The global importance scores of each molecular feature are sorted in descending order to obtain a feature importance ranking list (which can be visualized, for example, in the form of a SHAP summary chart), and the top K (5 in this example) key molecular features are selected.
[0071] For the molecular property features among the key molecular features, extract the numerical distribution of the molecular property features, and use the kernel density estimation method to calculate the numerical range of the molecular property features at a specified confidence level (the range is recorded in the format of [lower limit value - upper limit value]), which is used as the numerical optimization range of the molecular property features.
[0072] For the molecular fingerprint structure features in the key molecular features, the contribution direction of the fingerprint site to the predicted activity of the target molecule is obtained based on the mean SHAP value of the fingerprint site with a value of "1" (a positive mean is considered a positive contribution, and a negative mean is considered a negative contribution). The fingerprint site and its contribution direction are mapped using a pre-set text dictionary containing constraints and descriptions of chemical structures to generate a natural language description.
[0073] In one embodiment, S140 may include:
[0074] Based on the interpretability analysis results, a pre-set template is used to generate structured molecular design constraints for the molecule to be tested. The pre-set template includes the following parts:
[0075] Task definition: Clearly specify the type of molecules to be designed;
[0076] Property constraints: List each molecular property feature and its numerical optimization range in the key molecular features in the form of an item list;
[0077] Structural recommendations: List the chemical structural modifications or avoidance items extracted from the molecular fingerprint structural features of each key molecular feature in the form of an item list;
[0078] Output format specification: The generated molecules are required to be output as a standard SMILES string.
[0079] In one embodiment, S140 may include:
[0080] The system outputs molecular design constraints to the user terminal and provides a one-click copy function, allowing users to directly use them for subsequent large language model-driven molecular generation or as a reference for artificial design. Users can fine-tune parameters such as the number of molecules and confidence level in the molecular design constraints according to their actual needs.
[0081] The following three embodiments will further illustrate the molecular design constraint generation method based on interpretable artificial intelligence provided by the present invention.
[0082] Example 1
[0083] This embodiment demonstrates the process of training and analyzing a model using only molecular property features and generating design prompts based on property constraints. The training dataset used is the publicly available sweet molecule dataset (bittersweet).
[0084] 1. Data preparation: A dataset containing 1241 active (sweet) molecules and 1102 inactive molecules was obtained from a public database. All molecules were represented in SMILES format and had clear activity tags.
[0085] 2. Feature Selection and Calculation: Through the system configuration interface, users can specify that only molecular property features are used. The system automatically calculates 13 core physicochemical and structural properties, forming a feature set, including: molecular weight, lipid-water partition coefficient, number of hydrogen bond donors, number of hydrogen bond acceptors, topological polar surface area, number of rotatable bonds, number of rings, number of aromatic rings, drug similarity score, synthetic accessibility score, natural product similarity score, total number of heavy atoms, and the proportion of CSP3 hybridized carbon atoms.
[0086] 3. Model Training and Evaluation: The data was divided into training, validation, and test sets in a 7:1:2 ratio, and a stratified random partitioning method was used to ensure a consistent proportion of samples in each class. The number of independent models was set to N=10, and the hyperparameters of the random forest were optimized using a grid search method, ultimately resulting in 10 independent random forest classification models.
[0087] 4. Model performance: such as Figure 3As shown, the average ROC-AUC of the 10 models on the test set was 0.93 ± 0.01, the average accuracy was 0.85 ± 0.01, and the average F1 score was 0.86 ± 0.01, indicating that the model based solely on physicochemical properties can robustly distinguish between sweet and non-sweet molecules.
[0088] 5. Interpretability Analysis: SHAP analysis was performed on all 10 models, and the average global importance score for each feature was calculated. The results showed that molecular weight and the number of hydrogen bond donors were the two features with the highest importance ranking.
[0089] 6. Cue Word Generation: Based on the above analysis results, design cue words containing quantitative property constraints are automatically generated, such as... Figure 6 As shown.
[0090] Example 2
[0091] The overall process of this embodiment is the same as that of embodiment 1. The difference is that the feature selection and calculation steps only specify the use of MACCS fingerprint structure features.
[0092] Model performance such as Figure 4 As shown, the average ROC-AUC of the 10 models on the test set reached 0.94 ± 0.01, the average accuracy was 0.88 ± 0.01, and the average F1 score was 0.88 ± 0.01, indicating that the model based solely on MACCS fingerprint features can also robustly distinguish between sweet and non-sweet molecules. Based on the key structural features and their contribution directions determined by SHAP analysis, specific structural design constraint texts (hint words) are generated from a predefined constraint mapping dictionary (see Table 1), such as... Figure 6 As shown.
[0093] Table 1 shows a partial example of the correspondence between MACCS key fingerprint features and molecular design hints in this embodiment.
[0094]
[0095] Example 3
[0096] The overall process of this embodiment is the same as that of embodiment 1. The difference is that the feature selection and calculation steps adopt the fusion feature of molecular properties and MACCS fingerprint.
[0097] Model performance such as Figure 5As shown, the average ROC-AUC of the 10 models on the test set was 0.95 ± 0.01, the average accuracy was 0.88 ± 0.01, and the average F1 score was 0.89 ± 0.01. The dual-modal feature fusion model demonstrated the best performance. Based on the SHAP analysis results of the fused features, composite design cue words with both quantitative property constraints and qualitative structural guidance were generated, such as... Figure 6 As shown.
[0098] The comparison of the three embodiments above demonstrates that the method of the present invention can flexibly adapt to different feature input modes (property only, fingerprint only, fusion features), and can generate clear and specific molecular design prompts based on SHAP interpretability analysis. Among them, the bimodal feature fusion (Example 3) shows the best performance in model prediction and generates the most comprehensive prompts, providing powerful automated tool support for efficient and rational molecular design. Figure 7 The method for using the prompt words generated by the method of the present invention. Figure 8 The verification experiments shown further demonstrate that the new molecules generated by the prompt word-driven molecular generation model of this invention have a high efficiency and an acceptable success rate, proving the practical value of the generated prompt words.
[0099] The molecular design constraint generation system based on interpretable artificial intelligence provided by this invention is described below. The molecular design constraint generation system based on interpretable artificial intelligence described below can be referred to in correspondence with the molecular design constraint generation method based on interpretable artificial intelligence described above.
[0100] This invention provides a molecular design constraint generation system based on interpretable artificial intelligence, which may include:
[0101] The data acquisition module is used to: acquire molecular structure data of the target molecule type, wherein the molecular structure data is tagged with active / inactive tags;
[0102] The data processing module is used to: obtain molecular features based on molecular structure data, wherein the molecular features include molecular property features and / or molecular fingerprint structural features;
[0103] The model training module is used to: train a model to learn the molecular characteristics of active and inactive molecules based on the activity / inactivity labels of molecular features and molecular structure data, using machine learning algorithms, and obtain a molecular activity prediction model.
[0104] The molecular design constraint generation module is used to: apply the molecular activity prediction model to the molecular activity prediction of the target molecular type of the molecule, perform interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generate molecular design constraints for the molecule based on the interpretability analysis results. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction.
[0105] This invention provides an electronic device that may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory to execute the following steps:
[0106] Obtain molecular structure data for the target molecule type, wherein the molecular structure data is tagged with active / inactive tags;
[0107] Based on molecular structure data, molecular characteristics are obtained, including molecular property characteristics and / or molecular fingerprint structural characteristics.
[0108] Based on the activity / inactivity labels of molecular characteristics and molecular structure data, machine learning algorithms are used to train a model to learn the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model.
[0109] The molecular activity prediction model is applied to predict the molecular activity of the target molecule type. The molecular activity prediction results of the molecular activity prediction model are analyzed for interpretability. Based on the results of the interpretability analysis, molecular design constraints for the target molecule are generated. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction.
[0110] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0111] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and the computer program being executed by a processor, enabling the computer to perform the following steps:
[0112] Obtain molecular structure data for the target molecule type, wherein the molecular structure data is tagged with active / inactive tags;
[0113] Based on molecular structure data, molecular characteristics are obtained, including molecular property characteristics and / or molecular fingerprint structural characteristics.
[0114] Based on the activity / inactivity labels of molecular characteristics and molecular structure data, machine learning algorithms are used to train a model to learn the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model.
[0115] The molecular activity prediction model is applied to predict the molecular activity of the target molecule type. The molecular activity prediction results of the molecular activity prediction model are analyzed for interpretability. Based on the results of the interpretability analysis, molecular design constraints for the target molecule are generated. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction.
[0116] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0117] Obtain molecular structure data for the target molecule type, wherein the molecular structure data is tagged with active / inactive tags;
[0118] Based on molecular structure data, molecular characteristics are obtained, including molecular property characteristics and / or molecular fingerprint structural characteristics.
[0119] Based on the activity / inactivity labels of molecular characteristics and molecular structure data, machine learning algorithms are used to train a model to learn the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model.
[0120] The molecular activity prediction model is applied to predict the molecular activity of the target molecule type. The molecular activity prediction results of the molecular activity prediction model are analyzed for interpretability. Based on the results of the interpretability analysis, molecular design constraints for the target molecule are generated. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction.
[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating molecular design constraints based on interpretable artificial intelligence, characterized in that, include: Obtain molecular structure data for the target molecule type, wherein the molecular structure data is tagged with active / inactive tags; Based on molecular structure data, molecular characteristics are obtained, including molecular property characteristics and / or molecular fingerprint structural characteristics. Based on the activity / inactivity labels of molecular characteristics and molecular structure data, machine learning algorithms are used to train a model to learn the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model. The molecular activity prediction model is applied to predict the molecular activity of the target molecular type of the test molecule. The interpretability analysis of the molecular activity prediction results of the molecular activity prediction model is performed, and the molecular design constraints of the test molecule are generated based on the interpretability analysis results. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction. The process of applying the molecular activity prediction model to predict the molecular activity of the target molecule type, performing interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generating molecular design constraints for the target molecule based on the interpretability analysis results includes: The interpretability analysis of the molecular activity prediction results of the molecular activity prediction model was performed using the Shapley additive interpretation method, and the global importance score of each molecular feature for predicting the activity of the analyte molecule was obtained. Key molecular features are selected based on the global importance score of each molecular feature; For the molecular property features among the key molecular features, extract the numerical distribution of the molecular property features, and use the kernel density estimation method to calculate the numerical range of the molecular property features at a specified confidence level, which is used as the numerical optimization range of the molecular property features. For the molecular fingerprint features in the key molecular features, the contribution direction of the fingerprint site to the prediction of the activity of the target molecule is obtained, and the fingerprint site and its contribution direction are generated into a natural language description using a text dictionary with preset constraints. Also includes: Based on the interpretability analysis results, a pre-set template is used to generate structured molecular design constraints for the molecule to be tested. The pre-set template includes the following parts: Task definition: Clearly specify the type of molecules to be designed; Property constraints: List each molecular property feature and its numerical optimization range in the key molecular features in the form of an item list; Structural recommendations: List the chemical structural modifications or avoidable items extracted from each molecular fingerprint feature of the key molecular characteristics in the form of an item list; Output format specification: The generated molecules are required to be output as a standard SMILES string.
2. The method for generating molecular design constraints based on interpretable artificial intelligence according to claim 1, characterized in that, Molecular properties include any one or any combination of the following: molecular weight, lipid-water partition coefficient, number of hydrogen bond donors, number of hydrogen bond acceptors, topological polar surface area, number of rotatable bonds, ring system characteristics, drug similarity score, synthetic accessibility score, natural product similarity score, total number of heavy atoms, and proportion of CSP3 hybridized carbon atoms.
3. The method for generating molecular design constraints based on interpretable artificial intelligence according to claim 1, characterized in that, The molecular fingerprint structure is characterized by MACCS molecular fingerprint.
4. The method for generating molecular design constraints based on interpretable artificial intelligence according to any one of claims 1-3, characterized in that, The process involves using activity / inactivity labels based on molecular characteristics and molecular structure data, employing machine learning algorithms to train a model that learns the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model. This includes: When molecular features include molecular property features and molecular fingerprint structural features, the molecular property features and molecular fingerprint structural features are fused to obtain molecular fusion features; Based on the active / inactive labels of molecular fusion characteristics and molecular structure data, a machine learning algorithm is used to train a model to learn the molecular fusion characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model.
5. The method for generating molecular design constraints based on interpretable artificial intelligence according to claim 4, characterized in that, The process involves using activity / inactivity labels based on molecular characteristics and molecular structure data, employing machine learning algorithms to train a model that learns the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model. This includes: The activity / inactivity labels of molecular feature and molecular structure data are divided into training set, validation set and test set using random partitioning or hierarchical partitioning.
6. The method for generating molecular design constraints based on interpretable artificial intelligence according to claim 4, characterized in that, The process involves using activity / inactivity labels based on molecular characteristics and molecular structure data, employing machine learning algorithms to train a model that learns the molecular characteristics of active and inactive molecules, thereby obtaining a molecular activity prediction model. This includes: Based on the activity / inactivity labels of molecular features and molecular structure data, multiple models are trained using the random forest algorithm combined with multiple random seeds to learn the molecular features of active and inactive molecules, resulting in multiple molecular activity prediction models. The generalization performance of multiple molecular activity prediction models was evaluated.
7. The method for generating molecular design constraints based on interpretable artificial intelligence according to claim 1, wherein key molecular features are selected based on the global importance score of each molecular feature, characterized in that... include: All molecular features are sorted in descending order of global importance score, and the top K molecular features are selected as key molecular features.
8. A molecular design constraint generation system based on interpretable artificial intelligence, characterized in that, include: The data acquisition module is used to: acquire molecular structure data of the target molecule type, wherein the molecular structure data is tagged with active / inactive tags; The data processing module is used to: obtain molecular features based on molecular structure data, wherein the molecular features include molecular property features and / or molecular fingerprint structural features; The model training module is used to: train a model to learn the molecular characteristics of active and inactive molecules based on the activity / inactivity labels of molecular features and molecular structure data, using machine learning algorithms, and obtain a molecular activity prediction model. The molecular design constraint generation module is used to: apply the molecular activity prediction model to the molecular activity prediction of the target molecular type of the molecule, perform interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generate molecular design constraints for the molecule based on the interpretability analysis results. The molecular design constraints include molecular property characteristics and their numerical optimization range and / or molecular fingerprint structural characteristics and their contribution direction. The process of applying the molecular activity prediction model to predict the molecular activity of the target molecule type, performing interpretability analysis on the molecular activity prediction results of the molecular activity prediction model, and generating molecular design constraints for the target molecule based on the interpretability analysis results includes: The interpretability analysis of the molecular activity prediction results of the molecular activity prediction model was performed using the Shapley additive interpretation method, and the global importance score of each molecular feature for predicting the activity of the analyte molecule was obtained. Key molecular features are selected based on the global importance score of each molecular feature; For the molecular property features among the key molecular features, extract the numerical distribution of the molecular property features, and use the kernel density estimation method to calculate the numerical range of the molecular property features at a specified confidence level, which is used as the numerical optimization range of the molecular property features. For the molecular fingerprint features in the key molecular features, the contribution direction of the fingerprint site to the prediction of the activity of the target molecule is obtained, and the fingerprint site and its contribution direction are generated into a natural language description using a text dictionary with preset constraints. Also includes: Based on the interpretability analysis results, a pre-set template is used to generate structured molecular design constraints for the molecule to be tested. The pre-set template includes the following parts: Task definition: Clearly specify the type of molecules to be designed; Property constraints: List each molecular property feature and its numerical optimization range in the key molecular features in the form of an item list; Structural recommendations: List the chemical structural modifications or avoidable items extracted from each molecular fingerprint feature of the key molecular characteristics in the form of an item list; Output format specification: The generated molecules are required to be output as a standard SMILES string.
Citation Information
Patent Citations
Structure-property relationship analysis method and electronic equipment based on multidimensional interpretability
CN119785921A