Molecular generation and optimization system based on electrostatic characteristics and structure of protein

By constructing a molecular generation and optimization system based on the electrostatic characteristics and structure of proteins, the problem that the existing model fails to fully consider the electrostatic characteristics of the target protein is solved, and efficient generation and optimization of highly active chemical molecules are achieved, and the drug development cycle is shortened.

CN120473015APending Publication Date: 2025-08-12SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510568827.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing molecular generation model fails to fully consider the electrostatic characteristics of the target protein, resulting in limited molecular generation ability, difficulty in structural optimization, and difficulty in efficiently generating highly active chemical molecules.

Method used

A molecular generation and optimization system based on the electrostatic characteristics and structure of proteins is constructed. Through input modules, data processing modules, fulcrum atom selection modules, coordinate prediction modules, electrostatic feature modules, element type sampling modules and chemical bond type sampling modules, combined with chemical knowledge constraints, molecules are generated and optimized.

Benefits of technology

Based on the consideration of the electrostatic characteristics of target proteins, the generation and optimization of molecules have significantly improved the effectiveness and rationality of molecular generation, shortened the discovery cycle of leading compounds, and improved the efficiency of drug research and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120473015A_ABST
    Figure CN120473015A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of drug design, and particularly relates to a molecular generation and optimization system based on protein electrostatic characteristics and structures. The molecule generation and / or optimization system comprises an input module, a data processing module, a coordinate prediction module, an electrostatic characteristic module, an element type sampling module, a chemical bond type sampling module, a judgment module and an output module, according to the system, electrostatic characteristics and chemical knowledge rules are comprehensively considered, the functions of molecule de novo generation and lead compound optimization can be achieved at the same time, quantitative estimation of molecular effectiveness and drug-likeness and synthetic accessibility indexes are excellent, and meanwhile the fat-water distribution coefficient can be kept within a correct and reasonable range; a dry / wet experiment proves that a good effect is achieved, the discovery period of a lead compound is greatly shortened, and the method has a wide prospect in the field of medicine research and development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug design, and in particular relates to a system for molecule generation and optimization based on protein electrostatic characteristics and structure. Background Art

[0002] The development of new drugs is a complex process with high costs, high risks, and long cycles. It involves multiple steps, such as lead compound discovery, lead compound optimization, preclinical testing, and clinical research.

[0003] In the early stages of drug development, efficiently screening for promising lead compounds from a vast chemical space is a key breakthrough in accelerating the development of new drugs and reducing R&D costs and cycles. Currently, this process is primarily achieved through high-throughput screening of known compound libraries, using methods including physical experimental screening and computer-generated virtual screening. However, the limited structural diversity of existing compound libraries, coupled with repeated screening efforts by various R&D institutions, makes the discovery of active compounds with novel scaffolds increasingly challenging.

[0004] In addition, the structural optimization of lead compounds is also one of the main bottlenecks in current drug development. Lead compound optimization is generally carried out by adding, deleting and replacing groups on the compound, and its main goal is to further improve the properties of the compound in various aspects (mainly the activity of the compound against the corresponding target). However, this process is very costly because it requires the synthesis of a large number of compounds and the corresponding testing. At the same time, this process is highly dependent on the expertise of medicinal chemistry researchers, so this process is very difficult and costly.

[0005] In recent years, with the rapid development of artificial intelligence (AI), deep generative models have brought new technological breakthroughs and development opportunities to drug discovery. Deep generative models can capture the underlying patterns in data distribution. By using a subset of samples in chemical space as a training set, the generated model can be adapted to the actual data distribution, resulting in compounds that closely resemble the data. This approach can effectively guide the exploration of vast chemical space based on existing data, facilitating the discovery of compounds with novel molecular structures and desirable physicochemical properties.

[0006] Compared to traditional structure-based drug design methods, deep generative models are driven entirely by data and algorithms, independent of any predefined chemical rules. They utilize neural networks to learn the probabilistic distribution of structural information from a large amount of data on small molecule compounds or small molecule-protein complex structures, thereby generating potentially active small chemical molecules. Given that drug molecules must bind to specific target proteins in the body to exert their therapeutic effects, innovative deep generative models that fully utilize target protein structural information to generate molecules are being designed. This allows for the design of drug molecules based on arbitrary targets, significantly shortening the time required to acquire and optimize lead compounds.

[0007] AI-driven drug development has become an international research hotspot in medicinal chemistry. The advantages of AI-based generative models in drug discovery are becoming increasingly prominent, primarily due to their ability to deeply explore the compound space and generate medicinal chemical molecules with novel structures and potential activities. This technological breakthrough not only improves the efficiency of drug development but also provides new ideas and methods for addressing bottlenecks in traditional drug development.

[0008] Currently, there are a number of molecular generation models available internationally, such as Pocket2Mol and TargetDiff, which are used to generate molecules within protein binding pockets. Some models are already capable of generating chemically effective molecules, however, their ability to generate biologically active molecules remains limited. This problem may arise for a variety of reasons, with neglecting the electrostatic characteristics of the target protein being a primary factor. Electrostatic characteristics, including electric fields and potentials, are crucial for the interaction between ligands and proteins. Existing molecular generation models do not take electrostatic characteristics into account, thus limiting their molecular generation capabilities. Furthermore, most models are only capable of de novo generation of small molecule compounds and are unable to perform structural optimization of small molecules. However, in general, whether drug design is AI-based or not, it is difficult to directly obtain the final ideal drug molecule in a single step. Structural optimization is a crucial step in medicinal chemistry, and current AI-based molecular generation models rarely perform structural optimization of molecules.

[0009] Therefore, it is urgently needed in this field to construct a molecular generation system that comprehensively considers electrostatic characteristics and chemical knowledge rules and can perform structural optimization. Summary of the Invention

[0010] In response to the deficiencies of the prior art, the present invention provides a system for molecular generation and optimization based on the electrostatic characteristics and structure of proteins, with the aim of efficiently generating highly active chemical molecules.

[0011] The present invention provides a molecule generation or optimization system, comprising:

[0012] The input module is configured to input the three-dimensional structural information of the molecular structure and the charge values of all atoms in the protein pocket; the molecular structure is the protein pocket, or the molecular structure is a complex of the protein pocket and a small molecule fragment;

[0013] a data processing module configured to preprocess and encode the three-dimensional structure information to obtain encoded features;

[0014] A pivot atom selection module is configured to predict pivot atoms using the encoded features through a neural network;

[0015] a coordinate prediction module configured to predict relative coordinates of a new atom relative to the pivot atom using the encoded features, and obtain absolute coordinates of the new atom using the coordinates of the pivot atom, which are the coordinates of the new atom, and the new atom is used to construct the small molecule fragment;

[0016] The electrostatic feature module is configured to calculate the electric field sum and electric potential sum of the charges carried by all atoms in the protein pocket at the coordinates of the new atom, and use a neural network to obtain vector features containing electrostatic features and scalar features containing electrostatic features.

[0017] An element type sampling module is configured to sample the element type of the new atom by using the vector features including the electrostatic features and the scalar features including the electrostatic features through an autoregressive flow model;

[0018] a chemical bond type sampling module configured to sample the type of a new chemical bond connected to the new atom based on an autoregressive flow;

[0019] The judgment module is configured to: determine whether the type of new atoms and new chemical bonds is successfully generated based on preset constraints, and whether to add new atoms to the generated or existing small molecule fragments to form updated small molecule fragments. The system repeatedly runs at least one of the data processing module, the fulcrum atom selection module, the coordinate prediction module, the electrostatic feature module, the element type sampling module, the chemical bond type sampling module, and the output module.

[0020] an output module configured to output a de novo generated molecule or an optimized molecule, wherein the de novo generated molecule or the optimized molecule is the updated small molecule fragment;

[0021] Furthermore, the judgment module integrates a chemical knowledge constraint judgment module, a sampling times judgment module, and an atomic number judgment module;

[0022] The chemical knowledge constraint judgment module is configured as follows:

[0023] Checking whether the valence number of the new atom exceeds the maximum valence number limit of the new atom; if the valence number of the new atom is less than or equal to the maximum valence number of the new atom, the chemical knowledge constraint judgment module determines that the judgment result is "yes", and the type of the new atom and the new chemical bond is successfully generated;

[0024] The sampling times judgment module is configured as follows:

[0025] Determine whether the sampling number is less than a preset threshold; the sampling number is defined as the number of times the chemical bond type sampling module is repeatedly run before the new chemical bond type is successfully sampled. If the sampling number is greater than the preset threshold, the judgment result of the sampling number judgment module is "yes";

[0026] The atomic number determination module is configured as follows:

[0027] Determining whether the number of all atoms in the updated molecular fragment reaches a preset maximum value, if the number of all atoms is greater than the preset maximum value, the determination result of the atom number determination module is "yes";

[0028] The judgment module is configured to:

[0029] The element type of the new atom and the type of the new chemical bond are sent to the chemical knowledge constraint judgment module for judgment. If the chemical knowledge constraint judgment module judges that the result is "yes", the coordinates, element type and new chemical bond type of the new atom are added to the three-dimensional structure information. The three-dimensional structure information is preprocessed by the data processing module. The number of atoms is judged by the atomic number judgment module. Otherwise, the sampling number judgment module is used to perform the sampling number judgment.

[0030] When the determination result of the atom number determination module is "yes", the output module outputs the result; otherwise, the data processing module obtains the encoded feature;

[0031] When the judgment result of the sampling number judgment module is "yes", the new atom is abandoned and the result is outputted by the output module; otherwise, the chemical bond type sampling module is used to resample the new chemical bond type.

[0032] Furthermore, in the input module, the charge values of all atoms in the pocket of the protein are obtained from the three-dimensional structure of the protein using a molecular force field preparation;

[0033] And / or, in the data processing module, the preprocessing step includes: representing the three-dimensional structure information data into a K-nearest neighbor graph, wherein the nodes of the graph are atoms of the protein binding pocket and the small molecule fragment, and the edges of the graph are chemical bonds; when the molecular structure is a protein pocket, the nodes of the graph are atoms of the protein pocket;

[0034] And / or, in the data processing module, the encoding is implemented by an encoder.

[0035] Furthermore, in the pivot atom selection module, the pivot atom is selected according to the molecular structure type in the encoded feature. If the molecular structure type in the encoded feature is a protein pocket, the pivot atom is selected from the atoms in the protein pocket; if the molecular structure type in the encoded feature is a complex of a protein pocket and a small molecule fragment, the pivot atom is selected on the small molecule fragment;

[0036] and / or, when no atom is predicted as a fulcrum atom in the fulcrum atom selection module, outputting the result using an output module;

[0037] And / or, in the fulcrum atom selection module, the neural network is selected from the GDBP neural network.

[0038] Furthermore, in the coordinate prediction module, the step of predicting the relative coordinates of the new atom includes: using a mixture density network to implement a Gaussian mixture model to model the coordinates of the atom, taking the expectations of all Gaussian components as candidates for relative coordinates, selecting coordinates whose distance from the support atom is less than 2 angstroms and greater than 1 angstrom from the candidate relative coordinates, sampling the Gaussian components through a polynomial distribution, and taking the expectations of the Gaussian components of the sampling results as the relative coordinates of the new atom.

[0039] Furthermore, in the electrostatic feature module, the vector features containing electrostatic features and the scalar features containing electrostatic features are obtained using the following methods: based on Coulomb's law, the electric field sum and the potential sum of the charges carried by all atoms in the protein pocket at the coordinates of the predicted new atom are calculated; the electric field sum is input into the first VN-MLP neural network, fused with other vector features in the encoded features, and input into the second VN-MLP neural network again to obtain the vector features containing electrostatic features; the potential sum is input into the first MLP neural network, fused with other scalar features in the encoded features, and input into the second MLP neural network again to obtain the scalar features containing electrostatic features.

[0040] Furthermore, in the element type sampling module, the step of sampling the element type of the new atom includes: sampling from a Gaussian distribution, and obtaining the element type of the new atom after processing the vector features including the electrostatic features and the scalar features including the electrostatic features through a model based on autoregressive flow;

[0041] And / or, the element type is selected from one of C, N, O, F, P, S, Cl, Br, and I.

[0042] Furthermore, the chemical bond type sampling module integrates a geometric constraint module and a sampling module;

[0043] The geometric constraint module is configured to select candidate connecting atoms that form chemical bonds with new atoms from atoms in the small molecule fragment using a geometric constraint mechanism, thereby guiding the sampling module to sample chemical bonds; the geometric constraint mechanism includes a mechanism for selecting atoms within a range of 4 angstroms from the support atom as candidate connecting atoms and a triangular attention mechanism;

[0044] The sampling module is configured to sample from a Gaussian distribution, process the features obtained by the geometric constraint module through a model based on autoregressive flow, and obtain a new chemical bond type, wherein the new chemical bond type is selected from one of a single bond, a double bond, a triple bond and no bond, and the no bond indicates that there is no chemical bond between the new atom and the candidate connecting atom.

[0045] The chemical bond type sampling module is configured to input feature data into a geometric constraint module to obtain candidate connecting atoms and corresponding features; input the candidate connecting atoms and corresponding features into a sampling module to obtain a new chemical bond type connected to the new atom; the feature data refers to the encoded features output by the encoder, the features output by the fulcrum atom selection module, the features output by the coordinate prediction module, and the features output by the element type sampling module;

[0046] If the encoded feature is data of a protein pocket, the new chemical bond type connected to the new atom is sampled as no bond.

[0047] The present invention provides a molecular optimization system, which comprises any of the above-described molecular generation or optimization systems, wherein in the input module, the molecular structure is a complex of a protein pocket and a small molecule fragment, and in the output module, the optimized molecule is output.

[0048] Furthermore, in the input module, the method for obtaining the three-dimensional structure information is selected from at least one of structural biology methods and computational simulation methods; the structural biology method is selected from at least one of X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryogenic electron microscopy; and the computational simulation method is selected from at least one of molecular docking, molecular dynamics simulation, and AI prediction.

[0049] The term "protein pocket" refers to a concave area on the surface or inside a protein, which has a certain spatial structure and chemical environment and can accommodate and bind specific small molecules.

[0050] In the description of the present invention, the use of ordinal terms such as "first" and "second" is only for descriptive purposes to indicate the relative position or order of different technical features, and is not used to limit the importance, function or role of these features. The use of these ordinal terms is to facilitate understanding and explanation of the technical solutions of the present invention and does not represent any form of priority or functional difference. Those skilled in the art should understand that the use of these ordinal terms is to simplify the description and does not constitute any form of limitation on the scope of protection of the present invention.

[0051] The system for molecular generation or optimization of the present invention includes an input module, a data processing module, a pivot atom selection module, a coordinate prediction module, an electrostatic feature module, an element type sampling module, a chemical bond type sampling module, a judgment module, and an output module. The system of the present invention takes into account the electrostatic features of proteins in the process of generating and optimizing chemical molecules, and can simultaneously realize the functions of de novo molecular generation and lead compound optimization. The present invention explicitly considers the role of the electric field of the protein target in molecular generation for the first time; explicitly incorporates chemical domain knowledge into three-dimensional molecular generation to provide knowledge guidance to ensure the effectiveness and rationality of generated molecules; considers the geometric topological knowledge of proteins and ligands to obtain intermolecular interactions; uses transfer learning in model training, thereby alleviating the lack of small molecule-protein data and improving model performance. The system of the present invention performs excellently in quantitatively evaluating drug-like properties, molecular effectiveness, and synthetic accessibility indicators, while maintaining the lipid-water partition coefficient within a correct and reasonable range; dry / wet experiments have verified good results, greatly shortening the discovery cycle of lead compounds, and has broad prospects in the field of drug research and development.

[0052] Obviously, based on the above contents of the present invention, according to common technical knowledge and customary means in this field, without departing from the above basic technical ideas of the present invention, other various forms of modifications, replacements or changes can be made.

[0053] The following further describes the above content of the present invention in detail through specific embodiments in the form of examples. However, this should not be construed as limiting the scope of the above subject matter of the present invention to the following examples. All technologies implemented based on the above content of the present invention fall within the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flow chart of de novo molecule generation by the molecule generation system of Example 1;

[0055] Figure 2 Schematic diagram of the module for modeling electrostatic features;

[0056] Figure 3 This is a flow chart of lead compound optimization by the molecular optimization system of Example 2. DETAILED DESCRIPTION

[0057] It should be noted that the algorithms for data acquisition, transmission, storage, and processing steps not specifically described in the examples, as well as the hardware structures and circuit connections not specifically described, can all be implemented using the publicly available content in the prior art. In the following examples, reagents and materials not specifically described are all commercially available.

[0058] Example 1 A molecule generation system based on protein electrostatic characteristics and structure

[0059] 1. The system of this embodiment is as follows Figure 1 As shown, specifically as follows:

[0060] (1) An input module is configured to input the three-dimensional structural information of the protein pocket and the charge values of all atoms in the protein pocket; the charge values of all atoms in the protein pocket are obtained by using the molecular force field preparation of the three-dimensional structure of the protein.

[0061] (2) Data processing module, which integrates data preprocessing module and encoding module.

[0062] (2.1) A data preprocessing module is configured to represent the three-dimensional structural information data as a K-nearest neighbor graph, wherein the nodes of the graph are the atoms of the protein pockets and small molecule fragments, and the edges of the graph are chemical bonds; the small molecule fragments are obtained using the molecule generation system of this embodiment; when there are no small molecule fragments, the nodes of the graph are the atoms of the protein pockets.

[0063] (2.2) An encoding module is configured to encode the encoded features obtained by the data preprocessing module using an encoder.

[0064] (3) A pivot atom selection module is configured to predict pivot atoms using the encoded features through a GDBP neural network.

[0065] If the encoded feature is data of a protein pocket, the fulcrum atom is selected from the atoms in the protein pocket; if the encoded feature is data of a complex of a protein pocket and a small molecule fragment, the fulcrum atom is selected on the small molecule fragment; if all atoms are not predicted to be fulcrum atoms, generation is stopped and the results are output using an output module.

[0066] (4) A coordinate prediction module is configured to use a mixture density network (Mixture Density Networks) to implement a Gaussian mixture model for the encoded features to model the coordinates of the atoms, take the expectations of all Gaussian components as candidates for the relative coordinates of the new atom (relative coordinates relative to the fulcrum atom), select coordinates with a distance from the fulcrum atom less than 2 angstroms and greater than 1 angstrom from the candidate coordinates, sample the Gaussian components through a polynomial distribution, take the expectations of the Gaussian components of the sampling results as the predicted relative coordinates of the new atom, and obtain the absolute coordinates of the predicted new atom through the coordinates of the fulcrum atom, which are the coordinates of the new atom; the fulcrum atom is obtained using the above-mentioned fulcrum atom selection module;

[0067] The new atoms are used to construct small molecule fragments.

[0068] (5) An electrostatic feature module is configured to calculate the electric field sum and electric potential sum of the charge values of all atoms in the pocket of the protein at the coordinates of the predicted new atom based on Coulomb's law; the electric field sum is input into the first VN-MLP neural network, fused with other vector features in the encoded features, and input into the second VN-MLP neural network again to obtain a vector feature containing an electrostatic feature; the electric potential sum is input into the first MLP neural network, fused with other scalar features in the encoded features, and input into the second MLP neural network again to obtain a scalar feature containing an electrostatic feature (such as Figure 2 ).

[0069] (6) An element type sampling module is configured to sample from a Gaussian distribution, process the vector features containing electrostatic features and the scalar features containing electrostatic features obtained by the electrostatic feature module through an autoregressive flow model, and obtain the element type of the new atom, wherein the element type is selected from one of C, N, O, F, P, S, Cl, Br, and I.

[0070] (7) Chemical bond type sampling module, which integrates the geometric constraint module and the sampling module.

[0071] (7.1) The geometric constraint module is configured to use a geometric constraint mechanism to select candidate connecting atoms that form chemical bonds with new atoms from the atoms in the existing small molecule fragments, and guide the sampling module to sample chemical bonds; there are two geometric constraint mechanisms, one is the mechanism for selecting atoms within 4 angstroms from the support atom as candidate connecting atoms, and the other is the triangular attention mechanism (the triangular attention mechanism can guide the sampling module to sample new chemical bonds).

[0072] The triangular attention mechanism comes from AlphaFold's technology, which uses the constraint relationship between the three sides of the triangle formed by atoms and chemical bonds to constrain the type of chemical bonds.

[0073] (7.2) A sampling module is configured to sample from a Gaussian distribution and process the features obtained from the geometric constraint module through an autoregressive flow-based model to obtain a new chemical bond type; wherein the chemical bond type is selected from one of a single bond, a double bond, a triple bond, and no bond, and the no bond indicates that there is no chemical bond between the new atom and the candidate connecting atom.

[0074] (7.3) The chemical bond type sampling module is configured to input the feature data into the geometric constraint module to obtain candidate connecting atoms and corresponding features, and input the candidate connecting atoms and corresponding features into the sampling module to obtain a new chemical bond type connected to the new atom;

[0075] The feature data refers to the encoded features output by the encoder, the features output by the fulcrum atom selection module, the features output by the coordinate prediction module, and the features output by the element type sampling module.

[0076] If the encoded feature is data of a protein pocket, the new chemical bond type connected to the new atom is directly sampled as no bond.

[0077] (8) Judgment module, which integrates chemical knowledge constraint judgment module, sampling times judgment module, and atomic number judgment module.

[0078] (8.1) A chemical knowledge constraint judgment module is configured to check whether the valence number of the new atom exceeds the maximum valence number of the new atom. If the valence number of the new atom is less than or equal to the maximum valence number of the new atom, it indicates that the chemical rules are met, the type of the new atom and the new chemical bond are successfully generated, and the judgment result of the chemical knowledge constraint judgment module is "yes";

[0079] When the chemical knowledge constraint judgment module determines that the result is "yes", the coordinates of the new atom, the element type, and the type of the new chemical bond are added to the three-dimensional structure information. If the encoded feature is the data of a protein pocket, the coordinates of the new atom, the element type, and the chemical bond are constructed into a small molecule fragment. If the encoded feature is the data of a complex of a protein pocket and a small molecule fragment, the coordinates of the new atom, the element type, and the chemical bond are added to the generated small molecule fragment to form an updated small molecule fragment. Then the data preprocessing module is run, and then the atom number judgment module is run. Otherwise, the sampling number judgment module is run.

[0080] (8.2) an atom number determination module, configured to determine whether the number of all atoms in the small molecule fragment reaches a maximum value preset by the user (e.g., 100). If the number of all atoms in the small molecule fragment is greater than the preset maximum value, the determination result of the atom number determination module is "yes";

[0081] When the determination result of the atom quantity determination module is "yes", the output module is used to output the result; otherwise, the data encoding module is run to continue to try to generate the next new atom.

[0082] (8.3) a sampling number determination module configured to determine whether the sampling number is less than a preset threshold; the sampling number is defined as the number of times the chemical bond type sampling module is repeatedly executed before the new chemical bond type is successfully sampled; if the sampling number is greater than the preset threshold, the determination result of the sampling number determination module is "yes";

[0083] When the judgment result of the sampling number judgment module is "yes", the new atom is abandoned and the output module is used to output the result; otherwise, the chemical bond type sampling module of the chemical bond is used to resample the new chemical bond type.

[0084] (9) An output module configured to output a de novo generated molecule, wherein the de novo generated molecule is the updated small molecule fragment.

[0085] 2. Training of the System in This Example

[0086] To alleviate the problem of insufficient data on small molecule-protein complexes, this example employs a transfer learning strategy. The model is trained using a pre-training-fine-tuning model. First, 10 million drug-like molecules were selected from the ZINC dataset for pre-training. Since the ZINC dataset does not contain proteins, the protein representation section was set to empty. The model was then fine-tuned using the protein-ligand complex data set CrossDocked2020, which contains 150,000 molecules.

[0087] The goal of pre-training is to allow the model to learn the high-dimensional distribution of chemical small molecule data. The goal of fine-tuning is to allow the model to learn the high-dimensional distribution of small molecule-protein complex data. During fine-tuning, the charge of each atom in the protein pocket must be calculated. This is done by applying a molecular force field to the protein to obtain the charge of each atom in the binding pocket with the small molecule.

[0088] In the element type sampling module, a model based on autoregressive flow is used to construct a reversible transformation from the basic distribution to the element type distribution. At the same time, the generation of each element type depends on the previously generated element types. This sequential dependency enables the model to effectively capture the complex relationships between element types.

[0089] In the chemical bond type sampling module, a model based on autoregressive flow is used to construct a reversible transformation from the basic distribution to the chemical bond type distribution; at the same time, the generation of each chemical bond type depends on the previously generated chemical bond type. This sequential dependence enables the model to effectively capture the complex relationship between chemical bond types.

[0090] 3. Effect Verification

[0091] 1. Dry experiment

[0092] The dry experiment test was performed on 10 target proteins commonly used in testing: 1bvr, 1u0f, 1zyu, 2ah9, 2ati, 2hw1, 4bnw, 4i91, 5g3n, 5lvq. The results are shown in Table 1.

[0093] Table 1 Dry test results

[0094]

[0095] Note: All data are the average values of all generated molecules.

[0096] As shown in Table 1, the system of this embodiment performs best in quantitatively assessing drug-likeness, molecular effectiveness, and synthetic accessibility, while maintaining the lipid-water partition coefficient within a reasonable range. These results demonstrate that the molecule generation system of this embodiment, while fully considering the electrostatic characteristics of the target protein, can effectively output new molecules that meet the requirements.

[0097] 2. Wet test

[0098] This example further validated the wet-bed assay using YTHDC2, an important target that has been shown to play a crucial role in various biological processes, including RNA metabolism, cellular function, and development. It is also implicated in the pathology of numerous diseases, such as hepatitis C virus, prostate cancer, and rheumatoid arthritis. However, no potent inhibitors targeting YTHDC2 have been reported.

[0099] In this example, the three-dimensional structural information of the protein pocket of YTHDC2 and the charge of each atom after molecular force field preparation are input. The system of this example outputs a de novo generated molecule a, whose structural formula is shown below:

[0100]

[0101] Wet assay method: Fluorescence polarization experiments were performed in a pH 7.5 buffer containing 20 mM HEPES and 180 mM NaCl. GST-tagged protein (650 nM) was incubated with 7 nM probe (FAM-m6A-mRNA). The FAM-m6A-mRNA sequence is: UUCUUCUGUGG-(m6A)-CUGUG (SEQ ID NO. 1), with a FAM group modified at the 5' end. "m6A" indicates methylation at the 6th nitrogen atom of the adenosine (A) in the RNA molecule. Assays were performed in 384-well black plates, and fluorescence polarization (FP) readings were performed using a CLARIOstar PLUS plate reader. To determine the effect of molecule a on YTHDC2 enzyme activity, molecule a was incubated with protein for 30 minutes, followed by addition of substrate to the assay system and a further incubation of 2 hours. Data were fitted using GraphPad Prism software.

[0102] The results of wet test showed that the IC 50 =16.84 μM. This indicates that the system of this example can effectively produce molecules with high activity against the target protein.

[0103] Example 2 A molecular optimization system based on protein electrostatic characteristics and structure

[0104] 1. System of this embodiment

[0105] The system of Example 1 is configured as follows: Figure 3 As shown, the difference is that the input module and the output module are different.

[0106] The input module of this embodiment is configured to input the three-dimensional structural information of the complex conformation of the small molecule fragment and the protein pocket and the charge values of all atoms in the protein pocket; the small molecule fragment is a lead compound; there are two ways to obtain the complex conformation, one is through structural biology methods, and the other is through computational simulation methods; the structural biology methods are selected from X-ray crystallography, nuclear magnetic resonance spectroscopy, cryogenic electron microscopy, and small-angle X-ray scattering; the computational simulation methods are selected from molecular docking, molecular dynamics simulation, and AI prediction, such as molecular docking series software, GROMACS software, and DiffDock model.

[0107] The output module is configured to output an optimized molecule, where the optimized molecule is the updated small molecule fragment.

[0108] In this example, the three-dimensional structural information of the complex conformation of molecule a and YTHDC2 and the charge of each atom after molecular force field preparation are input. The system of this example outputs the optimized molecule b, whose structural formula is shown below:

[0109]

[0110] 2. Wet test effect verification

[0111] The wet test was carried out according to the method of Example 1.

[0112] The IC of molecule b was determined by wet test. 50 =0.56 μM. IC of molecule b 50 It is significantly lower than molecule a, indicating that the system of this embodiment can optimize molecules with higher activity based on existing molecules.

[0113] Through the above embodiments and experimental examples, it can be seen that the present invention explicitly considers the role of the electrostatic characteristics of protein pockets in molecule generation for the first time; explicitly adds chemical field knowledge in three-dimensional molecule generation, provides knowledge guidance, and ensures the effectiveness and rationality of generated molecules; the system of the present invention takes into account the geometric topological knowledge of proteins and ligands to obtain the interactions between molecules; transfer learning is used in model training, thereby alleviating the problem of insufficient small molecule-protein data and improving the performance of the model; it can simultaneously perform molecular de novo generation and lead compound optimization. The system of the present invention performs well in quantitatively evaluating drug-like properties, molecular effectiveness, and synthetic accessibility indicators, while the lipid-water partition coefficient can be maintained within a correct and reasonable range; the good results have been verified by dry / wet experiments, greatly shortening the lead compound discovery cycle, and has broad prospects in the field of drug research and development.

Claims

1. A molecular generation or optimization system, characterized in that include: The input module is configured to input the three-dimensional structural information of the molecular structure and the charge values of all atoms in the pocket of the protein; The molecular structure is a pocket of a protein, or the molecular structure is a complex of a protein pocket and a small molecule fragment; a data processing module configured to preprocess and encode the three-dimensional structure information to obtain encoded features; A pivot atom selection module is configured to predict pivot atoms using the encoded features through a neural network; a coordinate prediction module configured to predict relative coordinates of a new atom relative to the pivot atom using the encoded features, and obtain absolute coordinates of the new atom using the coordinates of the pivot atom, which are the coordinates of the new atom, and the new atom is used to construct the small molecule fragment; The electrostatic feature module is configured to calculate the electric field sum and electric potential sum of the charges carried by all atoms in the protein pocket at the coordinates of the new atom, and use a neural network to obtain vector features containing electrostatic features and scalar features containing electrostatic features. An element type sampling module is configured to sample the element type of the new atom by using the vector features including the electrostatic features and the scalar features including the electrostatic features through an autoregressive flow model; a chemical bond type sampling module configured to sample the type of a new chemical bond connected to the new atom based on an autoregressive flow; The judgment module is configured to: determine whether the type of new atoms and new chemical bonds is successfully generated based on preset constraints, and whether to add new atoms to the generated or existing small molecule fragments to form updated small molecule fragments. The system repeatedly runs at least one of the data processing module, the fulcrum atom selection module, the coordinate prediction module, the electrostatic feature module, the element type sampling module, the chemical bond type sampling module, and the output module. The output module is configured to output a de novo generated molecule or an optimized molecule, wherein the de novo generated molecule or the optimized molecule is the updated small molecule fragment.

2. The system for generating or optimizing molecules according to claim 1, wherein: The judgment module integrates a chemical knowledge constraint judgment module, a sampling number judgment module, and an atomic number judgment module; The chemical knowledge constraint judgment module is configured as follows: Checking whether the valence number of the new atom exceeds the maximum valence number limit of the new atom; if the valence number of the new atom is less than or equal to the maximum valence number of the new atom, the chemical knowledge constraint judgment module determines that the judgment result is "yes", and the type of the new atom and the new chemical bond is successfully generated; The sampling times judgment module is configured as follows: Determine whether the sampling number is less than a preset threshold; the sampling number is defined as the number of times the chemical bond type sampling module is repeatedly run before the new chemical bond type is successfully sampled. If the sampling number is greater than the preset threshold, the judgment result of the sampling number judgment module is "yes"; The atomic number determination module is configured as follows: Determine whether the number of all atoms in the updated molecular fragment reaches a preset maximum value. If the number of all atoms is greater than the preset maximum value, the determination result of the atom number determination module is "yes"; The judgment module is configured to: The element type and the new chemical bond type of the new atom are sent to the chemical knowledge constraint judgment module for judgment. If the chemical knowledge constraint judgment module judges that the result is "yes", the coordinates, element type and the new chemical bond type of the new atom are added to the three-dimensional structure information. The three-dimensional structure information is preprocessed by the data processing module. The atomic number judgment module is used to judge the atomic number. Otherwise, the sampling number judgment module is used to judge the sampling number. When the determination result of the atom number determination module is "yes", the output module outputs the result; otherwise, the data processing module obtains the encoded feature; When the judgment result of the sampling number judgment module is "yes", the new atom is abandoned and the result is outputted by the output module; otherwise, the chemical bond type sampling module is used to resample the new chemical bond type.

3. The system for generating or optimizing molecules according to claim 1, wherein: In the input module, the charge values of all atoms in the pocket of the protein are obtained from the three-dimensional structure of the protein using a molecular force field preparation; And / or, in the data processing module, the preprocessing step includes: representing the three-dimensional structure information data into a K-nearest neighbor graph, where the nodes of the graph are atoms of the protein binding pocket and the small molecule fragment, and the edges of the graph are chemical bonds; When the molecular structure is a pocket of a protein, the nodes of the graph are the atoms of the protein pocket; And / or, in the data processing module, the encoding is implemented by an encoder.

4. The molecular generation or optimization system according to claim 1, characterized in that In the pivot atom selection module, the pivot atom is selected according to the molecular structure type in the encoded feature. If the molecular structure type in the encoded feature is a protein pocket, the pivot atom is selected from the atoms in the protein pocket. If the molecular structure type in the encoded feature is a complex of a protein pocket and a small molecule fragment, the pivot atom is selected on the small molecule fragment. and / or, when no atom is predicted as a fulcrum atom in the fulcrum atom selection module, outputting the result using an output module; And / or, in the fulcrum atom selection module, the neural network is selected from the GDBP neural network.

5. The molecular generation or optimization system according to claim 1, characterized in that In the coordinate prediction module, the step of predicting the relative coordinates of the new atom includes: using a mixture density network to implement a Gaussian mixture model to model the coordinates of the atom, taking the expectations of all Gaussian components as candidates for relative coordinates, selecting coordinates whose distance from the support atom is less than 2 angstroms and greater than 1 angstrom from the candidate relative coordinates, sampling the Gaussian components through a polynomial distribution, and taking the expectations of the Gaussian components of the sampling results as the relative coordinates of the new atom.

6. The molecular generation or optimization system according to claim 1, characterized in that In the electrostatic feature module, the vector features containing electrostatic features and the scalar features containing electrostatic features are obtained using the following methods: based on Coulomb's law, the electric field sum and the electric potential sum of the charges carried by all atoms in the protein pocket at the coordinates of the predicted new atom are calculated; the electric field sum is input into the first VN-MLP neural network, fused with other vector features in the encoded features, and then input into the second VN-MLP neural network again to obtain the vector features containing electrostatic features; the electric potential sum is input into the first MLP neural network, fused with other scalar features in the encoded features, and then input into the second MLP neural network again to obtain the scalar features containing electrostatic features.

7. The molecular generation or optimization system according to claim 1, characterized in that In the element type sampling module, the step of sampling the element type of the new atom includes: sampling from a Gaussian distribution, and processing the vector features including the electrostatic features and the scalar features including the electrostatic features through a model based on autoregressive flow to obtain the element type of the new atom; And / or, the element type is selected from one of C, N, O, F, P, S, Cl, Br, and I.

8. The molecular generation or optimization system according to claim 1, characterized in that The chemical bond type sampling module integrates a geometric constraint module and a sampling module; The geometric constraint module is configured to select candidate connecting atoms that form chemical bonds with new atoms from atoms in the small molecule fragment using a geometric constraint mechanism, thereby guiding the sampling module to sample chemical bonds; the geometric constraint mechanism includes a mechanism for selecting atoms within a range of 4 angstroms from the support atom as candidate connecting atoms and a triangular attention mechanism; The sampling module is configured to sample from a Gaussian distribution, process the features obtained by the geometric constraint module through a model based on autoregressive flow, and obtain a new chemical bond type, wherein the new chemical bond type is selected from one of a single bond, a double bond, a triple bond and no bond, and the no bond indicates that there is no chemical bond between the new atom and the candidate connecting atom. The chemical bond type sampling module is configured to input feature data into a geometric constraint module to obtain candidate connecting atoms and corresponding features; input the candidate connecting atoms and corresponding features into a sampling module to obtain a new chemical bond type connected to the new atom; the feature data refers to the encoded features output by the encoder, the features output by the fulcrum atom selection module, the features output by the coordinate prediction module, and the features output by the element type sampling module; If the encoded feature is data of a protein pocket, the new chemical bond type connected to the new atom is sampled as no bond.

9. A molecular optimization system, characterized in that: The molecular optimization system comprises the molecular generation or optimization system according to any one of claims 1 to 8, wherein, in the input module, the molecular structure is a complex of a protein pocket and a small molecule fragment, and in the output module, the optimized molecule is output.

10. The molecular optimization system according to claim 9, characterized in that: In the input module, the method for obtaining the three-dimensional structure information is selected from at least one of structural biology methods and computational simulation methods; the structural biology method is selected from at least one of X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryogenic electron microscopy; the computational simulation method is selected from at least one of molecular docking, molecular dynamics simulation, and AI prediction.