Drug micromolecule multi-attribute optimization method combining deep learning and evolutionary computation

By combining deep learning and evolutionary computing methods, the drug small molecule optimization problem is modeled as a multi-objective optimization problem. Deep learning models and evolutionary algorithms are used to search for high-quality molecules in continuous latent space, solving the problem of inefficient optimization in the existing technology and achieving efficient multi-attribute optimization.

CN120496674APending Publication Date: 2025-08-15ANHUI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510412225.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing drug molecule optimization methods have huge space in chemical decision-making and rely on researchers' experience, which leads to inefficient optimization and makes it difficult to explore target molecules that meet targeting requirements and drug design principles in a short period of time.

Method used

Combining deep learning and evolutionary calculation methods, the multi-attribute optimization problem of drug small molecules is modeled as a multi-objective optimization problem. The deep learning model is used to encode molecules into continuous latent space, and the molecules are divided into multiple subpopulations through non-dominant sorting and sparseness optimization, and different mating pool selection strategies and offspring generation strategies are implemented to generate high-quality drug small molecules.

Benefits of technology

The efficiency and quality of drug small molecule optimization is improved, ensuring that molecules with good attributes are searched in the vast chemical space, avoiding local optimization traps, and achieving efficient multi-attribute optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496674A_ABST
    Figure CN120496674A_ABST
Patent Text Reader

Abstract

The invention provides a drug micromolecule multi-attribute optimization method combining deep learning and evolutionary computation, which comprises the following steps: S1, modeling a drug micromolecule multi-attribute optimization problem into a multi-objective optimization problem, and screening drug micromolecules meeting required characteristics from a DugBank database to obtain an initial molecular population; s2, fragmenting each drug small molecule of the molecular population, embedding the drug small molecule into a continuous submerged space, and carrying out molecular coding and marking; s3, determining the current optimal sparseness; s4, three different sub-populations are obtained; executing different mating pool selection strategies and offspring molecule generation strategies to obtain an offspring molecule set; s5, decoding the offspring molecular set back to a discrete chemical space, calculating an attribute value of the offspring molecular, mixing the parent molecular population with the offspring molecular population, and executing an environment selection strategy on the mixed population to obtain a next-generation molecular set; and S6, repeating the steps S2-S5 until a termination condition is met, and outputting an optimal molecular set. According to the method, the solving efficiency of the drug micromolecule multi-attribute optimization problem is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug property optimization, and specifically to a multi-attribute optimization method for small drug molecules that combines deep learning and evolutionary computing. Background Art

[0002] The discovery of new drugs is crucial for disease treatment, and molecular optimization is a core component of both chemical synthesis and drug discovery. Through systematic structural modification and functional improvement of lead molecules, the goal is to identify molecular structures that both conform to chemical constraints and possess ideal properties. However, while molecular optimization can improve the physicochemical properties of molecules to a certain extent, the optimized molecules often still exhibit certain pharmacological deficiencies that limit their clinical application. Therefore, molecular optimization in drug development is a challenging task.

[0003] Traditional molecular optimization methods usually divide molecules into a group of molecular fragments, each of which has a specific function in the overall molecular structure. On this basis, researchers rely on quantitative structure-activity relationship (QSAR) models to manually add, delete or replace molecular fragments to gradually optimize the target properties of the molecule. However, this trial-and-error optimization method faces huge challenges: the chemical decision space is usually extremely large, containing millions or even more potential molecular variants. Screening out target molecules that meet both targeting requirements and drug design principles not only requires a lot of time and computing resources, but also relies heavily on the experience of researchers.

[0004] To address this challenge, the introduction of artificial intelligence (AI) technology in recent years has provided new solutions for molecular optimization, enabling the exploration of a wider chemical space in a shorter timeframe. Based on the different molecular encoding methods, current mainstream approaches to multi-attribute optimization for small drug molecules fall into two categories: those based on discrete space search and those based on continuous space search.

[0005] 1. Multi-attribute optimization of small drug molecules based on discrete space search

[0006] Multi-attribute optimization methods for small drug molecules based on discrete space search leverage techniques such as evolutionary algorithms and reinforcement learning to search and optimize directly in discrete space. Based on the properties of the new molecules, molecules with better properties are selected for the next round of search, gradually yielding small drug molecules with improved performance. The advantage of these methods is that they can directly manipulate the discrete structure of molecules, avoiding information loss during the mapping process. However, searching in discrete space is complex and relatively inefficient.

[0007] 2. Multi-attribute optimization of small drug molecules based on continuous space search

[0008] Multi-attribute optimization methods for small drug molecules based on continuous space search utilize machine learning models to map small drug molecules, encoded discretely using molecular graphs or SMILES sequences, from discrete space to a continuous latent space. These methods then use search strategies such as gradient descent and Bayesian optimization to search for new molecular representations in the continuous space. Finally, they decode the newly generated molecules from the continuous space back into discrete chemical space to generate specific molecular structures and perform molecular property evaluation and screening. Compared to discrete space search, these methods offer a relatively stable search process and higher search efficiency. However, continuous space embedding is highly dependent on the quality of machine learning training, and the model's mapping capabilities directly impact the optimization results.

[0009] Therefore, an efficient method for drug molecule optimization is urgently needed. Summary of the Invention

[0010] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a multi-attribute optimization method for small drug molecules that combines deep learning and evolutionary computing to solve the existing problems.

[0011] In order to achieve the above object, the technical solution of the present invention is as follows:

[0012] The present invention is achieved through the following technical solution: a multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing, comprising the following steps:

[0013] S1: The multi-attribute optimization problem of drug small molecules is modeled as a multi-objective optimization problem. For the lead drug small molecules to be optimized, N drug small molecules with similarities with the lead drug small molecules higher than a preset threshold and attribute values at a preset level are screened from the DrugBank database to obtain the initial molecular population;

[0014] S2. fragmenting each drug small molecule in the molecular population, embedding the sequence code of the drug small molecule into a continuous latent space after the processing; obtaining a continuous vector representation of the molecule; and marking the processed fragments with 0 and 1 variables to obtain a mask vector of the molecule;

[0015] S3. performing non-dominated sorting on the current molecular population according to the attribute values of the drug small molecules, and determining the current optimal sparsity based on the continuous vector representation;

[0016] S4. Dividing the molecular population according to the non-dominated sorting result and the optimal sparsity to obtain three different subpopulations; executing different mating pool selection strategies and offspring molecule generation strategies for the three different subpopulations to obtain offspring molecule sets;

[0017] S5, using a variational autoencoder to decode the newly generated offspring from the continuous latent space back to the discrete chemical space; calculating the property values of the offspring molecules, mixing the parent molecular population of S2 with the offspring molecular population of S4 to obtain a mixed population; executing the environmental selection strategy in the genetic algorithm NSGA-II on the mixed population to obtain the next generation molecular set consisting of N molecules;

[0018] S6, determine whether the termination condition is met,

[0019] If so, output the optimal molecular set;

[0020] If not, return to S2.

[0021] Furthermore, N drug small molecules having similarity with the lead drug small molecule higher than a preset threshold and attribute values at a preset level are screened from the DrugBank database to obtain an initial molecular population including:

[0022] S1.1. Use RDKit to calculate the property values of each drug small molecule in the DrugBank database, and calculate the Tanimoto similarity between each drug small molecule and the lead molecule based on the molecular Morgan fingerprint. The drug small molecule properties include drug toxicity, solubility, and drug activity;

[0023] S1.2. Select N drug small molecules from the molecular database DrugBank whose attribute values are at a preset level and whose Tanimoto similarity is higher than a preset threshold to obtain an initial molecular population.

[0024] Furthermore, the S2 specifically includes:

[0025] S2.1. fragmenting each small drug molecule in the molecular population to obtain backbone fragments and non-backbone fragments of the molecule;

[0026] S2.2. Based on the backbone fragments and non-backbone fragments, a variational autoencoder is used to embed the sequence encoding of each small drug molecule in the current molecular population into a continuous latent space to obtain a continuous vector representation of the molecule; and 1 and 0 marking variables are set for the backbone fragments and non-backbone fragments respectively to obtain a mask vector of the molecule.

[0027] Furthermore, the S2.1 specifically includes:

[0028] S2.11. For each drug small molecule in the current molecule population, use the BRICS method in the RDKit toolkit to break the SMILES sequence of the molecule from left to right;

[0029] S2.12. Remove redundant side chains from each molecule to extract the Murcko backbone of the molecule, where the backbone consists of a ring structure and a linker.

[0030] Said S2.2 specifically includes:

[0031] S2.21. Use a variational autoencoder (FragVAE) to encode the molecular fragments of each small drug molecule in the molecular population from a discrete chemical space to a continuous latent space, where the length of the continuous vector is fixed to D.

[0032] S2.22. Design a two-layer encoding representation for a small drug molecule, where the upper layer is a mask vector and the lower layer is a real-valued vector, and the upper and lower layers of the vectors have the same length. Specifically, for any small drug molecule x, use the following formula:

[0033] (x1,…,x i )=(mask1×dec1,…,mask i ×dec i )

[0034] In the formula, x i is the i-th decision variable of molecule x in the continuous latent space, i = 1, 2, ... D, mask i is the i-th binary variable, which determines whether the value of the i-th decision variable is zero, dec i is the i-th real-valued variable, representing the specific value of the i-th decision variable; therefore, any molecule x can be represented by a mask vector Mask and a real-valued vector Dec, that is, x = Mask × Dec;

[0035] S2.23. First, set the mask vector of each drug small molecule to an all-zero vector, and then determine whether each real-valued variable is a continuous latent variable of the backbone fragment. If so, flip the mask variable at the corresponding position to 1; otherwise, keep it at 0.

[0036] Furthermore, the S3 specifically includes:

[0037] S3.1. Perform non-dominated sorting on the current molecule population based on molecule attribute values. Molecules that are better on all objectives have smaller non-dominated frontier numbers. Therefore, based on the dominance relationship, the molecules on the first frontier are divided into the non-dominated molecule set. For all molecules in the non-dominated molecule set, the number of non-zero decision variables is accumulated and the average is calculated as the sparsity of the molecule to obtain the sparsity vector sv.

[0038] S3.2. Perform k-means clustering on the sparsity vector sv. The number of clusters k is determined by the number of unique values in the sparsity vector sv, and the median of these k cluster centers is used as the current optimal sparsity cosparsity.

[0039] Furthermore, the molecular population is divided according to the non-dominated sorting result and the optimal sparsity to obtain three different sub-populations, including:

[0040] According to the distribution of the Pareto frontier, all molecules in the current molecular population that are outside the first frontier are divided into the dominated molecule set;

[0041] Divide all molecules in the non-dominated molecule set into the Winner population;

[0042] For all molecules in the dominated molecule set, the sparsity of the binary vector of the dominated molecules is calculated, and according to the size of the sparsity, the dominated molecules with sparsity less than the current optimal sparsity cosparsity are divided into the Loser1 population, and the dominated molecules with sparsity greater than the current optimal sparsity cosparsity are divided into the Loser2 population.

[0043] Furthermore, for the three different subpopulations, different mating pool selection strategies and offspring molecule generation strategies are executed to obtain offspring molecule sets, which specifically include:

[0044] Using the new offspring molecular generation strategy, different offspring generation strategies are executed according to the subpopulation size: specifically, the three different subpopulations are recorded as Winner population, Loser1 population, and Loser2 population, namely P winner 、P loser1 、P loser2 ;

[0045] State 1: When P loser1 The number of molecules in is greater than P loser2 The number of molecules in and greater than When P loser1 The molecules in execute OpLoser1 operator and then merge P loser2 and P winner , execute the OpWinner operator on the molecules in the merged population;

[0046] State 2: When P loser2 The number of molecules in is greater than P loser1 The number of molecules in and greater than When P loser2 The molecules in execute OpLoser2 operator and then merge P loser1 and P winner , execute the OpWinner operator on the molecules in the merged population;

[0047] When state 3, state 1 and state 2 are not satisfied, merge P loser1 、P loser2 、P winner, and perform the OpWinner operator on the molecules in the merged population;

[0048] Merge the three offspring subpopulations to obtain the offspring molecular set.

[0049] Furthermore, the specific steps of S5 are:

[0050] S5.1. Evaluate the property values of the progeny molecules using the RDKit toolkit;

[0051] S5.2. Mix the parent molecule population of S2 and the child molecule population of S4, perform non-dominated sorting on the molecules according to the attribute values, and calculate the crowding distance; sort them in ascending order according to the non-dominated frontier number, and under the condition that the frontier number is the same, sort them in descending order according to the crowding distance, and finally retain the top N drug small molecules to the next generation molecule population, that is, the next generation molecule set.

[0052] A device for optimizing the multi-attribute properties of small drug molecules by combining deep learning and evolutionary computing, comprising a processor and a memory; the memory is used to store a program; the processor executes the program to implement any of the methods described above.

[0053] A computer-readable storage medium stores a program, wherein the program is executed by a processor to implement any one of the above methods.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] The multi-attribute optimization method for small drug molecules of the present invention, which combines deep learning and evolutionary computing, uses a deep learning model to obtain a continuous latent search space that is smoother than a discrete chemical space, and performs a population-based evolutionary search with strong global exploration capabilities in the continuous latent space to ensure that the optimized small drug molecules have better quality.

[0056] With the help of the sparse optimization algorithm, drug small molecules are divided into multiple sub-populations, forming a co-evolutionary search paradigm in which low-quality drug molecule populations learn from high-quality drug molecule populations. Then, a new evolutionary operator is used to efficiently generate high-quality drug small molecules, ensuring the efficiency of solving the multi-attribute optimization problem of drug small molecules. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The disclosure of the present invention is described with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them:

[0058] Figure 1 This is a flow chart of the multi-attribute optimization method for small drug molecules that combines deep learning and evolutionary computing.

[0059] Figure 2 This is a schematic diagram of the fragmentation of a small drug molecule in an embodiment of the present invention;

[0060] Figure 3 Schematic diagram of the drug small molecule encoding process in an embodiment of the present invention;

[0061] Figure 4 The figure is a flow chart of determining the current optimal sparsity in an embodiment of the present invention. DETAILED DESCRIPTION

[0062] It is easy to understand that according to the technical solution of the present invention, without changing the essential spirit of the present invention, a person skilled in the art can propose a variety of interchangeable structural modes and implementation modes. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical solution of the present invention and should not be regarded as the entire invention or as a limitation or restriction of the technical solution of the present invention.

[0063] The present invention provides a multi-attribute optimization method for small drug molecules that combines deep learning and evolutionary computing. The process is as follows: Figure 1 As shown, the following steps are included:

[0064] S1. Model the multi-attribute optimization problem of drug small molecules into a multi-objective optimization problem. For the lead drug small molecules to be optimized, select N drug small molecules from the DrugBank database whose similarity with the lead drug small molecules is higher than a preset threshold and whose attribute values are at a preset level, and obtain the initial molecular population;

[0065] Specifically, the multi-attribute optimization problem of small drug molecules is modeled as a multi-objective optimization problem, and the mathematical representation of the problem is shown in formula (1).

[0066]

[0067] In formula (1), the decision vector x represents a small drug molecule, Ω represents the search space, and M represents the number of molecular attributes.

[0068] Screening out N drug small molecules from the DrugBank database whose similarity to the lead molecule is higher than a preset threshold and whose attribute values are at a preset level, and obtaining an initial molecular population, including screening out N drug small molecules from databases such as DrugBank that are similar to the lead molecule (i.e., similarity is higher than a preset threshold) and have better attribute values (attribute values are at a preset level), and obtaining an initial molecular population, denoted as P;

[0069] The specific steps are:

[0070] S1.1. Calculate the property values of each drug small molecule in databases such as DrugBank using RDKit, and calculate the Tanimoto similarity between each drug small molecule and the lead molecule based on the molecular Morgan fingerprint. The drug small molecule properties include drug toxicity, solubility, and drug activity;

[0071] S1.2. Select N drug small molecules from the molecular database DrugBank whose attribute values are at a preset level and whose Tanimoto similarity is higher than a preset threshold, to obtain an initial molecular population, denoted as P.

[0072] For example, the Tanimoto similarity ranges from 0 to 1. The closer the value is to 1, the more similar the chemical structures of the two molecules are. The higher the pLC50 for biological activity, the better. The lower the LD50 for toxicity, the better. The higher the logS for solubility, the better.

[0073] Assume that lead molecule A has the following characteristics:

[0074] Structure: Contains benzene ring and carboxyl group;

[0075] Properties: pIC50 = 7.5 (biological activity), LD50 = 1.6 (toxicity), logS (solubility) = 9.5.

[0076] When screening from the DrugBank database:

[0077] Calculate Tanimoto similarity: Screen out molecules with similar structure to lead molecule A (Tanimoto similarity > 0.7).

[0078] Calculate property values: screen out pIC50>7.0, LD50<1.6, logS>9.5.

[0079] Comprehensive screening: Finally, a group of molecules that are similar in structure to the lead molecule A and have better properties such as biological activity, toxicity and solubility are obtained as the initial molecular population.

[0080] S2. fragmenting each drug small molecule in the molecular population, embedding the sequence code of the drug small molecule into a continuous latent space after the processing; obtaining a continuous vector representation of the molecule; and marking the processed fragments with 0 and 1 variables to obtain a mask vector of the molecule;

[0081] 2.1. Fragmenting each drug small molecule in the molecular population to obtain molecular backbone fragments and non-backbone fragments;

[0082] The specific steps are as follows:

[0083] S2.11. For each drug small molecule in the current molecular population, use the BRICS (Breaking of Retrosynthetically Interesting Chemical Substructures, a molecular fragmentation method based on retrosynthetic analysis) method in the RDKit toolkit to break the SMILES (Simplified Molecular Input Line Entry System, a chemical language that uses strings to represent molecular structures) sequence of the molecule from left to right; the fragmentation process is shown in the figure below. Figure 2 As shown; the leader molecule shown in the figure is broken into four fragments.

[0084] S2.12. Remove the redundant side chains of each molecule and extract the Murcko backbone of the molecule, where the backbone consists of a ring structure and a linker.

[0085] S2.2, based on the skeleton fragments and non-skeleton fragments, the sequence encoding of each drug small molecule in the molecular population is embedded into a continuous latent space using a variational autoencoder (FragVAE) to obtain a continuous vector representation of the molecule; and the skeleton fragments and non-skeleton fragments are marked with 1 and 0 variables respectively to obtain a mask vector of the molecule; the encoding process is as follows Figure 3 As shown; for the four fragments obtained by fragmentation, the skeleton fragment mark (Mask) variable is 1, and the non-skeleton fragment is marked as 0, where Dec represents the specific value of the corresponding variable.

[0086] The specific steps are:

[0087] S2.21. Encode the molecular fragments of each small drug molecule in the current molecular population P from the discrete chemical space to the continuous latent space using a variational autoencoder. The length of the continuous vector is fixed to D. For example, the size of D is set to 512 bits.

[0088] S2.22. Design a two-layer encoding representation for drug small molecules, where the upper layer is a mask vector and the lower layer is a real-valued vector, and the lengths of the upper and lower layers are the same. Specifically, for any drug small molecule x, it can be expressed using formula (2):

[0089] (x1,…,x i )=(mask1×dec1,…,mask i ×dec i ) (2)

[0090] In formula (2), x i is the i-th decision variable of molecule x in the continuous latent space, i = 1, 2, ... D, mask iis the i-th binary variable, which determines whether the value of the i-th decision variable is zero, dec i is the i-th real-valued variable, representing the specific value of the i-th decision variable. Therefore, any molecule x can be represented by a mask vector Mask and a real-valued vector Dec, that is, x = Mask × Dec;

[0091] S2.23. First, set the mask vector of each drug small molecule to an all-zero vector, and then determine whether each real-valued variable is a continuous latent variable of the backbone fragment. If so, flip the mask variable at the corresponding position to 1; otherwise, keep it at 0.

[0092] S3, perform non-dominated sorting on the current molecular population according to the attribute values of the drug small molecules, and determine the current optimal sparsity based on the continuous vector representation; the calculation process is as follows Figure 4 As shown, specifically:

[0093] S3.1. Perform a non-dominated sort on the molecule population P based on its attribute values. Molecules that are better on all objectives have smaller non-dominated frontier numbers. Therefore, based on the dominance relationship, the molecules on the first frontier are divided into a non-dominated molecule set. For all molecules in the non-dominated molecule set, the number of non-zero decision variables is accumulated and the average is calculated as the sparsity of the molecule to obtain the sparsity vector sv.

[0094] S3.2. Perform k-means clustering on the sparsity vector sv. The number of clusters k is determined by the number of unique values in the sparsity vector sv, and the median of these k cluster centers is used as the current optimal sparsity cosparsity.

[0095] S4. Dividing the molecular population according to the non-dominated sorting result and the optimal sparsity to obtain three different subpopulations; executing different mating pool selection strategies and offspring molecule generation strategies for the three different subpopulations to obtain offspring molecule sets;

[0096] Specifically, the molecular population is divided according to the non-dominated sorting result and the optimal sparsity to obtain three different sub-populations; they are respectively recorded as Winner population, Loser1 population, and Loser2 population, namely P winner 、P loser1 、P loser2 The specific steps are:

[0097] According to the distribution of the Pareto frontier, all molecules in the current molecular population P that are outside the first frontier are divided into the dominated molecule set;

[0098] Divide all molecules in the non-dominated molecule set into the Winner population;

[0099] For all molecules in the dominated molecule set, the sparsity of the binary vector of the dominated molecules is calculated, and according to the size of the sparsity, the dominated molecules with sparsity less than the current optimal sparsity cosparsity are divided into the Loser1 population, and the dominated molecules with sparsity greater than the current optimal sparsity cosparsity are divided into the Loser2 population.

[0100] According to the mask vector, different mating pool selection strategies and offspring molecule generation strategies are executed for the three different subpopulations to obtain offspring molecule sets; the specific steps are:

[0101] First, the novel progeny molecule generation strategy proposed in the present invention implements different progeny generation strategies according to the size of the subpopulation: specifically, the following:

[0102] State 1: When P loser1 The number of molecules in is greater than P loser2 The number of molecules in and greater than When P loser1 The molecules in execute OpLoser1 operator and then merge P loser2 and P winner , execute the OpWinner operator on the molecules in the merged population;

[0103] State 2: When P loser2 The number of molecules in is greater than P loser1 The number of molecules in and greater than When P loser2 The molecules in execute OpLoser2 operator and then merge P loser1 and P winner , execute the OpWinner operator on the molecules in the merged population;

[0104] When state 3, state 1 and state 2 are not satisfied, merge P loser1 、P loser2 、P winner , and perform the OpWinner operator on the molecules in the merged population;

[0105] The specific details of the OpLoser1 operator, OpLoser2 operator, and OpWinner operator are as follows:

[0106] (1) The OpLoser1 operator first uses the binary league method to obtain the loser1 and P winner Each choice |P loser1| molecules, construct two parent mating pools of equal size, then randomly select a parent molecule from each of the two parent mating pools, and perform uniform crossover on the mask vectors of the parent molecules. If the sparsity of the mask vector obtained after crossover is still less than the current optimal sparsity cosparsity, then calculate the difference between the sparsity of the mask vector obtained after crossover and the current optimal sparsity cosparsity. The degree of sparsity difference will determine how many zero variables in the mask vector obtained after crossover need to be flipped into non-zero variables, that is, the number of flipped zero variables is In the mask vector obtained after the intersection, randomly select Zero variables are set to 1 in their corresponding positions to obtain the modified offspring mask vector. After obtaining the modified offspring mask vector, simulated binary crossover and polynomial mutation are performed on the parent real-valued vector to generate the offspring real-valued vector. Finally, the offspring mask vector and the real-valued vector are combined to obtain the offspring molecule. The above two parent molecules are deleted from the parent mating pool. The parent molecule selection and crossover mutation operations are repeated until both parent mating pools are empty. All offspring molecules are summed to obtain the offspring subpopulation.

[0107] (2) The OpLoser2 operator first uses the binary league method to obtain the loser2 and P winner Each choice |P loser2 | molecules, construct two parent mating pools of equal size, then randomly select a parent molecule from each of the two parent mating pools, and perform uniform crossover on the mask vectors of the parent molecules. If the sparsity of the mask vector obtained after crossover is still greater than the current optimal sparsity cosparsity, then calculate the difference between the sparsity of the mask vector obtained after crossover and the current optimal sparsity cosparsity. The degree of sparsity difference will determine how many non-zero variables in the mask vector obtained after crossover need to be flipped to zero variables, that is, the number of flipped non-zero variables is In the mask vector obtained after the intersection, randomly select Non-zero variables are set to 0 in their corresponding positions to obtain the revised offspring mask vector. After obtaining the revised offspring mask vector, simulated binary crossover and polynomial mutation are performed on the parent real-valued vector to generate the offspring real-valued vector. Finally, the offspring mask vector and the real-valued vector are combined to obtain the offspring molecule. The above two parent molecules are deleted from the parent mating pool. The parent molecule selection and crossover mutation operations are repeated until both parent mating pools are empty. All offspring molecules are summed to obtain the offspring subpopulation.

[0108] (3) The OpWinner operator first uses the binary league method to obtain the merged population P best Select 2|Pbest | molecules, construct a parent mating pool, then randomly select two parent molecules from this mating pool. Perform uniform crossover and bitwise mutation on the mask vectors of the parent molecules, and simulate binary crossover and polynomial mutation on the real-valued vectors of the parent molecules. Finally, combine the offspring mask vector and real-valued vector to obtain the offspring molecules. Delete the two parent molecules from the parent mating pool, repeat the parent molecule selection and crossover mutation operations until both parent mating pools are empty, and sum all offspring molecules to obtain the offspring subpopulation.

[0109] Next, the three offspring subpopulations are merged to obtain the offspring molecular population, which is the offspring molecular set, denoted as P'.

[0110] S5, using the variational autoencoder to decode the newly generated offspring from the continuous latent space back to the discrete chemical space; calculating the property values of the offspring molecules, mixing the parent molecular population of S2 with the offspring molecular population of S4 to obtain a mixed population, executing the environmental selection strategy in the genetic algorithm NSGA-II on the mixed population to obtain a next generation molecular set consisting of N molecules; also denoted as P, and looping back to S2 when necessary;

[0111] Specifically:

[0112] S5.1. Evaluate the property values of the progeny molecules using the RDKit toolkit;

[0113] S5.2. Mix the parent molecule population P and the child molecule set P', perform a non-dominated sort on the molecules based on their attribute values, and calculate the crowding distance. Sort them in ascending order by non-dominated front number. If the front numbers are the same, sort them in descending order by crowding distance. Finally, retain the top N drug small molecules for the next generation of molecules, which is denoted as P.

[0114] S6. Determine whether the termination condition is met (the termination condition includes but is not limited to the number of iterations and the convergence of the attribute value);

[0115] If so, output the optimal molecular set;

[0116] If not, return to S2.

[0117] In the molecular optimization process, the space of chemical decision variables is huge, usually involving a large number of atoms and their chemical bond combinations. At the same time, the backbone fragments that play a decisive role in the key properties of the molecule account for a smaller proportion in the molecular structure, while the non-backbone fragments account for a larger proportion in the molecular structure. Based on this feature, the molecular optimization process usually adopts a strategy: keep the fragment structure of the molecular skeleton unchanged, and systematically optimize the non-backbone fragments and the connection sequences between the backbone fragments and the non-backbone fragments. This strategy has a significant impact on the molecular properties by changing a few key structures, thereby achieving an improvement in the target attribute value. This process has obvious sparse characteristics. Therefore, the idea of the present invention is to regard the molecular optimization problem as a large-scale sparse multi-objective optimization problem, and then use the sparse optimization idea to achieve an efficient solution to the multi-attribute optimization problem of small drug molecules.

[0118] The multi-attribute optimization method for small drug molecules of the present invention, which combines deep learning and evolutionary computing, uses a deep learning model to obtain a continuous latent search space that is smoother than a discrete chemical space, and performs a population-based evolutionary search with strong global exploration capabilities in the continuous latent space to ensure that the optimized small drug molecules have better quality.

[0119] Specifically, the present invention uses a deep learning model (such as a modified variational autoencoder FragVAE) to map a discrete chemical space (such as a SMILES sequence or a molecular graph) to a continuous latent space. This mapping makes the search space smoother and facilitates the operation of the optimization algorithm. The deep learning model can greatly improve the search efficiency in a wide range of chemical spaces, and the search process is relatively more stable.

[0120] In a continuous latent space, an evolutionary algorithm is used for global search. Evolutionary algorithms have strong global exploration capabilities, enabling them to identify high-quality potential molecules across a vast chemical space. A population-based optimization strategy ensures a diverse and global search process, avoiding local optima.

[0121] With the help of the sparse optimization algorithm, drug small molecules are divided into multiple sub-populations, forming a co-evolutionary search paradigm in which low-quality drug molecule populations learn from high-quality drug molecule populations. Then, a new evolutionary operator is used to efficiently generate high-quality drug small molecules, ensuring the efficiency of solving the multi-attribute optimization problem of drug small molecules.

[0122] Based on the non-dominated sorting results and sparsity of the molecules, the drug molecules are divided into multiple subpopulations. The low-quality population learns from the high-quality population, gradually improving the quality of the entire population through information exchange and co-evolution.

[0123] Three novel evolutionary operators (including crossover, mutation, and selection) are used to efficiently generate new drug molecules within subpopulations. These operators effectively combine the strengths of high-quality populations to guide the generation of molecules with improved properties. Through the rational design of evolutionary operators, the diversity of the population is maintained and premature convergence is avoided.

[0124] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this document.

[0125] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0126] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices, or units, or can be an electrical, mechanical, or other form of connection.

[0127] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments herein.

[0128] In addition, the functional units in the various embodiments herein may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0129] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this article is essentially or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this article. The aforementioned storage medium includes: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0130] This article uses specific embodiments to illustrate the principles and implementation methods of this article. The description of the above embodiments is only used to help understand the methods and core ideas of this article. At the same time, for those skilled in the art, based on the ideas of this article, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation to this article.

Claims

1. A multi-attribute optimization method for small drug molecules that combines deep learning and evolutionary computing, characterized by: Including the following step: S1: The multi-attribute optimization problem of drug small molecules is modeled as a multi-objective optimization problem. For the lead drug small molecules to be optimized, N drug small molecules with similarities with the lead drug small molecules higher than a preset threshold and attribute values at a preset level are screened from the DrugBank database to obtain the initial molecular population; S2. fragmenting each drug small molecule in the molecular population, embedding the sequence code of the drug small molecule into a continuous latent space after the processing; obtaining a continuous vector representation of the molecule; and marking the processed fragments with 0 and 1 variables to obtain a mask vector of the molecule; S3. performing non-dominated sorting on the current molecular population according to the attribute values of the drug small molecules, and determining the current optimal sparsity based on the continuous vector representation; S4. Dividing the molecular population according to the non-dominated sorting result and the optimal sparsity to obtain three different subpopulations; executing different mating pool selection strategies and offspring molecule generation strategies for the three different subpopulations to obtain offspring molecule sets; S5, using a variational autoencoder to decode the newly generated offspring from the continuous latent space back to the discrete chemical space; calculating the property values of the offspring molecules, mixing the parent molecular population of S2 with the offspring molecular population of S4 to obtain a mixed population; executing the environmental selection strategy in the genetic algorithm NSGA-II on the mixed population to obtain the next generation molecular set consisting of N molecules; S6, determine whether the termination condition is met, If so, output the optimal molecular set; If not, return to S2.

2. The multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing according to claim 1, characterized in that: The method of screening N drug small molecules from the DrugBank database, whose similarity to the lead drug small molecule is higher than a preset threshold and whose attribute values are at a preset level, to obtain an initial molecular population including: S1.

1. Use RDKit to calculate the property values of each drug small molecule in the DrugBank database, and calculate the Tanimoto similarity between each drug small molecule and the lead molecule based on the molecular Morgan fingerprint. The drug small molecule properties include drug toxicity, solubility, and drug activity; S1.

2. Select N drug small molecules from the molecular database DrugBank whose attribute values are at a preset level and whose Tanimoto similarity is higher than a preset threshold to obtain an initial molecular population.

3. The multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing according to claim 1, characterized in that: The S2 specifically includes: S2.

1. fragmenting each small drug molecule in the molecular population to obtain backbone fragments and non-backbone fragments of the molecule; S2.

2. Based on the backbone fragments and non-backbone fragments, a variational autoencoder is used to embed the sequence encoding of each small drug molecule in the current molecular population into a continuous latent space to obtain a continuous vector representation of the molecule; and 1 and 0 marking variables are set for the backbone fragments and non-backbone fragments respectively to obtain a mask vector of the molecule.

4. The multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing according to claim 3, characterized in that: The S2.1 specifically includes: S2.

11. For each drug small molecule in the current molecule population, use the BRICS method in the RDKit toolkit to break the SMILES sequence of the molecule from left to right; S2.

12. Remove redundant side chains from each molecule to extract the Murcko backbone of the molecule, where the backbone consists of a ring structure and a linker. Said S2.2 specifically includes: S2.

21. Use a variational autoencoder (FragVAE) to encode the molecular fragments of each small drug molecule in the molecular population from a discrete chemical space to a continuous latent space, where the length of the continuous vector is fixed to D. S2.

22. Design a two-layer encoding representation for a small drug molecule, where the upper layer is a mask vector and the lower layer is a real-valued vector, and the upper and lower layers of the vectors have the same length. Specifically, for any small drug molecule x, use the following formula: (x1,…,x i )=(mask1×dec1,…,mask i ×Dec i ) In the formula, x i is the i-th decision variable of molecule x in the continuous latent space, i = 1, 2, ... D, mask i is the i-th binary variable, which determines whether the value of the i-th decision variable is zero, dec i is the i-th real-valued variable, representing the specific value of the i-th decision variable; therefore, any molecule x can be represented by a mask vector Mask and a real-valued vector Dec, that is, x = Mask × Dec; S2.

23. First, set the mask vector of each drug small molecule to an all-zero vector, and then determine whether each real-valued variable is a continuous latent variable of the backbone fragment. If so, flip the mask variable at the corresponding position to 1; otherwise, keep it at 0.

5. The multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing according to claim 1, characterized in that: The S3 specifically includes: S3.

1. Perform non-dominated sorting on the current molecule population based on molecule attribute values. Molecules that are better on all objectives have smaller non-dominated frontier numbers. Therefore, based on the dominance relationship, the molecules on the first frontier are divided into the non-dominated molecule set. For all molecules in the non-dominated molecule set, the number of non-zero decision variables is accumulated and the average is calculated as the sparsity of the molecule to obtain the sparsity vector sv. S3.

2. Perform k-means clustering on the sparsity vector sv. The number of clusters k is determined by the number of unique values in the sparsity vector sv, and the median of these k cluster centers is used as the current optimal sparsity cosparsity.

6. The multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing according to claim 1, characterized in that: The molecular population is divided according to the non-dominated sorting result and the optimal sparsity to obtain three different sub-populations, including: According to the distribution of the Pareto frontier, all molecules in the current molecular population that are outside the first frontier are divided into the dominated molecule set; Divide all molecules in the non-dominated molecule set into the Winner population; For all molecules in the dominated molecule set, the sparsity of the binary vector of the dominated molecules is calculated, and according to the size of the sparsity, the dominated molecules with sparsity less than the current optimal sparsity cosparsity are divided into the Loser1 population, and the dominated molecules with sparsity greater than the current optimal sparsity cosparsity are divided into the Loser2 population.

7. The multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing according to claim 6, characterized in that: The method of executing different mating pool selection strategies and offspring molecule generation strategies for the three different subpopulations to obtain offspring molecule sets specifically includes: Using the new offspring molecular generation strategy, different offspring generation strategies are executed according to the subpopulation size: specifically, the three different subpopulations are recorded as Winner population, Loser1 population, and Loser2 population, namely P winner 、P loser1 、P loser2 ; State 1: When P loser1 The number of molecules in is greater than P loser2 The number of molecules in and greater than When P loser1 The molecules in execute OpLoser1 operator and then merge P loser2 and P winner , execute the OpWinner operator on the molecules in the merged population; State 2: When P loser2 The number of molecules in is greater than P loser1 The number of molecules in and greater than When P loser2 The molecules in execute OpLoser2 operator and then merge P loser1 and P winner , execute the OpWinner operator on the molecules in the merged population; When state 3, state 1 and state 2 are not satisfied, merge P loser1 、P loser2 、P winner , and perform the OpWinner operator on the molecules in the merged population; Merge the three offspring subpopulations to obtain the offspring molecular set.

8. The multi-attribute optimization method for small drug molecules combining deep learning and evolutionary computing according to claim 1, characterized in that: The specific steps of S5 are: S5.

1. Evaluate the property values of the progeny molecules using the RDKit toolkit; S5.

2. Mix the parent molecule population of S2 and the child molecule population of S4, perform non-dominated sorting on the molecules according to the attribute values, and calculate the crowding distance; sort them in ascending order according to the non-dominated frontier number, and under the condition that the frontier number is the same, sort them in descending order according to the crowding distance, and finally retain the top N drug small molecules to the next generation molecule population, that is, the next generation molecule set.

9. A multi-attribute optimization device for small drug molecules that combines deep learning and evolutionary computing, characterized by: The method comprises a processor and a memory; the memory is used to store a program; the processor executes the program to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • High-throughput screening compound molecular similarity identification optimization method

    CN121687283A