Novel drug molecule design method and device fusing convex optimization and evolutionary learning
By combining convex optimization and evolutionary learning methods, combined with GAN and convolutional neural networks, the challenges of the existing technology in generating high-quality molecules are solved, and small molecules with high similarity, greater effectiveness and innovation are generated with the original samples, improving the efficiency and quality of molecular generation.
Patent Information
- Application Number
- CN202510010311.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing molecular generation model has challenges in generating high-quality molecules, especially in high-dimensional chemical spaces, which are difficult to balance the diversity and practicality of molecules, and cannot effectively take into account the drug properties, diversity and innovation of molecules.
A new drug molecule design method that combines convex optimization and evolutionary learning is adopted, combined with generative adversarial networks (GANs) and convolutional neural networks based on convex optimization improvements, small molecules with high similarity, greater effectiveness and innovation to the original samples are generated through technologies such as word embedding, attention mechanisms and strategic gradient methods.
It improves the diversity and quality of the generated molecules, especially in the generation of specific functional molecules, and solves the limitations of traditional methods in the generation of molecules in terms of diversity, stability, and physical properties.
Smart Images

Figure CN119943205A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and in particular relates to a new drug molecule design method and device integrating convex optimization and evolutionary learning. Background Art
[0002] In the field of small molecule generation, as the demand for drug development and material design grows, traditional methods based on chemical theory and empirical formulas have gradually shown their limitations. These traditional methods rely on high-precision quantum chemical calculations and laboratory synthesis screening. Although they can provide reliable molecular structures, they are inefficient and cannot effectively cope with the exploration of massive chemical space. Traditional molecular generation methods usually face the challenges of high complexity and high computational cost.
[0003] In recent years, with the development of machine learning and artificial intelligence technologies, data-driven molecular generation methods have gradually attracted attention. Generative models based on large-scale molecular databases, such as variational autoencoders and autoregressive models, have been used to generate small molecule structures with specific chemical properties. By learning the distribution of a large number of known molecules, this type of method can efficiently generate small molecules that conform to specific chemical rules, significantly improving the efficiency of molecular screening. However, these generative models still have limitations in the effectiveness of generated molecules, the similarity and innovation of the original data, especially when it is necessary to generate molecules with specific functions, it is difficult to take into account the effectiveness of the generated small molecules, the similarity to the original data, the innovation and drugability.
[0004] However, these generative models still have limitations in terms of the diversity and quality of generated molecules, especially when specific functional molecules need to be generated. Existing methods are difficult to take into account the drugability, molecular sample diversity, and innovation of the molecules. Despite this, some studies have begun to try to improve the efficiency and quality of molecule generation by introducing deep learning techniques (such as generative adversarial networks, GAN). This type of model relies on the learning process of known molecules and can generate molecular structures that meet the target distribution. However, existing generative models still face challenges in generating high-quality molecules, especially when faced with high-dimensional chemical space. How to balance the diversity and practicality of generated molecules is still an urgent problem to be solved. Therefore, how to further improve the performance of the generative model, especially in improving the accuracy of specific targets such as the stability and functionality of the generated molecules, remains a research focus in the field of small molecule generation. Summary of the invention
[0005] In response to the shortcomings of the existing technology, this application proposes a new drug molecule design method and device that integrates convex optimization and evolutionary learning, which combines convex optimization and the GAN (Generative Adversarial Networks) model with reinforcement learning. Compared with other methods, this method can obtain the similarity with the original sample, the effectiveness of generating small molecules, and whether there are small molecules with higher innovation and higher drugability for the original data set.
[0006] In the first aspect, the present application proposes a new drug molecule design method integrating convex optimization and evolutionary learning, comprising:
[0007] Step S1: obtaining a small molecule initial data set, wherein each element in the small molecule initial data set is a character string of a small molecule;
[0008] Step S2: using a word embedding method to map each element in the small molecule initial data set, and converting the small molecule string into a small molecule vector;
[0009] Step S3: In the generative adversarial network, the mapping result of the small molecule initial data set is input into the generator to generate a small molecule sequence, and the loss function of the generator is calculated, wherein the generator is established by adding an attention mechanism model to the long short-term memory network;
[0010] Step S4: using a discriminator to evaluate the authenticity of the small molecule sequence and calculating the loss function of the discriminator, wherein the discriminator is established by using a convolutional neural network improved based on convex optimization;
[0011] Step S5: Taking the loss function of the generator as the objective function, the policy gradient method is used to update the parameters of the generator;
[0012] Step S6: Taking the loss function of the discriminator as the objective function, the parameters of the discriminator are updated using the gradient descent method;
[0013] Step S7: when the current number of iterations does not reach the iteration number threshold, return to step S3: regenerate the small molecule sequence using the generator corresponding to the parameters of the updated generator;
[0014] Step S8: When the current number of iterations reaches the iteration number threshold, output the final parameters of the generator and the parameters of the discriminator;
[0015] Step S9: Generate the final small molecule sequence using the generator corresponding to the parameters of the final generator.
[0016] The obtaining of the small molecule initial data set includes: extracting a SMILES data set describing the structure and characteristics of the small molecule from the QM9 data set, and preprocessing the SMILES data set to obtain the small molecule initial data set.
[0017] The word embedding method is used to map each element in the small molecule initial data set, and convert the string of the small molecule into a vector of the small molecule, including:
[0018] Constructing a vocabulary, the vocabulary including: all unique characters appearing in the SMILES dataset and corresponding numeric indices;
[0019] Using the correspondence between characters and numeric indexes in the vocabulary, each element in the small molecule initial data set is mapped to obtain a one-hot vector of the small molecule initial data set;
[0020] Performing matrix multiplication on the one-hot vector and a preset embedding matrix to obtain an embedding vector;
[0021] The multiple embedding vectors are combined into a batch embedding vector, and the batch embedding vector is used as a vector of the small molecule.
[0022] The generator includes: multiple long short-term memory networks, an attention mechanism model and a fully connected layer. Each long short-term memory network corresponds to an input vector of the generator. The output of each long short-term memory network is multiplied by the attention weight. All product results are used as the input of the attention mechanism model, the output of the attention mechanism model is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the generator.
[0023] The discriminator is established by using a convolutional neural network improved based on convex optimization, and the establishment process includes:
[0024] Embedding the small molecule initial data set and the small molecule sequence to obtain a small molecule embedding matrix;
[0025] The Nystroem method is used to perform kernel approximation on the small molecule embedding matrix to obtain the characteristic matrix;
[0026] Input the feature matrix into the convolutional neural network for convolution operation;
[0027] The convolution result is input into the maximum pooling layer for dimensionality reduction to obtain the dimensionality reduction result;
[0028] Highway network is used to process the dimension reduction results;
[0029] Optimize the output of the Highway network through a convex optimization layer;
[0030] The optimized structure is mapped to the classification result through a fully connected layer to obtain the output of the discriminator.
[0031] In the convex optimization layer, the objective function of convex optimization is to minimize the distance between the small molecule sequence and the small molecule initial data set and to maximize the consistency between the diversity of the small molecule sequence and the small molecule initial data set. The constraint condition is that the characteristics of the small molecule sequence are close to the characteristics of the small molecule initial data set.
[0032] The objective function of the convex optimization is calculated as follows:
[0033]
[0034] Among them, hh i is the output of the ith Highway network, hh j is the output of the j-th Highway network, hh i ' is the feature of the output of the i-th element of the discriminator Highway network in the initial data set of small molecules, γ is the bandwidth parameter of the kernel function in the Nystroem method, div is the function for generating sample diversity, coh is the function for generating coherence between samples, sim(hh i ,hh j ) is the normalized similarity measure Tanimoto coefficient, whose value is between 0 and 1, ||hh i -hh j || 2 is the Euclidean distance between samples, ε is a set value, which can be 0.1, δ is the regularization constraint parameter, and B is the sample batch size.
[0035] In the second aspect, the present application proposes a new drug molecule design device integrating convex optimization and evolutionary learning, comprising:
[0036] An initial sample acquisition module is used to acquire an initial data set of small molecules, each element of which is a string of small molecules;
[0037] A sample conversion module, used to map each element in the small molecule initial data set using a word embedding method, and convert the small molecule string into a small molecule vector;
[0038] A sample generation module is used to input the mapping result of the small molecule initial data set into the generator to generate a small molecule sequence in a generative adversarial network, and calculate the loss function of the generator, wherein the generator is established by adding an attention mechanism model to a long short-term memory network;
[0039] A sample evaluation module, used to evaluate the authenticity of the small molecule sequence using a discriminator and calculate the loss function of the discriminator, wherein the discriminator is established using a convolutional neural network improved based on convex optimization;
[0040] The first parameter updating module is used to update the parameters of the generator using the loss function of the generator as the objective function and the policy gradient method;
[0041] The second parameter updating module is used to update the parameters of the discriminator by using the loss function of the discriminator as the objective function and adopting the gradient descent method;
[0042] The loop iteration module is used to return to the sample generation module when the current number of iterations does not reach the iteration number threshold: the generator corresponding to the updated generator parameters is used to regenerate the small molecule sequence; when the current number of iterations reaches the iteration number threshold, the parameters of the final generator and the parameters of the discriminator are output;
[0043] The result output module is used to generate the final small molecule sequence using the generator corresponding to the parameters of the final generator.
[0044] In a third aspect, the present application proposes an electronic device comprising: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the new drug molecule design method that integrates convex optimization and evolutionary learning.
[0045] In a fourth aspect, the present application proposes a computer-readable storage medium storing executable instructions, which, when executed, enable a processor to execute the new drug molecule design method that integrates convex optimization and evolutionary learning.
[0046] Beneficial effects:
[0047] This application proposes a new drug molecule design method and device that integrates convex optimization and evolutionary learning. Compared with traditional methods, this application can generate small molecules with high similarity to the original samples, stronger effectiveness and innovation, and improve the diversity and quality of generated molecules, especially in generating molecules with specific functions. This method not only improves the efficiency and quality of small molecule generation, but also solves the limitations of traditional methods in generating molecular diversity, stability, and physical properties, providing strong technical support for fields such as drug research and development and material design. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Flow chart of a new drug molecule design method integrating convex optimization and evolutionary learning in an embodiment of the present application
[0049] Figure 2 A schematic flow chart of a new drug molecule design method integrating convex optimization and evolutionary learning according to an embodiment of the present application;
[0050] Figure 3 Schematic diagram of the generator structure of the embodiment of the present application
[0051] Figure 4 Schematic diagram of the discriminator structure of the embodiment of the present application
[0052] Figure 5 Comparison chart of generated small molecules and sample small molecules in the embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the relevant drawings. The preferred embodiments of the present application are given in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thoroughly and comprehensively understood.
[0054] Small molecule generation is an indispensable part of pharmaceutical research and new drug discovery. In order to provide higher quality and more diverse small molecules, this application provides a new drug molecule design method and device that integrates convex optimization and evolutionary learning, which can generate innovative and high-quality small molecule datasets based on the original real small molecule datasets.
[0055] Embodiment 1:
[0056] This application proposes a new drug molecule design method that integrates convex optimization and evolutionary learning, such as Figure 1 As shown, including:
[0057] Step S1: obtaining a small molecule initial data set, wherein each element in the small molecule initial data set is a character string of a small molecule;
[0058] The method of obtaining the initial data set of small molecules includes: extracting a SMILES data set describing the structure and characteristics of small molecules from the QM9 data set, preprocessing the SMILES data set, including: removing duplicate data in the data set and using Rdkit to remove invalid molecules, to obtain the initial data set of small molecules.
[0059] In this example, SMILES describing the structure and characteristics of small molecules were extracted from the QM9 data set, and a data set of 130,000 SMILES was obtained after processing.
[0060] Step S2: using a word embedding method to map each element in the small molecule initial data set, and converting the small molecule string into a small molecule vector;
[0061] In this embodiment, the 130,000 data obtained in step S1 are subjected to a commonly used word embedding method in text data, thereby capturing and retaining the key features of the small molecule sequence, which is convenient for subsequent transmission to the neural network training, specifically including:
[0062] The word embedding method is used to map each element in the small molecule initial data set, and convert the string of the small molecule into a vector of the small molecule, including:
[0063] Step S2.1: construct a vocabulary, the vocabulary including: all unique characters appearing in the SMILES dataset and their corresponding numeric indices;
[0064] In this embodiment, in order to convert the SMILES dataset into an embedding vector, a vocabulary M needs to be constructed. The vocabulary M includes: all unique characters appearing in the SMILES dataset and their corresponding numeric indexes;
[0065] Step S2.2: using the correspondence between characters and numeric indexes in the vocabulary, mapping each element in the small molecule initial data set to obtain a one-hot vector of the small molecule initial data set;
[0066] In this embodiment, the character-to-index mapping uses a mapping function f to find the index i in the vocabulary for each character c in the SMILES string, where i=1, 2, ..., N:
[0067] i=f(c)
[0068] Where f(c) is a lookup function that returns the position of character c in the vocabulary, and N is the total number of indices.
[0069] In this embodiment, the vocabulary corresponding to the characters is as follows:
[0070] {"=":"0","-":"1","]":"2","2":"3","H":"4","F":"5","+":"6",")":"7","#":"8","(":"9","N":"10",
[0071] "O":"11","4":"12","C":"13","[":"14","n":"15","1":"16","5":"17","c":"18","3":"19","o":"20"}
[0072] Step S2.3: performing matrix multiplication on the one-hot vector and a preset embedding matrix to obtain an embedding vector;
[0073] In this embodiment, for each SMILES character, according to the character index i∈[1,N], it is converted into a one-hot vector with the same size as the vocabulary, where the index position corresponding to the character is 1 and the rest of the positions are 0. Then, these one-hot vectors are combined with the predefined embedding matrix E∈RN×D Matrix multiplication is performed to efficiently convert the discrete one-hot representation into an embedding vector. Each row of the embedding matrix represents the embedding vector for a character in the vocabulary. When the one-hot vector is multiplied by the embedding matrix, only the positions that are 1 in the one-hot vector match the corresponding rows of the embedding matrix, and the result is that the embedding vector e of the character is selected and returned. t ;
[0074] e t =E(i)
[0075] Where N is the size of the vocabulary, D is the dimension of the embedding, and the embedding vector e t represents the vector of the i-th row in the embedding matrix E, that is, the embedding vector of the character with index i.
[0076] Step S2.4: Combine multiple embedding vectors into a batch embedding vector, and use the batch embedding vector as the vector of the small molecule.
[0077] In this embodiment, multiple embedding vectors e t Combined into a batch embedding matrix, the input sequences of different lengths are padded to ensure that all sequences have the same length L sequence representation x g ∈R L×D , the embedded matrix is X∈R B×L×D , where B is the batch size;
[0078] x g =[e1,e2,...,e l ]
[0079] X=[x1,x2,...,x B ]
[0080] Step S3: In the generative adversarial network, the mapping result of the small molecule initial data set is input into the generator to generate a small molecule sequence, and the loss function of the generator is calculated, wherein the generator is established by adding an attention mechanism model to the long short-term memory network;
[0081] In this embodiment, the generator, such as Figure 3 As shown, it includes: multiple long short-term memory networks, an attention mechanism model and a fully connected layer. Each long short-term memory network corresponds to an input vector of the generator. The output of each long short-term memory network is multiplied by the attention weight. All product results are used as the input of the attention mechanism model, the output of the attention mechanism model is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the generator.
[0082] Among them, LSTM (Long Short-Term Memory) is an improvement of recurrent neural network, which effectively solves the gradient explosion and gradient vanishing problems of RNN (Recurrent Neural Network). LSTM processes the embedding vector sequence of step S2 and captures the long-term dependencies of the sequence through its internal gating mechanism. At each time step, the LSTM unit considers the current input and the previous hidden state to calculate the new hidden state and cell state; the LSTM unit consists of a forget gate, an input gate, a cell state, and an output gate. The formulas for each gate and cell unit are as follows:
[0083] (Lstm formula at time t)
[0084] Forget gate: f t =σ(W f ·[h t-1 ,e t ]+b f )
[0085] Input gate: i t =σ(W i ·[h t-1 ,e t ]+b i )
[0086]
[0087] Unit status update:
[0088] Output gate: o t =σ(W o ·[h t-1 ,e t ]+b o )
[0089] h t =o t tanh(C t +A t )
[0090] Among them, f t is the output of the forget gate, i t is the output of the input gate, is the candidate cell state, C t is the current cell state, o t is the output of the output gate, h t is the current hidden state, W f , W i , W C , W oare the weight matrices of the forget gate, input gate, candidate cell state, and output gate, respectively, and b f , b i , b C , b o are the bias vectors of the forget gate, input gate, candidate cell state, and output gate, respectively. σ is the sigmoid function, which maps the input to between 0 and 1. t is the embedding vector representing the current input character, h t-1 is the hidden state of the previous time step, A t is the weighted information calculated by the self-attention mechanism based on the current input sequence.
[0091] Adding an attention mechanism to LSTM allows the model to focus on the most important parts of the sequence when generating new samples. It weights different parts of the input sequence by calculating attention weights, thereby obtaining a weighted context vector. The calculation of attention weights usually involves the conversion of queries, keys, and values, and is normalized by the softmax function; the attention calculation formula is as follows:
[0092]
[0093] Q t =W q x g
[0094] K t =W K x g
[0095] V t =W V x g
[0096] Among them, x g =[e1,e2,...,e l ] represents the input sequence, W q , W K and W V represents the learned weight matrix, Q t , K t and V t Denote query, key and value matrices d respectively k Represents the input feature dimension.
[0097] The weighted context vector is passed to a fully connected layer to generate the final output. The output layer produces a probability distribution for the next word in the sequence. New samples are generated by sampling words from the probability distribution produced by the output layer, and the words with the highest probability are selected;
[0098] Step S4: using a discriminator to evaluate the authenticity of the small molecule sequence and calculating the loss function of the discriminator, wherein the discriminator is established by using a convolutional neural network improved based on convex optimization;
[0099] In this embodiment, the discriminator is established by using a convolutional neural network based on convex optimization, and the discriminator is used to evaluate the authenticity of the generated small molecule sequence to ensure the high quality and authenticity of the generator output, such as Figure 4 As shown, the establishment process includes:
[0100] Step S4.1: embedding the small molecule initial data set and the small molecule sequence to obtain a small molecule embedding matrix;
[0101] In this embodiment, the generated SMILES data (i.e., small molecule sequence) and the real SMILES data (i.e., small molecule initial data set) are embedded, wherein the dimension of the embedding layer is set to 32;
[0102] In this embodiment, the generated SMILES data "CNc1c(oc[nH]c1N)F" and the real SMILES data "CC12OC3C(=O)C1C23C" have index numbers [13,13,16,3,11,13,19,13,16,13,3,19,13] and [13,10,18,16,18,15,4,18,16,10,5], and are transformed into the following matrix:
[0103] [[-0.2342 0.3453 0.4564 0.5675 0.6786 0.7897 0.8908 0.9010 0.0120 0.1231 0.2342 0.3453 0.4565 0.5676 0.6787 0.7898 0.8909 0.9019 0.0120 0.1231 0.2342 0.3453 0.4565 0.5676 0.6787 0.7898 0.8909 0.9019 0.0120 0.1231]
[0108] [-0.2342 0.3453 0.4565 0.5676 0.6787 0.7898 0.8909 0.9019 0.0120 0.1231 0.2342 0.3453 0.4565 0.5676 0.6787 0.7898 0.8909 0.9019 0.0120 0.1231 0.2342 0.3453 0.4565 0.5676 0.6787 0.7898 0.8909 0.9019 0.0120 0.1231] ... ]
[0115] Step S4.2: Use the Nystroem method to perform kernel approximation on the small molecule embedding matrix to obtain the feature matrix;
[0116] In this embodiment, the embedding matrix is X∈R L×D First, Nystroem kernel approximation is applied to each row of the embedding matrix (i.e., the embedding vector of each character), and the original input is mapped through kernel approximation, converting the original nonlinear space into a high-dimensional linear space, so as to facilitate the subsequent conversion of the non-convex optimization properties of the convolutional neural network into a convex optimization problem: The role of using the Nystroem method: RKF (Rank-k Factorization) can be used to map low-dimensional data to linear high-dimensional, and then map feature data to low-dimensional.
[0117] First, using the Nystroem method, a randomly selected reference point x j To approximate the kernel matrix of the high-dimensional linear space. Here, the Nystroem kernel approximation uses Gaussian RKF (Rank-k Factorization) to calculate the kernel matrix and reduce the dimension. The purpose is to efficiently approximate the large kernel matrix. The kernel function calculation formula is as follows:
[0118]
[0119] Among them, γ is the parameter of Gaussian kernel, x i is the input vector of the Nystroem method, x j The reference vector is the input to the Nystroem method. n components is the number of reference points, kf(x i ,x j ) is x i With x j The kernel function of .
[0120] In order to map data to a high-dimensional space, the kernel function assumes the existence of a mapping function φ, as shown in the following formula so that the kernel function can be represented by the inner product, where φ(x i ) is the input data x iThe Nystroem method avoids calculating all kernel inner products by low-rank approximation of the kernel matrix, and instead speeds up the process by randomly selecting reference points and approximate calculations.
[0121] kf(x i ,x j )=<φ(x i ),φ(x j )>
[0122] Perform singular value decomposition on the kernel matrix:
[0123] kf(x i ,x j )=UΣV T Among them, U represents the left singular matrix, Σ represents the diagonal matrix, and V represents the right singular matrix.
[0124] Then, low-dimensional approximate feature mapping is performed, and the calculated low-dimensional features are:
[0125]
[0126] Among them, X' is the low-dimensional feature, that is, the feature matrix obtained by the Nystroem method.
[0127] Step S4.3: Input the feature matrix into the convolutional neural network for convolution operation;
[0128] In this embodiment, the feature matrix X' of 4.2 is sent to the convolutional neural network for convolution operation:
[0129] C=Conv(X',K,S)
[0130] in, is the convolution kernel, K H and K W are the height and width of the convolution kernel, C I is the number of input channels C O is the number of output channels and S is the step size.
[0131] After the convolution operation, the nonlinear activation function Relu is applied to increase the expressiveness of the model.
[0132] h = ReLU(Conv(W·X'+b))
[0133] Among them, Relu is a nonlinear activation function, Conv is a convolution operation, W is the weight of the convolution, b is the bias of the convolution, and h is the output after activation.
[0134] Method of applying residual neural network after convolution:
[0135] h'=h+X'
[0136] Among them, h' is the output of the residual neural network:
[0137] Step S4.4: input the convolution result into the maximum pooling layer for dimensionality reduction to obtain a dimensionality reduction result;
[0138] In this embodiment, after the convolution operation, the output is reduced in dimension through maximum pooling to retain important features;
[0139] hp=MAX(h')
[0140] Among them, MAX represents the maximum pooling operation and hp is the dimensionality reduction result.
[0141] Step S4.4: Use Highway network to process the dimension reduction result;
[0142] In this embodiment, Highway Networks is used to process the pooled features. Highway Networks includes linear transformations with gating mechanisms.
[0143] g=f(W g ·hp+b g )
[0144] t=σ(W t ·hp+b t )
[0145] hh=t·g+(1-t)·hp
[0146] Among them, g is the output of the activation function, f is the transformation function, t is the gate unit, σ is the sigmoid function, and W g , W t are the weight matrices of the transformation function and the gating function, respectively, and b g , b t They represent the bias of the transformation function and the gating function respectively.
[0147] Step S4.5: Optimize the output of the Highway network through a convex optimization layer;
[0148] In this embodiment, the features passing through the Highway network are optimized through a convex optimization layer. The objective function of the convex optimization is to minimize the distance between the generated samples and the real samples and to maximize the diversity of the generated samples and the overall consistency of the samples. The constraint condition is that the generated samples are as close as possible to the real sample features, and an L2 regularization term is added to prevent overfitting:
[0149]
[0150] Among them, hh represents the characteristics of the highway layer output of the generated SMILES, hh' represents the characteristics of the real sample, and γ is the kernel function bandwidth parameter. div and coh are functions for calculating the diversity of generated samples and the coherence between samples. sim(hh i ,hh j ) is the normalized similarity metric Tanimoto calculation, and its value is between 0 and 1. ||hh i -hh j || 2 It represents the Euclidean distance between samples. The Euclidean distance obtained through samples represents the coherence between samples. ε is a very small value, which controls the allowable range of similarity between generated samples and target samples. δ is a regularization constraint parameter used to prevent overfitting. Here we set ε and δ to 0.1.
[0151] Step S4.6: Map the optimized structure to the classification result through the fully connected layer to obtain the output of the discriminator.
[0152] In this embodiment, the optimized features are mapped to the classification results through a fully connected layer. The softmax function is used to generate a probability distribution to obtain the predicted probability of the authenticity of the generated SMILES;
[0153] z=softmax(W out ·hh+b out )
[0154] Among them, z is the probability distribution of the output, that is, the output of the discriminator, W out , b out Represent the output layer weight matrix and bias respectively.
[0155] Step S5: Taking the loss function of the generator as the objective function, the policy gradient method is used to update the parameters of the generator;
[0156] Step S6: Taking the loss function of the discriminator as the objective function, the parameters of the discriminator are updated using the gradient descent method;
[0157] In this embodiment, the generator is pre-trained on real data using maximum likelihood estimation to learn basic sequence patterns; the pre-trained generator is used to generate negative samples, and the discriminator is trained in combination with positive samples of real data, and the parameters of the discriminator are optimized by minimizing the cross entropy loss; in the adversarial training stage, the discriminator evaluates the sequence generated by the generator, and the discriminator provides a reward signal reflecting the authenticity of the sequence for the generated sequence, and the generator parameters are updated using the policy gradient method to maximize the expected reward, and the Monte Carlo search is used to estimate the reward of the future state-action pair; specifically, the generator and discriminator of the generative adversarial network are pre-trained for 50 epochs, and the pre-training is to train the generator and discriminator separately based on real data. In this way, the generator can gradually learn to generate samples close to real data, while the discriminator can improve the ability to distinguish between real data and generated data.
[0158] The parameter update of the generator adopts the method based on PolicyGradient. The goal of the generator is to maximize the probability that the generated samples are considered as real samples by the discriminator, that is, the similarity between the generated samples and the real samples. The Adam (Adaptive Moment Estimation) algorithm in the gradient descent method is used to update its own parameters through the feedback of the discriminator in step 4;
[0159] The pre-training loss function of the generator is defined as:
[0160]
[0161] G θ (y t |y 1:t-1 ) represents the generator G θ Under the given historical conditions y 1:t-1 Generate the tth target y t probability.
[0162] The adversarial training loss function of the generator is defined as:
[0163]
[0164] Among them, L G (θ) is the objective function of the generator, θ represents the generator parameter, T is the length of the generated sequence, y 1:t-1 ~G represents the partial sequence generated under the generator strategy. R(y 1:T ) is the reward value for the complete sequence, which is calculated by the discriminator based on the quality of the generated samples. At time step t, according to the generator G θ For the previous sequence y 1:t-1 expected value.
[0165] For the sequence generated by the generator, the discriminator determines whether it is real data or generated data based on the input sample. Through the classification error, the discriminator will update its own parameters, so that it gradually improves in distinguishing true and false samples;
[0166] The loss function of the discriminator is the sum of the cross loss function, L2 regularization, and nuclear norm regularization.
[0167]
[0168] Among them, y i is the label of the data (1 represents real data, 0 represents generated data), D φ (y i ) is the predicted probability of the discriminator for the i-th sample, λ1 is the regularization coefficient, is the L2 norm regularization, λ2 is the nuclear norm regularization coefficient, ||W HS || * is the nuclear norm regularization.
[0169] Update the generator model parameters by calculating the policy gradient and Adam optimizer:
[0170]
[0171] Where η is the learning rate, Find the gradient G for the generator loss function θ (y t ∣y 1:t-1 ) is the conditional probability of the generator. Step S7: when the current number of iterations does not reach the iteration number threshold, return to step S3: use the generator corresponding to the updated generator parameters to regenerate the small molecule sequence;
[0172] In this embodiment, Figure 2 As shown, the number of iterations includes the total number of iterations and the batch number of iterations. When any of the total number of iterations and the batch number of iterations does not meet the corresponding iteration threshold, it is necessary to return to step S3. In this embodiment, adversarial training is performed for 100 eopch (total number of iterations). The generator and the discriminator are trained alternately. The goal of the generator is to generate more and more realistic samples, while the discriminator will strive to improve the ability to distinguish between true and false samples. The small molecule sample data generated by the generator is sent to the discriminator and the real sample data to identify the authenticity of the generated sample data based on the features. The reward mechanism of the discriminator is used to update the generator parameters, and the generator's generated samples and labels are used to update the discriminator parameters. Repeat the above steps, using the idea of the game between the generator and the discriminator, so that the small molecules generated by the generator are more realistic, and the discriminator has a stronger ability to distinguish true and false samples, so as to generate small molecule data that can be realistic with the original data set;
[0173] Step S8: When the current number of iterations reaches the iteration number threshold, output the final parameters of the generator and the parameters of the discriminator;
[0174] Step S9: Generate the final small molecule sequence using the generator corresponding to the parameters of the final generator.
[0175] In the present embodiment, RDKit is used to calculate the properties of the small molecule until the properties of the generated small molecule reach the target. RDKit is an open source chemical informatics software package that provides a wealth of data structures and algorithms to process chemical data, including the construction, analysis and operation of molecules, and is widely used in drug discovery, chemical education and research fields. Because BLEU (Bilingual Evaluation Understudy) score is a common indicator for evaluating the similarity between machine translation text and human translation text, it calculates accuracy by comparing the n-gram matching degree between the generated text and the reference text. In the sequence generation task of small molecule generation, the BLEU value can be used to measure the similarity between the generated SMILES string and the SMILES string of the real molecule. Finally, the measurement indicators Bleu and Bleu2 values in the field of text generation and the measurement indicators in the field of small molecule generation are used to evaluate the generation results of the small molecule generation method based on the above-mentioned convex optimization generation adversarial network.
[0176] In this embodiment, the small molecule generation method based on the convex optimization generative adversarial network of the present invention is applied to the QM9 organic small molecule data set. The QM9 data set is a database containing the quantum chemical properties of more than 130,000 organic small molecules, which are composed of carbon (C), hydrogen (H), oxygen (O), nitrogen (N) and fluorine (F). Each molecule contains complete spatial information for atoms in a single low-energy conformation for calculating chemical properties. The data set provides information such as the geometric structure, energy, electronic and thermodynamic properties of the molecule, including the geometric conformation with the minimum molecular energy, the corresponding harmonic frequency, dipole moment, polarizability, as well as energy, thermal enthalpy and free energy. Due to its scale and level of detail, the QM9 data set has become an important resource in the fields of chemical informatics, materials science and machine learning, especially in the fields of molecular property prediction and new drug development.
[0177] In the small molecule generation task, we use the evaluation indicators Bleu and Bleu2 in terms of text evaluation indicators and small molecule evaluation indicators validity, diversity, novelty and drug-likeness. Bleu is used to evaluate the similarity between the chemical structure of the generated molecule and the target molecule structure; Bleu2 pays special attention to the overlap between two consecutive words in the text, by counting the maximum number of times each 2-gram in the generated sequence appears in the reference sequence, normalizing it, and finally taking the geometric mean. Here Bleu2 can be used to measure the similarity between the generated molecular structure and the known valid molecular structure, thereby evaluating the accuracy and rationality of the generated molecule; molecular validity is a validity indicator to ensure that the generated small molecules are chemically feasible, that is, they conform to known chemical rules and reaction mechanisms; the molecular diversity index measures the diversity of molecules in the generated small molecule set. In small molecule generation, high diversity means the ability to explore a wider range of chemical space, thereby increasing the chances of discovering new drugs or new materials; the molecular novelty index measures the similarity between the generated small molecules and known molecules, and is calculated by the Simpson diversity index, which takes into account the richness and uniformity of species. High novelty means that the generated molecules are very different from known molecules in structure and properties. The drug-likeness index measures whether the generated small molecules have the physicochemical properties of known drugs. This index is used to predict whether a molecule is likely to become an effective drug candidate, including pharmacological activity, suitable physicochemical properties and ADMET properties. The QED score is obtained by calculating the weighted sum of these properties, where the weight of each attribute is determined based on maximizing Shannon entropy.
[0178] These indicators together constitute a multi-dimensional framework for evaluating the effectiveness of small molecule generation, and the relevant results are shown in Table 1. From the experimental results, it can be seen that the SMILES generated with a similarity of up to 0.79 with the real sample data has a great similarity with the characteristics of the small molecules in the original QM9 dataset; the effectiveness of the generated small molecules is 0.35, indicating that some of the generated samples are data that can form small molecules; the drugability is as high as 0.17, indicating that the new molecule has the potential to become an effective drug candidate. Figure 5 The comparison of the structure diagrams of some generated small molecules and QM9 small molecules drawn by Pymol software is shown in the figure. The comparison shows that the generated small molecules have great consistency in structure. This method has high value in the field of small molecule discovery in drug development.
[0179] Table 1 Generate small molecule evaluation results table
[0180]
[0181] This embodiment proposes a new drug molecule design method that integrates convex optimization and evolutionary learning. First, by extracting the QM9 organic small molecule data set, an initial data set containing n samples is obtained; then, the data set is preprocessed using word embedding technology to convert the molecular sequence into a vector representation, effectively capturing the key features of the molecule; then, an LSTM model based on the attention mechanism is constructed as a generator to generate high-quality small molecule sequences; at the same time, a convolutional neural network based on convex optimization is constructed as a discriminator to evaluate the authenticity of the generated small molecule sequence. In addition, the present invention uses maximum likelihood estimation to pre-train the generator on real data to learn basic sequence patterns, and through the adversarial training stage, the policy gradient method is used to update the generator parameters to maximize the expected reward, and Monte Carlo search is used to estimate the reward of future state-action pairs. Finally, the properties of the small molecule are calculated by RDKit until the target properties of the generated small molecule and the original data are more than 79% similar and the innovation is more than 85%.
[0182] Embodiment 2:
[0183] The present application proposes a new drug molecule design device integrating convex optimization and evolutionary learning, including: an initial sample acquisition module, a sample conversion module, a sample generation module, a sample evaluation module, a first parameter update module, a second parameter update module, a loop iteration module, and a result output module;
[0184] The initial sample acquisition module is connected to the sample conversion module, the sample conversion module is connected to the sample generation module, the sample generation module is connected to the sample evaluation module, the sample evaluation module is connected to the first parameter update module, the first parameter update module is connected to the second parameter update module, the second parameter update module is connected to the loop iteration module, and the loop iteration module is respectively connected to the sample generation module and the result output module;
[0185] An initial sample acquisition module is used to acquire an initial data set of small molecules, each element of which is a string of small molecules;
[0186] A sample conversion module, used to map each element in the small molecule initial data set using a word embedding method, and convert the small molecule string into a small molecule vector;
[0187] A sample generation module is used to input the mapping result of the small molecule initial data set into the generator to generate a small molecule sequence in a generative adversarial network, and calculate the loss function of the generator, wherein the generator is established by adding an attention mechanism model to a long short-term memory network;
[0188] A sample evaluation module, used to evaluate the authenticity of the small molecule sequence using a discriminator and calculate the loss function of the discriminator, wherein the discriminator is established using a convolutional neural network improved based on convex optimization;
[0189] The first parameter updating module is used to update the parameters of the generator using the loss function of the generator as the objective function and the policy gradient method;
[0190] The second parameter updating module is used to update the parameters of the discriminator by using the loss function of the discriminator as the objective function and adopting the gradient descent method;
[0191] The loop iteration module is used to return to the sample generation module when the current number of iterations does not reach the iteration number threshold: the generator corresponding to the updated generator parameters is used to regenerate the small molecule sequence; when the current number of iterations reaches the iteration number threshold, the parameters of the final generator and the parameters of the discriminator are output;
[0192] The result output module is used to generate the final small molecule sequence using the generator corresponding to the parameters of the final generator.
[0193] Embodiment 3:
[0194] This embodiment proposes an electronic device, including: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the new drug molecule design method that integrates convex optimization and evolutionary learning.
[0195] The electronic device may be a mobile phone, a computer or a tablet computer, etc., including a memory and a processor, wherein a computer program is stored on the memory, and when the computer program is executed by the processor, a new drug molecule design method integrating convex optimization and evolutionary learning as described in the embodiment is implemented. It is understood that the electronic device may also include an input / output (I / O) interface and a communication component.
[0196] The processor is used to execute all or part of the steps in the new drug molecule design method integrating convex optimization and evolutionary learning as described in the above embodiment. The memory is used to store various types of data, which may include instructions of any application or method in the electronic device, as well as data related to the application.
[0197] The processor can be an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor or other electronic components to execute the new drug molecule design method integrating convex optimization and evolutionary learning described in the above embodiments.
[0198] Embodiment 4:
[0199] This embodiment provides a computer-readable storage medium storing executable instructions. When the instructions are executed and implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0200] The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the new drug molecule design method that integrates convex optimization and evolutionary learning as described in the various embodiments of the present application.
[0201] The aforementioned storage media include: flash memory, hard disk, multimedia card, card-type memory (for example, SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR abbreviation, memory data register) memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, CD, server, APP (Application, abbreviation of application software) application store and other media that can store program verification codes, on which computer programs are stored. When the computer program is executed by the processor, the various steps of the new drug molecule design method that integrates convex optimization and evolutionary learning described above can be implemented.
[0202] The protection scope of the present application is not limited to the above-mentioned embodiments. Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the scope and spirit of the present disclosure. If these changes and modifications fall within the scope of the claims of the present disclosure and their equivalents, the intention of the present disclosure also includes these changes and modifications.
Claims
1. A new drug molecule design method integrating convex optimization and evolutionary learning, characterized in that: include: Step S1: obtaining a small molecule initial data set, wherein each element in the small molecule initial data set is a character string of a small molecule; Step S2: using a word embedding method to map each element in the small molecule initial data set, and converting the small molecule string into a small molecule vector; Step S3: In the generative adversarial network, the mapping result of the small molecule initial data set is input into the generator to generate a small molecule sequence, and the loss function of the generator is calculated, wherein the generator is established by adding an attention mechanism model to the long short-term memory network; Step S4: using a discriminator to evaluate the authenticity of the small molecule sequence and calculating the loss function of the discriminator, wherein the discriminator is established by using a convolutional neural network improved based on convex optimization; Step S5: Taking the loss function of the generator as the objective function, the policy gradient method is used to update the parameters of the generator; Step S6: Taking the loss function of the discriminator as the objective function, the parameters of the discriminator are updated using the gradient descent method; Step S7: when the current number of iterations does not reach the iteration number threshold, return to step S3: regenerate the small molecule sequence using the generator corresponding to the parameters of the updated generator; Step S8: When the current number of iterations reaches the iteration number threshold, output the final parameters of the generator and the parameters of the discriminator; Step S9: Generate the final small molecule sequence using the generator corresponding to the parameters of the final generator.
2. The new drug molecule design method integrating convex optimization and evolutionary learning according to claim 1, characterized in that: The obtaining of the small molecule initial data set includes: extracting a SMILES data set describing the structure and characteristics of the small molecule from the QM9 data set, and preprocessing the SMILES data set to obtain the small molecule initial data set.
3. The new drug molecule design method integrating convex optimization and evolutionary learning according to claim 1, characterized in that: The word embedding method is used to map each element in the small molecule initial data set, and convert the string of the small molecule into a vector of the small molecule, including: Constructing a vocabulary, the vocabulary including: all unique characters appearing in the SMILES dataset and corresponding numeric indices; Using the correspondence between characters and numeric indexes in the vocabulary, each element in the small molecule initial data set is mapped to obtain a one-hot vector of the small molecule initial data set; Performing matrix multiplication on the one-hot vector and a preset embedding matrix to obtain an embedding vector; The multiple embedding vectors are combined into a batch embedding vector, and the batch embedding vector is used as a vector of the small molecule.
4. The new drug molecule design method integrating convex optimization and evolutionary learning according to claim 1, characterized in that: The generator includes: multiple long short-term memory networks, an attention mechanism model and a fully connected layer. Each long short-term memory network corresponds to an input vector of the generator. The output of each long short-term memory network is multiplied by the attention weight. All product results are used as the input of the attention mechanism model, the output of the attention mechanism model is used as the input of the fully connected layer, and the output of the fully connected layer is used as the output of the generator.
5. The new drug molecule design method integrating convex optimization and evolutionary learning according to claim 1, characterized in that: The discriminator is established by using a convolutional neural network improved based on convex optimization, and the establishment process includes: Embedding the small molecule initial data set and the small molecule sequence to obtain a small molecule embedding matrix; The Nystroem method is used to perform kernel approximation on the small molecule embedding matrix to obtain the characteristic matrix; Input the feature matrix into the convolutional neural network for convolution operation; The convolution result is input into the maximum pooling layer for dimensionality reduction to obtain the dimensionality reduction result; Highway network is used to process the dimension reduction results; Optimize the output of the Highway network through a convex optimization layer; The optimized structure is mapped to the classification result through a fully connected layer to obtain the output of the discriminator.
6. The new drug molecule design method integrating convex optimization and evolutionary learning according to claim 1, characterized in that: In the convex optimization layer, the objective function of convex optimization is to minimize the distance between the small molecule sequence and the small molecule initial data set and to maximize the consistency between the diversity of the small molecule sequence and the small molecule initial data set. The constraint condition is that the characteristics of the small molecule sequence are close to the characteristics of the small molecule initial data set.
7. The new drug molecule design method integrating convex optimization and evolutionary learning according to claim 6, characterized in that: The objective function of the convex optimization is calculated as follows: Among them, hh i is the output of the ith Highway network, hh j is the output of the j-th Highway network, hh i ' is the feature of the i-th element in the initial data set of small molecules, γ is the bandwidth parameter of the kernel function in the Nystroem method, div is the function for generating sample diversity, coh is the function for generating inter-sample coherence, sim(hh i ,hh j ) is the normalized similarity measure Tanimoto coefficient, whose value is between 0 and 1, ||hh i -hh j || 2 is the Euclidean distance between samples, ε is a set value, δ is the regularization constraint parameter, and B is the sample batch size.
8. A new drug molecule design device integrating convex optimization and evolutionary learning, characterized in that: include: An initial sample acquisition module is used to acquire an initial data set of small molecules, each element of which is a string of small molecules; A sample conversion module, used to map each element in the small molecule initial data set using a word embedding method, and convert the small molecule string into a small molecule vector; A sample generation module is used to input the mapping result of the small molecule initial data set into the generator to generate a small molecule sequence in a generative adversarial network, and calculate the loss function of the generator, wherein the generator is established by adding an attention mechanism model to a long short-term memory network; A sample evaluation module, used to evaluate the authenticity of the small molecule sequence using a discriminator and calculate the loss function of the discriminator, wherein the discriminator is established using a convolutional neural network improved based on convex optimization; The first parameter updating module is used to update the parameters of the generator using the loss function of the generator as the objective function and the policy gradient method; The second parameter updating module is used to update the parameters of the discriminator by using the loss function of the discriminator as the objective function and adopting the gradient descent method; The loop iteration module is used to return to the sample generation module when the current iteration number does not reach the iteration number threshold: the generator corresponding to the parameters of the updated generator is used to regenerate the small molecule sequence; When the current number of iterations reaches the iteration threshold, the final parameters of the generator and the parameters of the discriminator are output; The result output module is used to generate the final small molecule sequence using the generator corresponding to the parameters of the final generator.
9. An electronic device, characterized in that: include: One or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the new drug molecule design method integrating convex optimization and evolutionary learning as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The device stores executable instructions, which, when executed, enable a processor to execute the new drug molecule design method integrating convex optimization and evolutionary learning as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image description confrontation generation method based on reinforcement learning
CN114022687A
Medical Chinese named entity recognition method based on generative model
CN115630649A
Cyclic adversarial neural network molecule generation method and system based on embedded bidirectional long-short term memory
CN117079745A
Molecular generation method and system based on comparative learning
CN118398112A
Rotor wing submarine-launched unmanned aerial vehicle state monitoring method
CN118568675A
Cited By
Metal material design method based on deep learning
CN120356570A
A metal material design method based on deep learning
CN120356570B
Power system extreme weather scene small sample generation and identification method based on multi-generator-evaluator cooperation mechanism and storage medium
CN122173937A