Attribute classification-based molecular de novo design method

Through the molecular de novo design method based on attribute classification, the encoder-decoder model and auxiliary classifier generate an adversarial network model is solved, and the GAN has limited generators, unstable training and lack of control in molecular de novo design, achieving high-precision, reasonable structure and biochemical molecular generation.

CN120015168APending Publication Date: 2025-05-16NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062717.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the molecular de novo design, the existing generative adversarial network (GAN) has problems in which generators only learn finite molecular structure, easily oscillate and do not converge during training, lack of explicit control to generate results and potential spatial interpretation capabilities.

Method used

Using a molecular de novo design method based on attribute classification, by dividing the molecular data into different categories according to attributes, an encoder-decoder model and auxiliary classifier are constructed to generate an adversarial network model, and trained to generate molecules that meet attribute categories.

Benefits of technology

The high-precision generation of molecules that meet the attribute category is achieved, ensuring that the generated molecules are not only rational in structure but also have the required biochemical properties, and avoiding the generation of irregular or invalid molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015168A_ABST
    Figure CN120015168A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical chemistry, in particular to an attribute classification-based molecular de novo design method, which comprises the following steps of: dividing molecular data into different types of molecular data according to molecular attributes; obtaining a motif and a motif vocabulary of the molecular data; constructing a molecular de novo design model, and training to obtain a trained target molecular de novo design model; and inputting the molecular graph structures and the motifs of the molecular data of different categories into the target molecular de novo design model, and outputting the reconstructed molecular graph structure of the molecular data of the category. According to the method, the characteristics of the properties in the molecules can be learned in the molecule generation process, so that the molecules conforming to the attribute categories are generated with high precision, the graph structure of the molecules is effective, and the encoder reconstruction rate is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medicinal chemistry, and in particular to a molecular de novo design method based on attribute classification. Background Art

[0002] De novo molecular design is a strategy based on chemical knowledge and computational methods, which aims to design and construct novel molecular structures with specific properties and functions. With the continuous advancement of computer hardware and software, de novo molecular design can use more powerful computing power and more efficient algorithms for simulation and optimization, improving the accuracy and efficiency of prediction and design. Through machine learning and artificial intelligence technology, de novo molecular design can learn from a large amount of research data and discover potential structure-property relationships, which helps to accelerate the design process and improve the success rate of design.

[0003] Molecular de novo design breaks through the limitations of traditional trial and error methods and can explore infinite possible molecular spaces theoretically and computationally. This allows researchers to create chemical structures and functions that have never been touched before, providing new possibilities for discovering novel compounds and solving practical problems. Traditional drug discovery is usually achieved through modification and optimization of known drugs, while molecular de novo design makes it possible to directly design highly specific molecules to meet specific drug targets. This does not require the use of existing compound libraries, and can find drug candidates more quickly, improving the efficiency of drug discovery. Through molecular de novo design, tedious trial and error processes and large-scale laboratory synthesis can be avoided, saving time, resources and money. Through simulation and optimization on a computer, various design schemes can be quickly evaluated to screen out the most promising and economically efficient molecular structures.

[0004] In order to better apply it in the field of pharmaceutical molecule design, deep learning networks require effective and reasonable transformation of molecular input forms. After the development of deep learning network structures, several mainstream molecular expression forms have gradually formed as follows: Based on the linear input specification of molecules; Based on the two-dimensional image representation of molecules Based on the graph structure representation of molecules. The Simplified Molecular Input Line Entry System (SMILES) is a chemical symbol system designed for modern chemical information processing. It is a molecular text representation based on the principle of molecular graph theory. Although the molecular linear input specification can ensure that the chemical structure diagram of the molecule is uniquely described, the SMILES syntax is not robust to small changes or errors, which greatly increases the difficulty of machine learning to understand molecular features. The generation of two-dimensional images has been deeply developed in the field of computer vision, and there are relatively mature corresponding deep learning model architectures for different types of image generation. However, there is an essential difference between molecules and images. The molecular structure is composed of atoms and chemical bonds between them. The number of connections and the type of atoms are variable, while the image is a two-dimensional immutable pixel matrix. Molecules can be mapped into image form through two-dimensional tiling, but the rotation, displacement, and different choices of the main chain of the same molecule will result in different mapped images, which makes the molecular preprocessing very complicated and also makes the information required to learn by the generation network very redundant. With the development of graph neural networks in recent years, network information transmission based on graph structures has begun to enter the field of researchers. The graph structure is a data structure composed of nodes (vertices) and edges (connections) between nodes, which is used to represent the relationships and connections in various practical phenomena and problems. This representation is very consistent with the molecular structure. At the same time, combined with the graph neural network, it can learn to obtain the fixed-length feature network of molecules, greatly improving the effectiveness of generating molecules. Therefore, the present invention uses molecular graph representation as the training input and generation output of the generation network.

[0005] In recent years, based on the research results in the field of pharmaceutical design at home and abroad, molecular generation models can be roughly divided into the following three categories: (1) Methods based on variational auto-encoders, such as the Chemical Variational Auto-Encoder (CVAE) method and the Syntax-Directed VAE (SD-VAE) method; (2) Generative adversarial network-based methods, such as Mol-CycleGAN and ORGAN (objective reinforced generative adversarial network). (3) Methods based on diffusion models, such as the MDM (Molecular diffusion model) method.

[0006] Among the methods introduced above, the method based on generative adversarial networks has the characteristics of realistic generated samples, strong diversity, and no need for prior knowledge. Compared with other methods, it has certain advantages in fusion effect.

[0007] There are some key problems when using generative adversarial networks (GANs) for de novo molecular design in existing technologies. The main problems are: the generator only learns to generate limited molecular structures and cannot cover the diversity of the entire molecular space. For complex molecules, the adversarial relationship between the generator and discriminator of GAN can easily lead to oscillations or even non-convergence during training; GANs are usually unable to explicitly control the generation results and lack the ability to interpret the latent space of generated samples. In molecular design, scientists hope to understand the characteristics of the latent space and associate them with the physical and chemical properties of the molecules, but the generation process of GANs is difficult to provide such a clear mapping. The molecular structures generated by GANs may be illegal or have no chemical meaning, such as incorrect number of molecular bonds or unreasonable atomic connections. The generated samples require further legitimacy detection and screening, resulting in reduced efficiency and effectiveness of generation.

[0008] Therefore, it is necessary to provide a molecular de novo design method based on attribute classification to solve the above problems. Summary of the invention

[0009] The present invention provides a molecular de novo design method based on attribute classification to solve the existing problems.

[0010] A molecular de novo design method based on attribute classification of the present invention adopts the following technical solution, including: Classifying molecular data into different categories of molecular data according to molecular properties; According to the motifs of the molecular data, a motif vocabulary is obtained to map all molecules into a tree structure; Construct a molecular de novo design model, which includes an encoder-decoder model and an auxiliary classifier generative adversarial network model; The encoder-decoder model is trained to obtain a trained target encoder-decoder model, wherein during the training process: the molecular graph structure and motif of the molecular data are used as inputs of the encoder of the encoder-decoder model, the encoder compiles the molecular graph structure into a molecular vector in a latent space based on the motif, the molecular vector in the latent space is input into the decoder of the encoder-decoder model, and the decoder outputs the reconstructed molecular graph structure; The auxiliary classifier generative adversarial network model is trained, wherein the auxiliary classifier generative adversarial network model includes: a generator, a discriminator and an auxiliary classifier. During the training process, the generator is used to generate samples according to preset class labels and noise, and the molecular vectors in the latent space corresponding to the real samples and their corresponding class labels are input into the discriminator together, and the discriminator performs discrimination; the classification training of the auxiliary classifier is carried out by comparing the class labels of the generated samples with the class labels of the real samples, and helping the generator to generate diverse generated samples until the cross entropy loss of the discriminator loss and the generator loss is minimized, thereby obtaining a trained target auxiliary classifier generative adversarial network model; The trained target encoder-decoder model and the target auxiliary classifier generative adversarial network model are combined to obtain a trained target molecule de novo design model; The molecular graph structures and motifs of different categories of molecular data are input into the target molecule de novo design model, and the molecular graph structure reconstructed from the molecular data of this category is output.

[0011] Preferably, the molecular property is any one of lipophilicity, QED or SAS.

[0012] Preferably, the steps of obtaining the motif of the molecular data are: A motif is defined as a subgraph of molecular data induced by atoms and chemical bonds; Get all bridge bonds in the molecule data; According to all bridge bonds in the molecular data, all bridge connections of the molecular data are separated from adjacent bridges to obtain a set of disconnected subgraphs; The conditions for extracting motifs are: the subgraph appears a set number of times in the training set, where the training set refers to a subset of the ZINC dataset; If a subgraph does not appear the set number of times in the training set, the subgraph is decomposed into rings and bonds, and the rings and bonds of the subgraph are selected as motifs.

[0013] Preferably, the objective function when the encoder-decoder model is trained is:

[0014] In the formula, Represents the objective function when training the encoder-decoder model; Represents the posterior distribution of the molecule vector in the latent space The expectation operation performed; represents the approximate posterior distribution, which is used to approximate the true posterior distribution of the molecule vector in the latent space; represents the true posterior distribution of the molecule vector in the latent space; Represents the KL divergence between the approximate posterior distribution and the prior distribution; represents the prior distribution; represents the prior distribution; Represents the parameters of the decoder, which are used to control how to reconstruct the input data x from the latent space Z; Represents the parameters of the encoder, which are used to control how to map the input data x to the latent space Z; Represents the optimal objective function (loss function) of the encoder-decoder model.

[0015] Preferably, the loss function of the discriminator of the auxiliary classifier generation adversarial network model is:

[0016] In the formula, Represents the loss function value of the discriminator of the auxiliary classifier generative adversarial network model; Represents real data The expectation that the discriminator will correctly classify it as a "real" sample during training; D(x) represents the probability of the discriminator output, that is, the probability that the input data is a "real" sample; represents the expectation that when training on sample data generated by the generator, the discriminator hopes to correctly classify it as a "fake" sample; Indicates that given the real data, the discriminator class Model the probability distribution of the data and correctly predict the expectation of the data.

[0017] Preferably, the loss function of the generator of the auxiliary classifier generative adversarial network model is:

[0018] In the formula, represents the loss function of the generator of the auxiliary classifier generative adversarial network model; represents the expectation that the generator maximizes the “realistic score” given by the discriminator to the generated data; Represents the predicted probability of the generator for the generated data category; represents the hyperparameter of the regularization term; represents the reconstruction loss of the generator; Represents the real data in the discriminator; Data representing a generator.

[0019] Preferably, a graph neural network and a long short-term memory network are used to construct the encoder and decoder.

[0020] Preferably, the generator and discriminator of the adversarial network are constructed using a convolutional network.

[0021] A molecular de novo design system based on attribute classification, comprising: A data processing module is used to classify the molecular data into different categories of molecular data according to the molecular attributes; and to obtain a motif vocabulary that maps all molecules into a tree structure according to the motifs of the molecular data; A molecular de novo design model building module is used to build a molecular de novo design model, which includes an encoder-decoder model and an auxiliary classifier generative adversarial network model; The molecular de novo design model training module is used to train the encoder-decoder model to obtain a trained target encoder-decoder model, wherein during the training process: the molecular graph structure and motif of the molecular data are used as the input of the encoder of the encoder-decoder model, the encoder compiles the molecular graph structure into a molecular vector in the latent space based on the motif, the molecular vector in the latent space is input into the decoder of the encoder-decoder model, and the decoder outputs the reconstructed molecular graph structure; the auxiliary classifier generative adversarial network model is trained, wherein the auxiliary classifier generative adversarial network model includes: a generator, a discriminator and an auxiliary classifier, and during the training process, The generator is used to generate samples according to preset class labels and noise, and input the molecular vectors in the latent space corresponding to the real samples and their corresponding class labels into the discriminator, which makes the discrimination. The classification training of the auxiliary classifier compares the class labels of the generated samples with the class labels of the real samples, and helps the generator to generate diverse generated samples until the cross entropy loss of the discriminator loss and the generator loss is minimized, thereby obtaining a trained target auxiliary classifier generative adversarial network model. The trained target encoder-decoder model and the target auxiliary classifier generative adversarial network model are combined to obtain a trained target molecule de novo design model. The molecular reconstruction module is used to input the molecular graph structure and motif of different categories of molecular data into the target molecule de novo design model, and output the molecular graph structure after reconstructing the molecular data of this category.

[0022] The beneficial effects of the present invention are: In the present invention, molecules that meet the attribute category can be generated with high accuracy during the molecule generation process, mainly due to the design of the encoder-decoder model and the combination of the attribute classification task. Through training, the model can learn structural features related to specific attributes from molecular data, which makes the generated molecules not only meet the structural requirements, but also have the required biochemical properties. The encoder maps the molecular graph structure to the latent space, learns the implicit relationship between the latent variables and the attributes, and ensures that the molecular representation in the latent space can pass enough information to the decoder, which generates the corresponding molecular structure according to these latent variables. At the same time, by maximizing the reconstruction error, the model ensures that the generated molecular structure can be accurately reconstructed, thereby ensuring the validity and rationality of the graph structure. The KL divergence term, as a regularization means, limits the difference between the approximate posterior distribution and the prior distribution, so that the molecular representation in the latent space remains stable and consistent, thereby avoiding the generation of irregular or invalid molecules. These design steps work together to ensure that the generated molecules are not only structurally effective, but also can efficiently generate molecules that meet the requirements under different attribute categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A flowchart of a molecular de novo design method based on attribute classification according to the present invention; Figure 2 A structural diagram of a molecular de novo design model in an embodiment of the present invention; Figure 3 A structural diagram of an encoder-decoder model in an embodiment of the present invention; Figure 4 Schematic diagram of the implementation of the molecular de novo design method in the embodiment of the present invention; Figure 5 A molecule instance image generated for specifying molecule attributes in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] An embodiment of a molecular de novo design method based on attribute classification of the present invention is as follows: Figure 1 As shown, including: S1, classifying the molecular data into different categories of molecular data according to molecular attributes; Taking the LogP property as an example of molecular properties, this property is divided into six categories, and a single property is divided into 6 categories, with a total of three properties: LogP is used to measure the lipophilicity (fat solubility) of molecular compounds, and there is a good linear relationship between the oil-water partition coefficient of a drug and its absorption rate constant; generally, when oral drugs penetrate by passive diffusion, a LogP of 0-3 is considered to be the optimal absorption range for the human body. High LogP compounds have poor water solubility, while low LogP compounds have poor lipid permeability. Therefore, the low value interval (-∞, 0) is classified into one category, the high value interval (4, +∞) is classified into one category, and the range of LogP with high bioavailability is more finely divided into six types of lipid-water partition coefficients. After classification, LogP can be calculated as: (1) in, Indicates the fat-water partition coefficient to be calculated; Indicates the molecule The number of atoms of the type, Indicates the molecule Contribution of type atoms.

[0027] Taking the QED attribute (quantitative estimation of drug similarity) as an example, the distribution analysis of some key physicochemical properties of approved drugs is used to find evidence of the possibility of a molecule becoming a drug in drug discovery. The value range of this attribute is (0,1), and the higher the value, the more the molecule evaluated by combining multiple molecular descriptors meets the drug specifications. Based on the demand for high drug similarity and the distribution of this attribute under the zinc 250K data set, in this example, the low drug similarity of (0,0.5) is divided into one category, and the high drug similarity interval set is divided into 0.1, which is divided into 6 categories in total.

[0028] The molecular attributes are explained by taking SAS attributes as an example. SAS is used to determine whether a molecular compound has important properties that may become a drug. It consists of fragmentScore and complexitpenalty. It is obtained by analyzing the common structural features of a large number of synthetic molecules to capture "historical synthesis knowledge" and combining the complex structural features present in the molecule. This method characterizes the accessibility of molecular synthesis as a score between 1 (easy to synthesize) and 10 (difficult to synthesize). Although SAscore is widely distributed in the range of [2,8] in nature, the distribution of synthetic accessibility of biological molecules and recorded molecules is mainly concentrated in the range of [2,4.5]. Therefore, in this embodiment, the extremely easy synthesis interval [1,2] is divided into one category, the difficult synthesis interval [4.5,10] is divided into one category, and the easy synthesis drug interval 0.5 is divided into one category; this attribute is divided into 7 categories, and the oversampling method is used to balance the data.

[0029] S2, obtain the motifs and motif vocabulary of molecular data; In one embodiment, the motif Defined by Atoms and Chemical bond-induced molecular data A subgraph of . Given a molecular data , extract its motif , requiring that their combination will cover the entire molecular graph: and The motif is extracted by breaking all bridges that do not violate chemical validity, converting the molecule Decompose into disconnected fragments. The specific steps for extracting motifs are as follows: Step 21, u and v represent the atomic nodes in the molecular graph. Each atom is connected to other atoms in the molecular graph through chemical bonds (edges). The "bridge bond" here refers to the specific chemical bond connecting two atoms in the molecule (i.e. u and v). Get molecular data All bridge keys in , where u and v both have degree ≥2, and u or v is part of a ring. The degree of an atom refers to the number of chemical bonds that the atom has with other atoms, that is, the number of adjacent atoms to which the atom is connected. For example, if an atom is connected to two other atoms, its degree is 2; if it is connected to three other atoms, its degree is 3.

[0030] and refers to the degrees of atoms u and v, and ≥2, which means that the degrees of u and v must be greater than or equal to 2, that is, these two atoms are connected to at least two or more other atoms. This condition ensures that atoms u and v are "internal contact points" within the molecule, rather than atoms at the ends of the molecule.

[0031] Step 22: Separate all bridge connections from adjacent bridges, then the molecular data It becomes a set of disconnected subgraphs That is, according to the bridge bonds in the molecular graph, all bridge structures are separated from the entire molecule, and the connection between the bridge and the adjacent bridge is cut off to form an independent subgraph. The purpose of this is to extract specific structures (motifs) in the molecule in order to more clearly represent the key features of the molecule and provide support for subsequent model training and motif analysis.

[0032] Step 23: Set the condition for entering the motif as a subgraph The number of occurrences in the training set exceeds 100, wherein the training set refers to a subset in the ZINC data set, and the "training set" based on which the motif is extracted refers to a subset selected from the public chemical database (ZINC database) for model training and motif extraction. By setting a threshold for the number of times the subgraph appears in the training set, it is ensured that the extracted motif has sufficient representativeness and statistical significance, while avoiding the problem of insufficient data due to rare structures. This training set is a key source of data in the model development process. The ZINC data set is a publicly available chemical molecule database developed by the Irwin and Shoichet laboratories of the University of California, San Francisco (UCSF). It contains millions of drug-like molecules, provides free commercial procurement information, and is widely used in drug development and computational chemistry research. The molecular data of the ZINC data set can be used to construct a training set to meet research needs by screening specific types of molecules (such as molecular weight, lipophilicity or functional group characteristics). Therefore, the training set used for model training in this embodiment is screened from the ZINC data set according to research objectives and specific conditions.

[0033] Step 24: If no sub-image is selected As a motif, we further transform the subgraph Decomposition into rings and bonds, and selection of subgraphs Rings and bonds as molecular data The motif in .

[0034] S3, constructing a molecular de novo design model, and performing training to obtain a trained target molecular de novo design model; In this embodiment, the molecular de novo design model includes: an encoder-decoder model and an auxiliary classifier generative adversarial network model.

[0035] The steps to train a molecular de novo design model are: S31. Train the encoder-decoder model: Specifically, the encoder-decoder model is trained to obtain a trained target encoder-decoder model, wherein, during the training process: the molecular graph structure and motif of the molecular data are used as inputs of the encoder of the encoder-decoder model, the encoder compiles the molecular graph structure into a molecular vector in a latent space based on the motif, the molecular vector in the latent space is input into the decoder of the encoder-decoder model, and the decoder outputs the reconstructed molecular graph structure.

[0036] The structure diagram of the encoder-decoder model in this embodiment is as follows Figure 3 As shown in the figure, the encoder and decoder of the encoder-decoder model are constructed using graph neural networks and long short-term memory networks. The encoder-decoder model is a probabilistic model through variational inference. It belongs to a generative model and is an unsupervised model. In variational inference, in addition to the known data (observation data, training data), there is also an implicit variable. Here, the known data set is recorded as N continuous variables or discrete variables The unobserved random variable is denoted as z. Variational inference introduces the posterior distribution To jointly model, according to Bayes Theorem, derive the posterior distribution.

[0037] Among them, the prior distribution is defined in advance (e.g., standard normal distribution), and the posterior distribution We use a deep network to learn. As implicit features, the deep network can be regarded as a probabilistic decoder. The variational estimate of , it also uses a deep network to learn, and this deep network can be regarded as a probabilistic encoder.

[0038] For the encoder-decoder model, the negative of the evidence lower bound (ELBO) is the training objective function to be minimized, that is, the objective function is: (2) In the formula, Represents the objective function when training the encoder-decoder model; Represents the posterior distribution of the molecule vector in the latent space The expected operation performed indicates that sampling is required on the latent space Z. The sampling process is completed by the encoder, which outputs the parameters of the distribution in the latent space (such as mean and variance), and then samples the latent variables z from these distributions and sends them to the decoder for reconstruction; represents the approximate posterior distribution, parameterized by The parameters of the controlled distribution are generated by the encoder network. Specifically, the parameters generated by the encoder network include: mean μ(x) and standard deviation σ(x), μ(x) represents the central position of the latent variable z in the latent space. The standard deviation σ(x) represents the degree of expansion of the latent variable z in the latent space. These parameters control the distribution of latent variables in the latent space, which in turn affects the generation ability of the model. By minimizing the objective function in the variational lower bound (ELBO), the encoder learns how to generate these parameters so that the model can generate data effectively. Used to approximate the true posterior distribution of the molecule vector in the latent space; represents the true posterior distribution of the molecule vector in the latent space; Represents the KL divergence of the approximate posterior distribution and the prior distribution. The KL divergence is a non-negative measure of the difference between two probability distributions. The larger the value, the greater the difference between the two distributions. represents the prior distribution; Represents the parameters of the decoder, which are used to control how to reconstruct the input data x from the latent space Z; Represents the parameters of the encoder, which are used to control how to map the input data x to the latent space Z; represents the optimal objective function (loss function) of the encoder-decoder model; Represents the reconstruction error term, which is used to generate the log-likelihood of the observed data x given the molecular vector latent variable z in the latent space of this patent, which is determined by the parameter is the decoder parameter Control, control the generation process from the latent space Z to the observation space x. The decoder is a neural network whose weights and biases are key parameters for model learning. These parameters determine how to reconstruct the original input data x from the latent variables; it should be noted that the latent vector z is a specific point in the latent space Z, which is used to represent the implicit features of a single input data (such as a molecule) and is a numerical vector of fixed length; and the latent space is a multidimensional continuous space containing all possible latent vectors, which is used to describe the distribution and characteristics of the entire data. Simply put, the latent vector z is an instance in the latent space Z, and the latent space is the collection of all these latent vectors.

[0039] In the molecular encoding process, the encoder encodes the three layers (MPN) in the molecular hierarchy graph successively, and the MPN encoding process is represented as a parameterized ψ of , with parameters ψIncludes all weights and biases involved in the decoder and encoder, which control the mapping process from the initial embedding vectors of atoms and chemical bonds to atomic representations. These parameters are optimized by minimizing objective functions (such as reconstruction error, KL divergence, etc.) during training to learn effective atomic and molecular representations. Represented as a multi-layer neural network, its input is and The concatenation of x: represents the embedding vector of the atoms in the molecule, usually a vector related to the atomic features. Each atom i in the molecular graph has a corresponding embedding vector , represents the characteristics of the atom (such as atom type, electron configuration). y represents the embedding vector of the chemical bond in the molecule, which is usually a vector related to the chemical bond information. Each chemical bond (i, j) has a corresponding embedding vector in the molecular graph , representing the characteristics of the bond (e.g., bond type, bond strength).

[0040] At the atomic level, the input of MPN is the embedding vector of each atom and the embedding vector of the chemical bond in the molecular data. Through message passing and iteration, the network updates the representation of each atom and outputs the final atomic representation of each atomic layer.

[0041] The formula is as follows: (3) in: is the final representation of the i-th atom; Indicated by the parameter A controlled neural network to update the atomic representation; is the representation of atom i after the tth iteration; represents the neighbor atom j connected to atom i in round t; is the embedding vector of the chemical bond between atoms i and j.

[0042] The attachment layer is used to describe the relative position relationship between an atom and its neighboring atoms. During the encoding process, the input feature of a node is the sum of its atomic vector and the attachment layer embedding vector. The input feature of each edge of the attachment layer is the embedding vector that describes the relative order between nodes.

[0043] The formula is as follows: (4) is the output of the attachment layer node; is the input embedding vector of the atomic layer, is the atomic representation of the attached layer. The relative order between nodes is encoded by the embedding vector.

[0044] Next, the motif is calculated through a message passing iterative process. After T rounds of iterations, the final representation of each node is calculated. The motif representation can be obtained through message passing.

[0045] The formula is as follows: (5) is the representation of the motif; Represents the message passing process of the motif layer.

[0046] The encoding process of the motif layer is similar to that of the atomic layer and the attachment layer. The input feature of the node is the concatenation of the output node representation of the previous layer and the embedding vector. After message passing, the representation of the motif is finally obtained.

[0047] The formula is as follows: (6) is the final representation of each node in the motif layer; is the embedding vector of the atomic layer; is the node vector of the motif layer.

[0048] Based on the final representation of the motif layer, the latent vector of the latent space is calculated through the reparameterization technique. The latent vector can be obtained by sampling by calculating the mean and log variance.

[0049] The formula is as follows: (7) Where z is the final latent vector, representing the numerator; μ and σ are the mean and logarithmic variance calculated by the network, respectively; is a noise vector sampled from a standard normal distribution.

[0050] For the molecular decoding process, given a feature vector, a molecule is constructed by gradually attaching new motifs to the newly constructed molecular parts. The decoder operates hierarchically in a coarse-to-fine manner and makes three key continuous predictions at each step: the new motif selection, which part of it is connected, and the contact point with the current molecule. These decisions are highly coupled and modeled as autoregression. In addition, each step decision is made by the motif layer encoder to create the molecular part, and the features learned in the association layer are guided by the corresponding structure of the molecule.

[0051] Step 32: Train the auxiliary classifier to generate an adversarial network model: Specifically, the auxiliary classifier generative adversarial network model is trained, wherein the auxiliary classifier generative adversarial network model includes: a generator, a discriminator and an auxiliary classifier. During the training process, the generator is used to generate samples according to preset class labels and noise, and the molecular vectors in the latent space corresponding to the real samples and their corresponding class labels are input into the discriminator together, and the discriminator performs discrimination; the classification training of the auxiliary classifier compares the class labels of the generated samples with the class labels of the real samples, and helps the generator to generate diverse generated samples until the cross entropy loss of the discriminator loss and the generator loss is minimized, thereby obtaining a trained target auxiliary classifier generative adversarial network model.

[0052] In this embodiment, the generator and discriminator of the adversarial network constructed by the convolutional network are used, and the training process of the AC-GAN model includes two conventional steps: adversarial training of the generator and the discriminator, and classification training of the auxiliary classifier. In adversarial training, the generator and the discriminator compete with each other, the discriminator tries to distinguish between real samples and generated samples, and the generator tries to generate more realistic samples to deceive the discriminator. The classification training of the auxiliary classifier helps the generator generate samples with diversity by comparing the category labels of the generated samples with the category labels of the real samples.

[0053] Set a class label , the generator is based on the class label and noise Generate samples , and the latent vector corresponding to the real molecular data The class labels corresponding to the real molecular data are input into the discriminator, which makes the judgment .

[0054] During training, the training loss function of the discriminator part of the adversarial network is: (8) In the formula, Represents the loss function value of the discriminator of the auxiliary classifier generative adversarial network model; Represents real data The expectation that the discriminator will correctly classify it as a "real" sample during training; D(x) represents the probability of the discriminator output, that is, the probability that the input data is a "real" sample; It represents the expectation that the discriminator will correctly classify the sample data generated by the generator as a "forged" sample when training. For the generated sample, Indicates that given the real data, the discriminator class The probability distribution is modeled so that it correctly predicts the class of the data.

[0055] During training, the training loss function of the generator part of the adversarial network is: (9) In the formula, express; represents the expectation that the generator maximizes the “realistic score” given by the discriminator to the generated data; Represents the predicted probability of the generator for the generated data category; Represents the hyperparameter of the regularization term, which is used to balance the generation quality and reconstruction quality of the generator; Represents the reconstruction loss of the generator, which is used to ensure that the generated data is similar to the real data under a certain metric. The similarity is expressed using the Euclidean distance. Represents the real data in the discriminator; Data representing a generator.

[0056] When training the auxiliary classifier to generate the adversarial network model, the Adam optimizer is used to iteratively update the weights of the auxiliary classifier to generate the adversarial network model based on the training data. and the loss of the generator to generate "real" molecules Calculation uses a binary cross entropy loss function; for the molecular category loss calculation of the discriminator and generator, multiple cross entropy processing is performed to determine how close the actual output is to the expected output. It should be noted that in order to determine whether the training is complete, it is usually necessary to comprehensively consider the losses of the discriminator and the generator: (1) Balanced loss: Ideally, the losses of the discriminator and the generator should be relatively balanced. During adversarial training, the generator and the discriminator compete with each other, and it is usually necessary to keep the losses of both within a reasonable range. If the loss of the discriminator is too low and the loss of the generator is still high, it means that the generator has not yet learned effective features and the generated molecules are of poor quality. Conversely, if the loss of the generator is too low and the loss of the discriminator is high, it may mean that the molecules generated by the generator are overfitted and the model may need to be adjusted. (2) Convergence signs: The generator loss gradually decreases and tends to a stable value, indicating that the generator has learned to generate better molecules. At this time, the generator loss decreases and tends to be stable; the discriminator loss should tend to a lower value, and if the discriminator loss remains low, it means that the discriminator has learned to effectively distinguish between true and false molecules. At this time, the discriminator loss tends to a lower value. (3) Category loss: In addition to the basic adversarial loss of the discriminator and the generator, the category loss (multi-class cross-entropy loss) needs to be observed. The generator not only needs to generate real molecules, but also molecules that meet the requirements of specific categories. Therefore, if the category loss is still high, it may mean that the generator has not yet learned to generate molecules of the correct category well. At this time, you can focus on the category loss.

[0057] Step 33: Combine the trained target encoder-decoder model and the target auxiliary classifier generative adversarial network model to obtain Figure 2 The trained de novo design model of the target molecule is shown.

[0058] S4. Input the molecular graph structure and motif of different categories of molecular data into the target molecule de novo design model, and output the molecular graph structure after reconstructing the molecular data of this category.

[0059] Figure 4 This is an implementation framework diagram of the molecular de novo design method of an embodiment of the present invention. The molecular de novo design model is finally given to the user. The specified attributes are input into the generation model to obtain the potential feature vector of the pre-generated molecule, and the encoder gradually reconstructs the potential feature vector of the molecule into a graph-structured molecule. This end-to-end drug design method effectively improves the efficiency of pharmaceutical molecule design. Figure 5 The generated molecule instance image for the specified molecule properties. Figure 5 (a) is the actual molecular structure of material molecules in real life. Figure 5 (b) shows that the model proposed in this embodiment can generate molecules with the same properties as actual molecules but different structures, which greatly accelerates the research and development process of new materials.

[0060] A molecular de novo design system based on attribute classification comprises: a data processing module, a molecular de novo design model construction module, a molecular de novo design model training module and a molecular reconstruction module; the data processing module is used to classify molecular data into different categories of molecular data according to molecular attributes; a motif vocabulary that maps all molecules into a tree structure is obtained according to the motifs of the molecular data; the molecular de novo design model construction module is used to construct a molecular de novo design model, the molecular de novo design model comprises: an encoder-decoder model and an auxiliary classifier generative adversarial network model; the molecular de novo design model training module is used to train the encoder-decoder model to obtain a trained target encoder-decoder model, wherein, during the training process: the molecular graph structure and motif of the molecular data are used as inputs of the encoder of the encoder-decoder model, the encoder compiles the molecular graph structure into a molecular vector in a latent space based on the motif, the molecular vector in the latent space is input into the decoder of the encoder-decoder model, and the decoder outputs a reconstructed molecular graph structure; training the auxiliary classifier generative adversarial network model, wherein the auxiliary classifier generative adversarial network model includes: a generator, a discriminator and an auxiliary classifier. During the training process, the generator is used to generate samples according to preset class labels and noise, and the molecular vectors in the latent space corresponding to the real samples and their corresponding class labels are input into the discriminator together, and the discriminator performs discrimination; the classification training of the auxiliary classifier is performed by comparing the class labels of the generated samples with the class labels of the real samples, and helping the generator to generate diverse generated samples until the cross entropy loss of the discriminator loss and the generator loss is minimized, thereby obtaining a trained target auxiliary classifier generative adversarial network model; the trained target encoder-decoder model and the target auxiliary classifier generative adversarial network model are combined to obtain a trained target molecule de novo design model; the molecular reconstruction module is used to input the molecular graph structure and motif of molecular data of different categories into the target molecule de novo design model, and output the molecular graph structure after the molecular data of this category is reconstructed.

[0061] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for de novo molecular design based on attribute classification, characterized in that: include: Classifying molecular data into different categories of molecular data according to molecular properties; According to the motifs of the molecular data, a motif vocabulary is obtained to map all molecules into a tree structure; Construct a molecular de novo design model, which includes an encoder-decoder model and an auxiliary classifier generative adversarial network model; The encoder-decoder model is trained to obtain a trained target encoder-decoder model, wherein during the training process: the molecular graph structure and motif of the molecular data are used as inputs of the encoder of the encoder-decoder model, the encoder compiles the molecular graph structure into a molecular vector in a latent space based on the motif, the molecular vector in the latent space is input into the decoder of the encoder-decoder model, and the decoder outputs the reconstructed molecular graph structure; The auxiliary classifier generative adversarial network model is trained, wherein the auxiliary classifier generative adversarial network model includes: a generator, a discriminator and an auxiliary classifier. During the training process, the generator is used to generate samples according to preset class labels and noise, and the molecular vectors in the latent space corresponding to the real samples and their corresponding class labels are input into the discriminator together, and the discriminator performs discrimination; the classification training of the auxiliary classifier is carried out by comparing the class labels of the generated samples with the class labels of the real samples, and helping the generator to generate diverse generated samples until the cross entropy loss of the discriminator loss and the generator loss is minimized, thereby obtaining a trained target auxiliary classifier generative adversarial network model; The trained target encoder-decoder model and the target auxiliary classifier generative adversarial network model are combined to obtain a trained target molecule de novo design model; The molecular graph structures and motifs of different categories of molecular data are input into the target molecule de novo design model, and the molecular graph structure reconstructed from the molecular data of this category is output.

2. A method for de novo molecular design based on attribute classification according to claim 1, characterized in that: The molecular property is any one of lipophilicity, QED or SAS.

3. A method for de novo molecular design based on attribute classification according to claim 1, characterized in that: The steps to obtain the motif of molecular data are: A motif is defined as a subgraph of molecular data induced by atoms and chemical bonds; Get all bridge bonds in the molecule data; According to all bridge bonds in the molecular data, all bridge connections of the molecular data are separated from adjacent bridges to obtain a set of disconnected subgraphs; The conditions for extracting motifs are: the subgraph appears a set number of times in the training set, where the training set refers to a subset of the ZINC dataset; If a subgraph does not appear the set number of times in the training set, the subgraph is decomposed into rings and bonds, and the rings and bonds of the subgraph are selected as motifs.

4. The method for de novo molecular design based on attribute classification according to claim 1, characterized in that: The objective function for training the encoder-decoder model is: In the formula, Represents the objective function when training the encoder-decoder model; Represents the posterior distribution of the molecule vector in the latent space The expectation operation performed; represents the approximate posterior distribution, which is used to approximate the true posterior distribution of the molecule vector in the latent space; represents the true posterior distribution of the molecule vector in the latent space; Represents the KL divergence between the approximate posterior distribution and the prior distribution; represents the prior distribution; Represents the parameters of the decoder, which are used to control how to reconstruct the input data x from the latent space Z; Represents the parameters of the encoder, which are used to control how to map the input data x to the latent space Z; Represents the optimal objective function (loss function) of the encoder-decoder model.

5. The method for de novo molecular design based on attribute classification according to claim 1, characterized in that: The loss function of the discriminator of the auxiliary classifier generation adversarial network model is: In the formula, Represents the loss function value of the discriminator of the auxiliary classifier generative adversarial network model; It represents the expectation that the discriminator will correctly classify the real data as a "real" sample when training it; D(x) represents the probability of the discriminator output, that is, the probability that the input data is a "real" sample; represents the expectation that when training on sample data generated by the generator, the discriminator hopes to correctly classify it as a "fake" sample; Indicates that given the real data, the discriminator class The probability distribution is modeled so that it correctly predicts the class of the data.

6. The method for de novo molecular design based on attribute classification according to claim 1, characterized in that: The loss function of the generator of the auxiliary classifier generative adversarial network model is: In the formula, represents the loss function of the generator of the auxiliary classifier generative adversarial network model; represents the expectation that the generator maximizes the "authenticity score" given by the discriminator to the generated data; Represents the predicted probability of the generator for the generated data category; represents the hyperparameter of the regularization term; represents the reconstruction loss of the generator; Represents the real data in the discriminator; Data representing a generator.

7. The method for de novo molecular design based on attribute classification according to claim 1, characterized in that: Graph neural network and long short-term memory network are used to construct encoder and decoder.

8. The method for de novo molecular design based on attribute classification according to claim 1, characterized in that: Generator and discriminator of an adversarial network built using convolutional networks.

9. A molecular de novo design system based on attribute classification, characterized in that: include: A data processing module, used for classifying the molecular data into different categories of molecular data according to molecular properties; According to the motifs of the molecular data, a motif vocabulary is obtained to map all molecules into a tree structure; A molecular de novo design model building module is used to build a molecular de novo design model, which includes an encoder-decoder model and an auxiliary classifier generative adversarial network model; The molecular de novo design model training module is used to train the encoder-decoder model to obtain a trained target encoder-decoder model, wherein during the training process: the molecular graph structure and motif of the molecular data are used as the input of the encoder of the encoder-decoder model, the encoder compiles the molecular graph structure into a molecular vector in the latent space based on the motif, the molecular vector in the latent space is input into the decoder of the encoder-decoder model, and the decoder outputs the reconstructed molecular graph structure; the auxiliary classifier generative adversarial network model is trained, wherein the auxiliary classifier generative adversarial network model includes: a generator, a discriminator and an auxiliary classifier, and during the training process, The generator is used to generate samples according to preset class labels and noise, and input the molecular vectors in the latent space corresponding to the real samples and their corresponding class labels into the discriminator, which makes the discrimination. The classification training of the auxiliary classifier compares the class labels of the generated samples with the class labels of the real samples, and helps the generator to generate diverse generated samples until the cross entropy loss of the discriminator loss and the generator loss is minimized, thereby obtaining a trained target auxiliary classifier generative adversarial network model. The trained target encoder-decoder model and the target auxiliary classifier generative adversarial network model are combined to obtain a trained target molecule de novo design model. The molecular reconstruction module is used to input the molecular graph structure and motif of different categories of molecular data into the target molecule de novo design model, and output the molecular graph structure after reconstructing the molecular data of this category.