Training method, device, equipment and medium for molecular generation model

By reconstructing and error training of the molecular generation model, the problem of insufficient diversity of candidate molecules is solved, and the effect of molecular generation model being able to optimize multiple attributes is achieved.

CN115116562BActive Publication Date: 2025-08-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210405951.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-08-12
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

In the prior art, when training molecular generation models, the diversity of candidate molecules is insufficient, resulting in the model that can only optimize 1 to 2 properties of the molecule and cannot meet the optimization needs of multiple properties.

Method used

Reconstructed molecules are generated by screening molecules, changing the chemical structure of the screening molecules, and using the errors between the generated molecules and the reconstructed molecules to train the molecular generation model to ensure the diversity of data during the training process and rich training labels.

Benefits of technology

The molecular generation model is implemented to optimize multiple attributes, and the generated molecules can meet the needs of multiple attributes, increasing the number and yield of generated molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116562B_ABST
    Figure CN115116562B_ABST
Patent Text Reader

Abstract

This application discloses a training method, apparatus, device, and medium for a molecular generation model, which is applied to the field of machine learning. The method comprises: screening candidate molecules according to molecular optimization conditions to obtain screening molecules, wherein the molecular optimization conditions are used to evaluate the properties of the candidate molecules; reconstructing the screening molecules to generate reconstructed molecules, wherein the reconstructed molecules meet the molecular optimization conditions and the chemical structure of the reconstructed molecules is different from the chemical structure of the screening molecules; adjusting the chemical structure of the screening molecules using the molecular generation model to obtain generated molecules; and training the molecular generation model based on the error between the generated molecules and the reconstructed molecules. The molecular generation model trained by this method can output molecules that meet multiple properties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning, and in particular to a method, apparatus, device, and medium for training a molecular generation model. Background Art

[0002] During the drug development phase, while maintaining the lead compound's target activity, one or more drugability characteristics are optimized to obtain a molecule with the potential to enter clinical research. A lead compound is a compound with specific biological activity and chemical structure. Structural modification and engineering of the lead compound can lead to the development of an ideal drug. Drugability refers to the characteristics of the lead compound that demonstrate potential for drug development, based on preliminary pharmacodynamic studies, pharmacokinetic properties, and early safety evaluations.

[0003] When training a molecular generative model, related technologies first use an attribute discriminator to screen candidate molecules that meet optimization criteria, generating screened molecules. These screened molecules are then fed into the molecular generative model, which outputs generated molecules. The error between the screened and generated molecules is then calculated, and the molecular generative model is trained based on this error.

[0004] Since the relevant technology only uses the information of candidate molecules during the training process, the diversity of candidate molecules is insufficient, and the amount of information that candidate molecules can provide is relatively limited, the resulting molecular generation model can only optimize 1 to 2 properties of the molecule. Summary of the Invention

[0005] The present application provides a method, apparatus, device, and medium for training a molecular generative model. The molecular generative model trained by this method can optimize multiple properties of molecules. The technical solution is as follows:

[0006] According to one aspect of the present application, a method for training a molecular generative model is provided, the method comprising:

[0007] screening candidate molecules according to molecular optimization conditions to obtain screening molecules, wherein the molecular optimization conditions are used to evaluate the properties of the candidate molecules;

[0008] reconstructing the screening molecule to generate a reconstructed molecule, wherein the reconstructed molecule satisfies the molecule optimization condition and the chemical structure of the reconstructed molecule is different from the chemical structure of the screening molecule;

[0009] Adjusting the chemical structure of the screening molecule using the molecular generation model to obtain a generated molecule;

[0010] The molecule generation model is trained according to the error between the generated molecule and the reconstructed molecule.

[0011] According to one aspect of the present application, a method for training a molecular generative model is provided, the method comprising:

[0012] According to another aspect of the present application, a computer device is provided, comprising: a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the training method of the molecular generation model as described above.

[0013] According to another aspect of the present application, a computer storage medium is provided, wherein the computer-readable storage medium stores at least one program code, and the program code is loaded and executed by a processor to implement the training method of the molecular generation model as described above.

[0014] According to another aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the method for training a molecular generative model as described above.

[0015] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:

[0016] When training the molecular generative model, the screened molecules are reconstructed, changing their chemical structure to produce reconstructed molecules that meet the molecular optimization criteria. The screened molecules are then fed into the molecular generative model to generate generated molecules, and the model is trained using the error between the generated and reconstructed molecules. Because the chemical structure of the reconstructed molecules differs from that of the screened molecules, the reconstructed molecules ensure data diversity during training, providing the molecular generative model with a rich set of training labels. These rich training labels provide training directions for a variety of optimized properties, allowing the molecules generated by this molecular generative model to meet a variety of properties. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 is a schematic diagram of a computer system provided by an exemplary embodiment of the present application;

[0019] Figure 2 is a schematic diagram of a method for training a molecular generative model provided by an exemplary embodiment of the present application;

[0020] Figure 3 is a schematic diagram of a process for training a molecular generation model provided by an exemplary embodiment of the present application;

[0021] Figure 4 is a schematic diagram of a reconstructed molecule provided by an exemplary embodiment of the present application;

[0022] Figure 5 is a schematic diagram of a VAE model provided by an exemplary embodiment of the present application;

[0023] Figure 6 1 is a flow chart of a method for training a molecular generation model provided by an exemplary embodiment of the present application;

[0024] Figure 7 1 is a flow chart of a method for training a molecular generation model provided by an exemplary embodiment of the present application;

[0025] Figure 8 1 is a schematic structural diagram of a training device for a molecular generation model provided by an exemplary embodiment of the present application;

[0026] Figure 9 It is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0028] First, the nouns involved in the embodiments of this application are introduced:

[0029] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0030] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0031] Lead compounds are compounds with specific biological activity and chemical structures obtained through various pathways and methods. These compounds are used for further structural modification and serve as the starting point for modern drug research. Because lead compounds often have certain drawbacks, such as insufficient activity, unstable chemical structure, high toxicity, poor selectivity, and unsuitable pharmacokinetic properties, they require chemical modification and further optimization to develop them into ideal drugs. This process is called lead compound optimization.

[0032] Druggability refers to the potential for drug development based on preliminary pharmacodynamic studies, pharmacokinetic properties, and early safety evaluations. In one implementation, druggability can be measured using the Five Rules of Druggability. These are empirical rules for estimating a compound's druggability based on its molecular weight, hydrogen bond donors, hydrogen bond acceptors, and calculated partition coefficient limits.

[0033] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the molecules involved in this application were obtained with full authorization.

[0034] Figure 1 FIG. 1 is a block diagram of a computer system according to an exemplary embodiment of the present application. The computer system 100 includes a terminal 120 and a server 140 .

[0035] Terminal 120 has an application program related to model generation installed. This application program can be a small program within an app (application), a dedicated application, or a web client. Terminal 120 is at least one of a smartphone, a tablet computer, an e-book reader, an MP3 player, an MP4 player, a laptop computer, and a desktop computer.

[0036] The terminal 120 is connected to the server 140 via a wireless network or a wired network.

[0037] The server 140 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server 140 is used to provide background services for the model generation application and send information related to the model generation to the terminal 120. Optionally, the server 140 undertakes the primary computing work and the terminal 120 undertakes the secondary computing work; alternatively, the server 140 undertakes the secondary computing work and the terminal 120 undertakes the primary computing work; alternatively, the server 140 and the terminal 120 both adopt a distributed computing architecture for collaborative computing.

[0038] Figure 2 FIG. 1 shows a schematic diagram of a method for training a molecular generation model provided by an exemplary embodiment of the present application. Figure 2 As shown, when training the molecule generation model 205 , the molecule reconstruction model 201 is needed.

[0039] The molecular reconstruction model 201 includes two functions: sampling generation and reconstruction generation. Sampling generation means that the molecular reconstruction model 201 generates candidate molecules 202 through sampling processing. Reconstruction generation means that the molecular reconstruction model 201 reconstructs the screening molecules 203 to generate reconstructed molecules 204. Optionally, the molecular reconstruction model 201 is at least one of a VAE model (Variational Auto Encoder), a GVAE model (Graph Variational Auto Encoder), a MolGAN model (Molecule Generative Adversarial Networks), a MolFlow model (MoleculeFlow-based, a flow-based chemical molecule generation model), a GraphAF model (GraphAutoregressive Flow-basedmodel, a flow-based image generation autoregressive model), and a GraphNVP model (GraphNon Volume Preserving, a reversible flow model for image generation). It should be noted that the embodiment of the present application does not specifically limit the type of the molecular reconstruction model 201.

[0040] The candidate molecules 202 are screened to obtain screening molecules 203 to obtain molecules that meet the molecular optimization conditions. The molecular optimization conditions are used to evaluate the properties of the candidate molecules 202. Exemplarily, the molecular optimization conditions include that the molecular activity of the candidate molecules 202 needs to be greater than a preset activity threshold, or the molecular optimization conditions include that the candidate molecules 202 meet druggability, or the molecular optimization conditions include that the candidate molecules 202 need to meet a preset a type of properties, where a is a positive integer, wherein the a type of property refers to an inherent characteristic of the molecule. Exemplarily, the a type of property can be that the biological activity of the molecule is greater than a preset activity threshold, or that the molecule can bind to the target, or that the molecule is non-toxic to organisms, or that the molecule is water-soluble, or that the molecule affects cell metabolism.

[0041] In an optional implementation, candidate molecules 202 are screened using an attribute judgment model to obtain screened molecules 203, or alternatively, candidate molecules 202 are screened using an activity judgment model to obtain screened molecules 203. The attribute judgment model is used to determine whether candidate molecule 202 satisfies a type of attribute. Exemplarily, the attribute judgment model processes data on candidate molecule 202 to obtain an attribute judgment result. In a specific example, the attribute judgment result is [1, 0, 1, 1, 1], where each digit of the attribute judgment result represents the judgment result for a specific attribute, with "1" indicating that candidate molecule 202 possesses that attribute, and "0" indicating that candidate molecule 202 does not possess that attribute. The activity judgment model is used to determine the activity of candidate molecule 202, where the activity value includes, but is not limited to, at least one of an enzyme activity value, a bactericidal activity value, an antibacterial activity value, and a protein activity value. Exemplarily, the activity judgment model processes data on candidate molecule 202 to obtain an activity value for candidate molecule 202. The aforementioned attribute judgment model or activity judgment model is an AI model or a computational chemistry model. In an optional implementation, a simulation experiment is performed on the candidate molecule 202 to obtain an experimental result; and the candidate molecule 202 is screened based on the experimental result to obtain the screened molecule 203 .

[0042] The screening molecule 203 is reconstructed to obtain a reconstructed molecule 204, which satisfies the molecular optimization conditions and has a chemical structure different from that of the screening molecule 203. In this embodiment of the present application, the screening molecule 203 is reconstructed using the molecular reconstruction model 201 to generate the reconstructed molecule 204.

[0043] In the model training stage, the molecule generation model 205 optimizes the model parameters of the molecule generation model 205 based on the screening molecules 203 and the reconstruction molecules 204. Exemplarily, the screening molecules 203 are subjected to data processing by the molecule generation model 205 to obtain generated molecules (not shown in the figure); the molecule generation model 205 is trained according to the error between the generated molecules and the reconstructed molecules 204 to optimize the model parameters of the molecule generation model 205. Optionally, the molecule generation model 205 is at least one of a VAE model, a GVAE model, a MolGAN model, a MolFlow model, a GraphAF model, and a GraphNVP model. It should be noted that the embodiment of the present application does not specifically limit the type of the molecule reconstruction model 201. Optionally, the types of the molecule reconstruction model 201 and the molecule generation model 205 are the same or different.

[0044] It should be noted that the process of training the molecule generation model 205 is an iterative process. In the i-th (i is a positive integer) iteration, after the molecule generation model 205 obtains the generated molecule, the molecule generation model 205 performs sampling processing on the generated molecule to obtain the candidate molecule 202 of the i+1th round.

[0045] Figure 3 FIG1 shows a flow chart of a method for training a molecular generation model provided by an exemplary embodiment of the present application. Figure 1 The computer system 100 shown is executed, and the method includes:

[0046] Step 302: Screening candidate molecules to obtain screened molecules based on the molecular optimization conditions, where the molecular optimization conditions are used to evaluate properties of the candidate molecules.

[0047] Optionally, the molecular optimization condition includes the optimization value of the screening molecule being the largest N optimization values among the candidate molecules, where N is a positive integer. Exemplarily, the optimization values of the candidate molecules are calculated; the candidate molecules corresponding to the N largest optimization values are used as the screening molecules. N is related to the number of screening molecules, for example, N = number of screening molecules × 40%. In a specific embodiment, the optimization value includes, but is not limited to, at least one of an enzyme activity value, a bactericidal activity value, an antibacterial activity value, and a protein activity value.

[0048] Optionally, the molecular optimization condition includes the optimization value of the candidate molecule being within a preset range. Exemplarily, the optimization value of the candidate molecule is calculated; and the candidate molecule corresponding to the optimization value within the preset range is determined as the screening molecule. The preset range can be set by a technician based on actual needs.

[0049] Optionally, the molecular optimization condition includes that the optimization value of the candidate molecule is greater than a preset optimization value. Exemplarily, the optimization value of the candidate molecule is calculated; and the candidate molecule corresponding to the optimization value greater than the preset optimization value is used as the screening molecule.

[0050] Optionally, screening candidate molecules to obtain screening molecules is achieved through an AI model. Exemplarily, the candidate molecules are screened by a screening model to obtain screening molecules, and the screening model is an attribute judgment model or a molecular activity judgment model. Exemplarily, the candidate molecules are subjected to data processing by the attribute judgment model to obtain an attribute judgment result. In a specific example, the attribute judgment result is [1, 0, 1, 1, 1], and each bit of the attribute judgment result is used to represent the judgment result of an attribute, "1" indicates that the candidate molecule has the attribute, and "0" indicates that the candidate molecule does not have the attribute.

[0051] In an optional implementation, the candidate molecule is a molecule sampled from a molecular space, where the molecular space includes molecules from a public dataset. Optionally, the public dataset includes chemical structures of at least two molecules.

[0052] In an optional implementation, Figure 2 As shown, in the first iteration of the training molecule generation model 205, the candidate molecule 202 is a molecule sampled from the second sampling space of the molecular reconstruction model 201. Alternatively, in the i-th iteration of the training molecule generation model 205, the candidate molecule 202 is a molecule generated in the i-1-th iteration.

[0053] Optionally, this step also includes molecular deduplication and legality verification. Molecular deduplication is used to remove duplicate molecules from candidate molecules. For example, duplicate molecules from candidate molecules are removed. Legality verification is used to determine whether the candidate molecules comply with the rules for molecular connection. For example, molecules that do not comply with chemical rules from candidate molecules are removed; or, molecules whose molecular graphs are not connected from candidate molecules are removed. For example, the carbon atoms in candidate molecule 1 form 5 covalent bonds with 5 hydrogen atoms respectively, but according to the bonding rules of carbon atoms, carbon atoms cannot form 5 covalent bonds. Therefore, candidate molecule 1 does not comply with the chemical rules.

[0054] Step 304: reconstructing the screening molecule to generate a reconstructed molecule, where the reconstructed molecule meets the molecular optimization condition and the chemical structure of the reconstructed molecule is different from the chemical structure of the screening molecule.

[0055] Alternatively, as Figure 2 As shown, the screening molecules are reconstructed using the molecular reconstruction model 201 to generate reconstructed molecules 204 .

[0056] Alternatively, molecular remodeling refers to adjusting other regions of the chemical structure while retaining the regions that meet the optimization criteria in the screening molecule. For example, the chemical structure of the screening molecule includes region A and region B, where region A enables the screening molecule to meet the optimization criteria, while region B does not affect the optimization criteria. Therefore, the chemical structure of region B is adjusted.

[0057] The screening molecule is composed of at least two atoms. Optionally, the method of reconstructing the screening molecule includes but is not limited to at least one of deleting an atom, replacing an atom, and adding an atom. For example, one H atom in the screening molecule is replaced with -OH (hydroxyl).

[0058] It should be noted that although the screening molecules and the reconstructed molecules have different chemical structures, the chemical structures of the screening molecules and the reconstructed molecules are similar overall, and the chemical structures of the screening molecules and the reconstructed molecules differ in some regions. For example, Figure 4 As shown, the overall similarity of the chemical structures of the screening molecule 401 and the reconstructed molecule 402 is high, but there are differences between the screening molecule 401 and the reconstructed molecule 402 at the fragments, and the atoms in the dotted boxes of the screening molecule 401 and the reconstructed molecule 402 are different.

[0059] Step 306: Adjust the chemical structure of the screened molecule using the molecular generation model to obtain a generated molecule.

[0060] Optionally, the molecular generative model is at least one of a VAE model, a GVAE model, a MolGAN model, a MolFlow model, a GraphAF model, and a GraphNVP model. It should be noted that the embodiment of the present application does not specifically limit the type of the molecular generative model.

[0061] Optionally, because the molecular generative model is still being trained, the adjustment of the chemical structure of the screened molecules by the molecular generative model is non-directional, compared to the molecular reconstruction described above. When adjusting the chemical structure of the screened molecules, the chemical structure of any region of the screened molecules is adjusted. As the molecular generative model is trained through iterative cycles, the process of adjusting the chemical structure of the screened molecules by the molecular generative model will approach the molecular reconstruction described above as the number of iterations increases.

[0062] Optionally, the manner of adjusting the chemical structure of the screening molecule includes but is not limited to at least one of deleting atoms, replacing atoms, and adding atoms.

[0063] Step 308: Train the molecule generation model based on the error between the generated molecule and the reconstructed molecule.

[0064] Optionally, the molecular generation model is trained using an error back-propagation algorithm based on the error between the generated molecules and the reconstructed molecules. In one optional implementation, the generated molecules are converted into generated molecule feature vectors, and the reconstructed molecules are converted into reconstructed molecule feature vectors; a normal distribution of the generated molecule feature vectors is fitted to obtain a generated molecule normal distribution, and a normal distribution of the reconstructed molecule feature vectors is fitted to obtain a reconstructed molecule normal distribution; and the error between the generated and reconstructed molecules is obtained based on the relative entropy (Kullback-Leibler divergence, KL divergence) between the generated and reconstructed molecule normal distributions.

[0065] In one optional implementation, when a training completion condition for the molecular generation model is met, the molecular generation model is used to sample from the first sampling space to obtain a target generation molecule meeting the molecular optimization condition. Alternatively, when a training completion condition for the molecular generation model is met, the molecular generation model is used to sample from the first sampling space to obtain a target vector; the target vector is decoded to obtain a target generation molecule meeting the molecular optimization condition. The target generation molecule is a lead compound.

[0066] In summary, when training the molecular generation model, the embodiment of the present application reconstructs the screening molecules, changes the chemical structure of the screening molecules, obtains reconstructed molecules, and the reconstructed molecules meet the molecular optimization conditions. The screening molecules are input into the molecular generation model to obtain generated molecules, and the molecular generation model is trained by the error between the generated molecules and the reconstructed molecules. Since the chemical structure of the reconstructed molecules is different from that of the screening molecules, the reconstructed molecules ensure the data diversity in the training process, provide rich training labels for the molecular generation model, and the rich training labels provide a variety of training directions for optimizing properties, so that the molecules generated by the molecular generation model can meet a variety of properties.

[0067] In the following embodiments, the structure of the first generative model will be introduced by taking the molecular generative model being a VAE model as an example. Figure 5 FIG. 5 is a schematic diagram of a molecular generation model provided by an exemplary embodiment of the present application. The molecular generation model 500 includes a tuning network 501 and a first decoding network 502 .

[0068] The tuning network 501 is used to generate a normal distribution 508 of latent features of the screening molecule 503, where the latent features are used to describe features in the screening molecule that are related to the molecular optimization conditions. Exemplarily, the tuning network 501 performs encoding and sampling processing on the screening molecule 503 to obtain a first sampling vector corresponding to the screening molecule 503. The tuning network 501 encodes the first encoding vector of the screening molecule based on the chemical structure of the screening molecule, and the first encoding vector is used to represent the features of the screening molecule. The tuning network 501 fits the mean of the first encoding vector to obtain a first mean 504, which is used to represent the mean of the latent features of the first encoding vector. The tuning network 501 fits the variance of the first encoding vector to obtain a first variance 505, which is used to represent the variance of the latent features of the first encoding vector. Thereafter, the normal distribution of the latent features is determined based on the first mean 504 and the first variance 505.

[0069] The tuning network 501 is further configured to sample from the normal distribution 508 of the latent features to obtain a first sampling vector 506. The first sampling vector 506 can be decoded into a generated molecule output by the molecular generation model. Optionally, the sampling method is random sampling or sampling according to a preset rule.

[0070] The first decoding network 502 is used to decode the first sampling vector 506 to obtain a generated numerator 507. In an optional implementation, the first decoding network 502 is a fully connected feed-forward network.

[0071] Exemplarily, after molecular generation model 500 is trained, it stores a normal distribution 508 of the aforementioned latent features. Because latent features are used to describe features in the screened molecules 503 that are relevant to the molecular optimization conditions, normal distribution 508 obtained by training molecular generation model 500 includes the latent features of the aforementioned molecular optimization conditions. Therefore, first sampling vector 506 sampled from normal distribution 508 also includes the latent features of the aforementioned molecular optimization conditions. Accordingly, generated molecules 507 decoded from first sampling vector 506 satisfy the molecular optimization conditions.

[0072] In the following examples, the molecular generation model is trained iteratively. In the first iteration, the candidate molecules are molecules sampled from the second sampling space of the molecular reconstruction model. In the i-th iteration, the candidate molecules are the molecules generated in the i-1th iteration, where i is an integer greater than 1. For example, the molecular reconstruction model and the molecular generation model are both VAE models. In this case, the molecular generation model includes a tuning network and a first decoding network:

[0073] Figure 6 FIG1 shows a flow chart of a method for training a molecular generation model provided by an exemplary embodiment of the present application. Figure 1The computer system 100 shown is executed, and the method includes:

[0074] Step 601: In the first iteration, according to the molecular optimization conditions, candidate molecules are screened to obtain screened molecules, where the candidate molecules are molecules sampled from the second sampling space of the molecular reconstruction model.

[0075] Optionally, the molecular reconstruction model is trained based on molecules in a public dataset. Exemplarily, the public dataset includes at least 6.8 million molecules.

[0076] Optionally, the molecular reconstruction model is at least one of a VAE model, a GVAE model, a MolGAN model, a MolFlow model, a GraphAF model, and a GraphNVP model. It should be noted that the embodiment of the present application does not specifically limit the type of the molecular reconstruction model.

[0077] For example, the candidate molecule is recorded as X0, and the molecular optimization condition is recorded as y0, then the screening molecule that meets the molecular optimization condition is Y0=0⊙X0, where ⊙ represents an exclusive OR operation.

[0078] Step 602: In the i-th iteration, according to the molecule optimization condition, the molecules generated in the previous round are screened to obtain screened molecules, where the molecules generated in the previous round are the molecules generated in the i-1-th iteration.

[0079] Optionally, the chemical structure of the screened molecule in the i-1th iteration is adjusted by the molecular generation model to obtain the generated molecule in the i-1th iteration.

[0080] During the iterative process, if the number of molecules obtained in this round of screening is too small, the number of screening molecules input into the molecule generation model will be too small, which is not conducive to the subsequent iterative training. Optionally, when the number of molecules screened in this round is less than the quantity threshold in the j-th round of iteration, the range of the molecular optimization condition is adjusted, and j is an integer less than i; wherein, in the iterative process after the j-th round of iteration, the range of the molecular optimization condition shrinks as the number of iterations increases. Exemplarily, the molecular optimization condition includes that when the optimization value of the screening molecule is the largest N1 optimization value among the candidate molecules, the number of screening molecules generated is less than the quantity threshold, and after adjusting N1 to N2 (N2 is less than N1), the number of screening molecules generated is greater than the quantity threshold. Exemplarily, the molecular optimization condition includes that when the screening molecule meets M1 preset attributes, the number of screening molecules generated is less than the quantity threshold, and after adjusting M1 to M2 (M2 is less than M1), the number of screening molecules generated is greater than the quantity threshold.

[0081] It should be noted that step 601 and step 602 are mutually exclusive. When step 601 is executed, step 602 is not executed; when step 602 is executed, step 601 is not executed.

[0082] Step 603: Perform encoding and sampling processing on the screening molecules through the reconstruction network in the molecular reconstruction model to obtain a second sampling vector corresponding to the screening molecules.

[0083] Optionally, this step includes the following sub-steps:

[0084] 1. By reconstructing the network, the second encoding vector of the screening molecule is encoded according to the chemical structure of the screening molecule.

[0085] Exemplarily, by reconstructing the network, according to the chemical structure of the screening molecule and the SMILES specification (Simplified Molecular Input Line Entry System), a second encoding vector of the screening molecule is encoded. Exemplarily, propionic acid is represented as CCC(=O)O, and cyclohexane is represented as C1CCCCC1.

[0086] 2. By reconstructing the network, a posterior distribution hypothesis is made on the second encoding vector to generate a second sampling space corresponding to the second encoding vector.

[0087] The posterior distribution hypothesis is used to assume that the latent features of the second encoding vector are normally distributed. In this case, the latent features are distributed in a second sampling space, which can be obtained by fitting the mean and variance of the second encoding vector.

[0088] Exemplarily, the second mean is obtained by fitting the mean of the latent features of the second coding vector through the second mean sub-network, and the second mean is used to represent the mean of the latent features of the second coding vector. Optionally, the second mean sub-network is at least one of a fully connected feedforward network, a transformer (a type of codec model), and an RNN model. It should be noted that the embodiment of the present application uses a machine learning model to fit the mean of the second coding vector, and the type of machine learning model can be determined by technicians according to actual needs.

[0089] Exemplarily, the variance of the second encoding vector is fitted through the second variance subnetwork to obtain the second variance, and the second mean is used to represent the variance of the latent features of the second encoding vector.

[0090] Exemplarily, a second normal distribution is determined based on the second mean and the second variance by generating a sub-network through the second distribution.

[0091] 3. Sample from the second sampling space to obtain a second sampling vector.

[0092] Optionally, sampling is performed randomly from the second sampling space to obtain a second sampling vector. Optionally, sampling is performed according to a sampling rule from the second sampling space to obtain a second sampling vector.

[0093] Step 604: Decode the second sampling vector using the second decoding network in the molecular reconstruction model to obtain a reconstructed molecule.

[0094] Exemplarily, the screening molecule Y0 is input into the molecular reconstruction model, and the reconstructed molecule Y0′ is generated.

[0095] Optionally, the second decoding network is a decoder. Optionally, the decoder is at least one of a fully connected feedforward network, a transformer model, and an RNN model.

[0096] Step 605: By tuning the network, encoding and sampling are performed on the screening molecules to obtain a first sampling vector corresponding to the screening molecules.

[0097] Optionally, this step includes the following sub-steps:

[0098] 1. By reconstructing the network, the first encoding vector of the screening molecule is encoded according to the chemical structure of the screening molecule.

[0099] Exemplarily, by reconstructing the network, a first encoding vector of the screening molecule is encoded according to the chemical structure of the screening molecule and the SMILES specification.

[0100] 2. By reconstructing the network, a posterior distribution hypothesis is made on the first coding vector to generate a first sampling space corresponding to the first coding vector.

[0101] The posterior distribution hypothesis is used to assume that the latent features of the first encoding vector are normally distributed. In this case, the latent features are distributed in a first sampling space, which can be obtained by fitting the mean and variance of the first encoding vector.

[0102] Exemplarily, the first mean is obtained by fitting the mean of the latent features of the first coding vector through the first mean sub-network, and the first mean is used to represent the mean of the latent features of the first coding vector. Optionally, the second mean sub-network is at least one of a fully connected feedforward network, a transformer (a type of codec model), and an RNN model. It should be noted that the embodiment of the present application uses a machine learning model to fit the mean of the first coding vector, and the type of machine learning model can be determined by technicians according to actual needs.

[0103] Exemplarily, the variance of the first encoding vector is fitted through the first variance subnetwork to obtain the first variance, and the first mean is used to represent the variance of the latent features of the first encoding vector.

[0104] Exemplarily, a first normal distribution is determined based on a first mean and a first variance by generating a subnetwork through a first distribution.

[0105] 3. Sample from the first sampling space to obtain a second sampling vector.

[0106] Optionally, sampling is performed randomly from the first sampling space to obtain a first sampling vector. Optionally, sampling is performed according to a sampling rule from the first sampling space to obtain a first sampling vector.

[0107] Step 606: Decode the first sampling vector using the first decoding network in the molecule generation model to obtain a generated molecule.

[0108] Optionally, the first decoding network is a decoder. Optionally, the decoder is at least one of a fully connected feedforward network, a transformer model, and an RNN model.

[0109] Step 607: Train the molecule generation model based on the error between the generated molecule and the reconstructed molecule.

[0110] Optionally, the reconstructed molecule Y0′ and the screened molecule Y0 are combined to form an input-output pair When training a molecular generative model, the input of the molecular generative model is the input-output pair The optimization goal of the molecular generation model is the input-output pair Output data.

[0111] Optionally, the molecule generation model is trained based on the error between the generated molecules and the reconstructed molecules through an error back propagation algorithm.

[0112] Alternatively, the molecular generative model is trained based on the error between the generated molecules and the reconstructed molecules, as well as the error between the generated molecules and the screened molecules. In this case, since the generated molecules are also compared with the screened molecules, when training the molecular generative model, the generated molecules will not be too far away from the screened molecules.

[0113] In an optional implementation, when the training completion condition of the molecule generation model is met, the molecule generation model is used to sample from the first sampling space to obtain a target generation molecule of the molecule optimization condition.

[0114] Optionally, the screening molecule Y0 is input into the molecule generation model, and the generated molecule X1 is output.

[0115] Step 608: Iterate the above six steps.

[0116] Optionally, the above six steps are iterated starting from step 602 .

[0117] Step 609: When the training completion condition of the molecule generation model is met in the i-th iteration, the training of the molecule generation model is completed.

[0118] Optionally, when the model parameters of the molecule generation model converge in the i-th iteration, the training of the molecule generation model is completed.

[0119] Optionally, when K rounds of iteration are completed, the training of the molecular generation model is completed, where K is a preset positive integer.

[0120] In an optional implementation, when the training completion condition of the molecular generation model is met, the molecular generation model is used to sample from the first sampling space to obtain a target generation molecule of the molecular optimization condition, wherein the target generation molecule is a lead compound.

[0121] In summary, when training the molecule generation model, the embodiment of the present application will reconstruct the screening molecules, change the chemical structure of the screening molecules, obtain reconstructed molecules, and the reconstructed molecules meet the molecular optimization conditions. The screening molecules are input into the molecule generation model to obtain generated molecules, and the molecule generation model is trained by the error between the generated molecules and the reconstructed molecules. Since the chemical structure of the reconstructed molecules is different from that of the screening molecules, the reconstructed molecules ensure the data diversity in the training process, provide rich training labels for the molecule generation model, and the rich training labels provide a variety of training directions for optimizing properties, so the molecules generated by the molecule generation model can meet a variety of properties. Moreover, the generated molecules are obtained by sampling the molecule generation model, which can ensure that the number of generated molecules is large and the output of generated molecules is high.

[0122] At the same time, the embodiment of the present application can use the error between the generated molecule and the screened molecule as a reward to ensure that the generated molecule is not too far away from the original screened molecule.

[0123] In the following embodiments, taking the generation of molecules satisfying a properties as an example, a is a positive integer greater than 2. When the method provided in the embodiments of the present application is adopted, the generated molecules can satisfy a properties at the same time.

[0124] Figure 7 FIG1 shows a flow chart of a method for training a molecular generation model provided by an exemplary embodiment of the present application. Figure 1 The computer system 100 shown is executed, and the method includes:

[0125] Step 701: In the first iteration, candidate molecules are screened according to a type of attributes to obtain screened molecules, where the candidate molecules are molecules sampled from the second sampling space of the molecular reconstruction model.

[0126] Optionally, property a refers to an inherent characteristic of a molecule, and property a can be set by a technician based on actual needs. For example, property a can include the molecule having a biological activity greater than a preset activity threshold, or the molecule being able to bind to a target, or the molecule being non-toxic to organisms, or the molecule being water-soluble, or the molecule affecting cellular metabolism.

[0127] Optionally, the molecular reconstruction model is trained based on molecules in a public dataset. Exemplarily, the public dataset includes at least 6.8 million molecules.

[0128] Optionally, the molecular reconstruction model is at least one of a VAE model, a GVAE model, a MolGAN model, a MolFlow model, a GraphAF model, and a GraphNVP model. It should be noted that the embodiment of the present application does not specifically limit the type of the molecular reconstruction model.

[0129] For example, the candidate molecule is recorded as X0, and the molecular optimization condition is recorded as y0, then the screening molecule that meets the molecular optimization condition is Y0=y0⊙X0, where ⊙ represents an exclusive OR operation.

[0130] Step 702: In the i-th iteration, according to a types of attributes, the molecules generated in the previous round are screened to obtain screened molecules, where the molecules generated in the previous round are the molecules generated in the i-1-th iteration.

[0131] Optionally, the chemical structure of the screened molecule in the i-1th iteration is adjusted by the molecular generation model to obtain the generated molecule in the i-1th iteration.

[0132] During the iterative process, if the number of molecules obtained in this round of screening is too small, the number of screened molecules input into the molecule generation model will be too small, which is not conducive to the subsequent iterative training. Optionally, when the number of molecules screened in this round is less than the quantity threshold in the j-th round of iteration, the range of the molecular optimization conditions is adjusted, where j is an integer less than i; wherein, in the iterative process after the j-th round of iteration, the range of the molecular optimization conditions decreases as the number of iterations increases. Exemplarily, the a type of attribute is downgraded to the b type of attribute (b is less than a) to increase the number of screened molecules, and the number of optimized attributes is increased in the subsequent iterative process until it increases to the a type of attribute.

[0133] It should be noted that step 701 and step 702 are mutually exclusive. When step 701 is executed, step 702 is not executed; when step 702 is executed, step 701 is not executed.

[0134] Step 703: encoding and sampling the screening molecules through the reconstruction network in the molecular reconstruction model to obtain a second sampling vector corresponding to the screening molecules.

[0135] Optionally, this step includes the following sub-steps. The specific implementation method can refer to step 603 of the above embodiment and will not be repeated here:

[0136] 1. By reconstructing the network, the second encoding vector of the screening molecule is encoded according to the chemical structure of the screening molecule.

[0137] 2. By reconstructing the network, a posterior distribution hypothesis is made on the second encoding vector to generate a second sampling space corresponding to the second encoding vector.

[0138] 3. Sample from the second sampling space to obtain a second sampling vector.

[0139] Step 704: Decode the second sampling vector using the second decoding network in the molecular reconstruction model to obtain a reconstructed molecule.

[0140] Exemplarily, the screening molecule Y0 is input into the molecular reconstruction model, and the reconstructed molecule Y0′ is generated.

[0141] Optionally, the second decoding network is a decoder. Optionally, the decoder is at least one of a fully connected feedforward network, a transformer model, and an RNN model.

[0142] Step 705: By tuning the network, encoding and sampling are performed on the screening molecules to obtain a first sampling vector corresponding to the screening molecules.

[0143] Optionally, this step includes the following sub-steps, the specific implementation method of which can refer to step 605 of the above embodiment and will not be repeated here:

[0144] 1. By reconstructing the network, the first encoding vector of the screening molecule is encoded according to the chemical structure of the screening molecule.

[0145] 2. By reconstructing the network, a posterior distribution hypothesis is made on the first coding vector to generate a first sampling space corresponding to the first coding vector.

[0146] 3. Sample from the first sampling space to obtain a first sampling vector.

[0147] Step 706: Decode the first sampling vector using the first decoding network in the molecule generation model to obtain a generated molecule.

[0148] Optionally, the first decoding network is a decoder. Optionally, the decoder is at least one of a fully connected feedforward network, a transformer model, and an RNN model.

[0149] Step 707: Train the molecule generation model based on the error between the generated molecule and the reconstructed molecule.

[0150] Optionally, the reconstructed molecule Y0′ and the screened molecule Y0 are combined to form an input-output pair When training a molecular generative model, the input of the molecular generative model is the input-output pair The optimization goal of the molecular generation model is the input-output pair Output data.

[0151] Optionally, the molecular generative model is trained using an error back-propagation algorithm based on the error between the generated molecules and the reconstructed molecules. In one optional implementation, the generated molecules are converted into generated molecule feature vectors; a normal distribution of the generated molecule feature vectors is fitted to obtain a generated molecule normal distribution; and the error between the generated molecule and the reconstructed molecule is obtained based on the relative entropy between the generated molecule normal distribution and the normal distribution corresponding to the first sampling space.

[0152] Alternatively, the molecular generative model is trained based on the error between the generated molecules and the reconstructed molecules, as well as the error between the generated molecules and the screened molecules. In this case, since the generated molecules are also compared with the screened molecules, when training the molecular generative model, the generated molecules will not be too far away from the screened molecules.

[0153] In an optional implementation, when the training completion condition of the molecule generation model is met, the molecule generation model is used to sample from the first sampling space to obtain a target generation molecule of the molecule optimization condition.

[0154] Optionally, the screening molecule Y0 is input into the molecule generation model, and the generated molecule X1 is output.

[0155] Step 708: Iterate the above six steps.

[0156] Optionally, the above six steps are iterated starting from step 702 .

[0157] Step 709: When the training completion condition of the molecule generation model is met in the i-th iteration, the training of the molecule generation model is completed.

[0158] Optionally, when the model parameters of the molecule generation model converge in the i-th iteration, the training of the molecule generation model is completed.

[0159] Optionally, when K rounds of iteration are completed, the training of the molecular generation model is completed, where K is a preset positive integer.

[0160] Step 710: Sampling from the first sampling space through the molecule generation model to obtain target generated molecules that meet a properties.

[0161] Alternatively, the molecular generation model randomly samples from the first sampling space to obtain target generated molecules that satisfy a properties. Alternatively, the molecular generation model samples according to a preset rule in the first sampling space to obtain target generated molecules that satisfy a properties.

[0162] Among them, the target generated molecule belongs to the lead compound.

[0163] In summary, when training the molecular generation model, the embodiment of the present application reconstructs the screening molecules, changes the chemical structure of the screening molecules, obtains reconstructed molecules, and the reconstructed molecules meet the molecular optimization conditions. The screening molecules are input into the molecular generation model to obtain generated molecules, and the molecular generation model is trained by the error between the generated molecules and the reconstructed molecules. Since the chemical structure of the reconstructed molecules is different from that of the screening molecules, the reconstructed molecules ensure the data diversity in the training process, provide rich training labels for the molecular generation model, and the rich training labels provide a variety of training directions for optimizing properties, so that the molecules generated by the molecular generation model can meet a variety of properties.

[0164] In an alternative embodiment, please refer to Figure 2 :

[0165] The molecular reconstruction model 202 first needs to go through a pre-training process. The data used for pre-training are all molecules from a public data set, and the molecules in the public data set are encoded using SMILES. The molecular reconstruction model 202 here uses at least one of the VAE model, GVAE model, MolGAN model, MolFlow model, GraphAF model, and GraphNVP model. After pre-training the molecular reconstruction model 202 on more than 6.8 million molecules, a trained molecular reconstruction model 202 can be obtained. The molecular reconstruction model 202 will not be trained again in subsequent operations. The molecular reconstruction model 202 can perform the following two molecular generation operations:

[0166] (1) Sampling generation: The molecular reconstruction model 202 does not need to input any reference molecules during the generation process, and it directly samples the generated candidate molecules X0 from the molecular space.

[0167] (2) Reconstruction Generation: The molecular reconstruction model 202 needs to input the reference screening molecule X during the generation process. Based on the screening molecule X, the molecular reconstruction model 202 samples a reconstructed molecule X′ that is close to X.

[0168] Molecular generation model 205 shares the same architecture as molecular reconstruction model 202, but further trains molecular generation model 202 based on each reinforcement learning reward. The initial parameters of molecular generation model 205 are obtained from molecular reconstruction model 202. Molecular generation model 205 also utilizes two of the same molecular generation operations as molecular reconstruction model 202: sampling generation and reconstruction generation. These operations will not be further detailed here; please refer to the description of molecular reconstruction model 202 above.

[0169] The screening process can be implemented using a screening model, which is a property or activity judgment model. Its input is the screening molecule X. The output is a judgment y indicating whether each molecule meets the optimization criteria. This model can be an artificial intelligence model, a computational chemistry model, or even experimental results. As long as the model can determine whether a given molecule meets the optimization criteria, it is sufficient.

[0170] Based on the above description, training the molecule generation model 205 may include the following steps:

[0171] 1. Molecular reconstruction model sampling and generation: At the beginning of the process, the first batch of candidate molecules are directly sampled and generated by the molecular reconstruction model 201, which is X0.

[0172] 2. Molecular deduplication and legality verification: remove duplicate molecules used to generate candidate molecules 202, as well as molecules that do not conform to chemical rules or disconnected molecular graphs.

[0173] 3. Molecular screening: Using multiple property or activity filters, select the screening molecules 203 that meet the optimization conditions: Y0 = y0 ⊙ X0 from the candidate molecules 202. The screening methods are as follows:

[0174] (a) The larger the optimization value, the better: directly screen out the 50% molecules with the largest optimization value;

[0175] (b) Optimization value within a certain range: Screening molecules by limiting the range;

[0176] (c) After filtering based on the conditions, very few molecules remain (e.g., <100 molecules): Lower the filtering conditions to allow enough molecules to be selected. Then, raise the filtering conditions in a subsequent reinforcement learning optimization cycle.

[0177] 4. Input the screening molecule 203 into the molecular reconstruction model 201 for reconstruction and generation, and obtain the reconstructed molecule 204, Y0'. Optionally, the screening result generated from the first sampling of the molecular reconstruction model 201 does not need this step.

[0178] 5. Combine the screening molecule 203 and the reconstructed molecule 204 to form Input-output pairs.

[0179] 6. Train and optimize the molecular generation model 205, whose input is The optimization goal is Output data.

[0180] 7. Use the trained and optimized molecular generation model 205 to perform sampling and generation to obtain a new batch of generated molecules X1, and use the new batch of generated molecules X1 as candidate molecules for the next iteration.

[0181] 8. Repeat steps 2-7 4-5 times, and finally obtain a molecular generation model 205 that can generate molecules that meet the optimization conditions.

[0182] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0183] Please refer to Figure 8 , which shows a block diagram of a training device for a molecular generation model provided by an embodiment of the present application. The above functions can be implemented by hardware or by hardware executing corresponding software. The device 800 includes:

[0184] A screening module 801 is configured to screen candidate molecules according to molecular optimization conditions to obtain screening molecules, wherein the molecular optimization conditions are used to evaluate the properties of the candidate molecules;

[0185] a reconstruction module 802 for reconstructing the screening molecule to generate a reconstructed molecule, wherein the reconstructed molecule satisfies the molecule optimization condition and the chemical structure of the reconstructed molecule is different from the chemical structure of the screening molecule;

[0186] An optimization module 803 is configured to adjust the chemical structure of the screened molecule using the molecular generation model to obtain a generated molecule;

[0187] The training module 804 is configured to train the molecule generation model according to the error between the generated molecule and the reconstructed molecule.

[0188] In an optional design of the present application, the molecular generation model includes a tuning network and a first decoding network; the optimization module 803 is further used to encode and sample the screening molecules through the tuning network to obtain a first sampling vector corresponding to the screening molecules; and to decode the first sampling vector through the first decoding network in the molecular generation model to obtain the generated molecule.

[0189] In an optional design of the present application, the optimization module 803 is also used to encode the chemical structure of the screening molecule through the tuning network to obtain a first encoding vector of the screening molecule; perform a posterior distribution hypothesis on the first encoding vector through the tuning network to generate a first sampling space corresponding to the first encoding vector; and sample from the first sampling space to obtain the first sampling vector.

[0190] In an optional design of the present application, the features in the first sampling space are normally distributed; the tuning network includes a first mean subnetwork, a first variance subnetwork and a first distribution generation subnetwork; the optimization module 803 is also used to fit the mean of the latent features of the first encoding vector through the first mean subnetwork to obtain a first mean; fit the variance of the latent features of the first encoding vector through the first variance subnetwork to obtain a first variance; and determine the first sampling space with a normal distribution according to the first mean and the first variance through the first distribution generation subnetwork.

[0191] In an optional design of the present application, the reconstruction module 802 is further used to encode and sample the screening molecule through the reconstruction network in the molecular reconstruction model to obtain a second sampling vector corresponding to the screening molecule; and to decode the second sampling vector through the second decoding network in the molecular reconstruction model to obtain the reconstructed molecule.

[0192] In an optional design of the present application, the reconstruction module 802 is also used to generate a second encoding vector of the screening molecule according to the chemical structure of the screening molecule through the reconstruction network; perform a posterior distribution hypothesis on the second encoding vector through the reconstruction network to generate a second sampling space corresponding to the second encoding vector; and sample from the second sampling space to obtain the second sampling vector.

[0193] In an optional design of the present application, the features in the second sampling space are normally distributed; the reconstruction network includes a second mean subnetwork, a second variance subnetwork and a second distribution generation subnetwork; the reconstruction module 802 is also used to fit the mean of the latent features of the second encoding vector through the second mean subnetwork to obtain a second mean; fit the variance of the latent features of the second encoding vector through the second variance subnetwork to obtain a second variance; and determine the second normal distribution that is normally distributed according to the second mean and the second variance through the second distribution generation subnetwork.

[0194] In an optional design of the present application, the screening module 801 is further used to remove duplicate molecules from the candidate molecules; or, to remove molecules from the candidate molecules that do not conform to chemical rules; or, to remove molecules from the candidate molecules that are not connected in the molecular graph.

[0195] In an optional design of the present application, the reconstruction module 802 is further used to calculate the optimization value of the candidate molecule; use the candidate molecules corresponding to the N maximum optimization values as the screening molecules, where N is a positive integer; or calculate the optimization value of the candidate molecule; and determine the candidate molecules corresponding to the optimization values in a preset range as the screening molecules.

[0196] In an optional design of the present application, in the first round of iteration, the candidate molecule is a molecule sampled from the second sampling space of the molecular reconstruction model; in the i-th round of iteration, the candidate molecule is the generated molecule in the i-1-th round of iteration, where i is an integer greater than 1; the training module 804 is also used to complete the training of the molecular generation model when the training completion conditions of the molecular generation model are met in the i-th round of iteration.

[0197] In an optional design of the present application, the training module 804 is further used to adjust the range of the molecular optimization conditions when the number of molecules screened in the current round is less than a quantity threshold in the j-th iteration, where j is an integer less than i; wherein, in the iterative process after the j-th iteration, the range of the molecular optimization conditions shrinks as the number of iterations increases.

[0198] In an optional design of the present application, the training module 804 is further used to train the molecular generation model based on the error between the generated molecule and the reconstructed molecule, and the error between the generated molecule and the screened molecule.

[0199] In an optional design of the present application, the optimization module 803 is further used to sample from the first sampling space through the molecular generation model to obtain the target generation molecule of the molecular optimization condition when the training completion condition of the molecular generation model is met.

[0200] In summary, when training the molecular generation model, the embodiment of the present application reconstructs the screening molecules, changes the chemical structure of the screening molecules, obtains reconstructed molecules, and the reconstructed molecules meet the molecular optimization conditions. The screening molecules are input into the molecular generation model to obtain generated molecules, and the molecular generation model is trained by the error between the generated molecules and the reconstructed molecules. Since the chemical structure of the reconstructed molecules is different from that of the screening molecules, the reconstructed molecules ensure the data diversity in the training process, provide rich training labels for the molecular generation model, and the rich training labels provide a variety of training directions for optimizing properties, so that the molecules generated by the molecular generation model can meet a variety of properties.

[0201] Figure 9 1 is a schematic diagram illustrating the structure of a computer device according to an exemplary embodiment. The computer device 900 includes a central processing unit (CPU) 901, a system memory 904 including a random access memory (RAM) 902 and a read-only memory (ROM) 903, and a system bus 905 connecting the system memory 904 and the CPU 901. The computer device 900 also includes a basic input / output system (I / O system) 906 for facilitating information transmission between various components within the computer device, and a mass storage device 907 for storing an operating system 913, application programs 914, and other program modules 915.

[0202] The basic input / output system 906 includes a display 908 for displaying information and an input device 909 such as a mouse and a keyboard for user input. The display 908 and the input device 909 are both connected to the central processing unit 901 via an input / output controller 910 connected to the system bus 905. The basic input / output system 906 may also include an input / output controller 910 for receiving and processing input from a variety of other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 910 also provides output to a display screen, a printer, or other types of output devices.

[0203] The mass storage device 907 is connected to the central processing unit 901 via a mass storage controller (not shown) connected to the system bus 905. The mass storage device 907 and its associated computer-readable medium provide non-volatile storage for the computer device 900. In other words, the mass storage device 907 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM drive.

[0204] Without loss of generality, the computer device readable medium may include computer device storage media and communication media. Computer device storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer device readable instructions, data structures, program modules or other data. Computer device storage media include RAM, ROM, Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), CD-ROM, Digital Video Disc (DVD) or other optical storage, tape cassettes, magnetic tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer device storage media are not limited to the above-mentioned ones. The above-mentioned system memory 904 and mass storage device 907 can be collectively referred to as memory.

[0205] According to various embodiments of the present application, the computer device 900 may also be connected to a remote computer device on a network such as the Internet for operation. That is, the computer device 900 may be connected to the network 911 via the network interface unit 912 connected to the system bus 905, or the network interface unit 912 may be used to connect to other types of networks or remote computer device systems (not shown).

[0206] The memory further includes one or more programs, which are stored in the memory. The central processing unit 901 implements all or part of the steps of the above-mentioned training method for molecular generation model by executing the one or more programs.

[0207] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one instruction, at least one program, code set or instruction set is stored. The at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the training method of the molecular generation model provided by each of the above-mentioned method embodiments.

[0208] The present application also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the training method of the molecular generation model provided in the above method embodiment.

[0209] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the method for training a molecular generative model as provided in the above embodiments.

[0210] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0211] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0212] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for training a molecular generative model, characterized in that: The method comprises: screening candidate molecules according to molecular optimization conditions to obtain screening molecules, wherein the molecular optimization conditions are used to evaluate the properties of the candidate molecules; Reconstructing the screening molecule to generate a reconstructed molecule, wherein the reconstructed molecule satisfies the molecular optimization condition and the chemical structure of the reconstructed molecule is different from the chemical structure of the screening molecule; wherein the molecular reconstruction is used to adjust other regions in the chemical structure of the screening molecule while retaining the region in the screening molecule that satisfies the molecular optimization condition; Adjusting the chemical structure of the screening molecule using the molecular generation model to obtain a generated molecule; The molecule generation model is trained according to the error between the generated molecule and the reconstructed molecule.

2. The method according to claim 1, characterized in that The molecular generation model includes a tuning network and a first decoding network; The step of adjusting the chemical structure of the screening molecule by using the molecular generation model to obtain the generated molecule comprises: Performing encoding and sampling processing on the screening molecule through the tuning network to obtain a first sampling vector corresponding to the screening molecule; The first sampling vector is decoded by a first decoding network in the molecule generation model to obtain the generated molecule.

3. The method according to claim 2, characterized in that The step of performing encoding and sampling processing on the screening molecule through the tuning network to obtain a first sampling vector corresponding to the screening molecule includes: Encoding the first encoding vector of the screening molecule according to the chemical structure of the screening molecule through the tuning network; Through the tuning network, a posterior distribution hypothesis is performed on the first encoding vector to generate a first sampling space corresponding to the first encoding vector; and sampling is performed from the first sampling space to obtain the first sampling vector.

4. The method according to claim 3, characterized in that The features in the first sampling space are normally distributed; the tuning network includes a first mean subnetwork, a first variance subnetwork and a first distribution generation subnetwork; The step of performing a posterior distribution hypothesis on the first encoding vector through the tuning network to generate a first sampling space corresponding to the first encoding vector includes: Fitting the mean of the latent features of the first encoding vector through the first mean sub-network to obtain a first mean; Fitting the variance of the latent features of the first encoding vector through the first variance sub-network to obtain a first variance; The first sampling space having a normal distribution is determined according to the first mean and the first variance through a first distribution generation subnetwork.

5. The method according to any one of claims 1 to 4, characterized in that The step of reconstructing the screening molecule to generate a reconstructed molecule comprises: Performing encoding and sampling processing on the screening molecule through a reconstruction network in a molecular reconstruction model to obtain a second sampling vector corresponding to the screening molecule; The second sampling vector is decoded by a second decoding network in the molecular reconstruction model to obtain the reconstructed molecule.

6. The method according to claim 5, characterized in that The encoding and sampling processing is performed on the screening molecule through the reconstruction network in the molecular reconstruction model to obtain a second sampling vector corresponding to the screening molecule, including: Encoding the second encoding vector of the screening molecule according to the chemical structure of the screening molecule through the reconstructed network; Performing a posterior distribution hypothesis on the second encoding vector through the reconstruction network to generate a second sampling space corresponding to the second encoding vector; Sampling is performed from the second sampling space to obtain the second sampling vector.

7. The method according to claim 6, characterized in that The features in the second sampling space are normally distributed; the reconstruction network includes a second mean subnetwork, a second variance subnetwork and a second distribution generation subnetwork; The step of performing a posterior distribution hypothesis on the second encoding vector through the reconstruction network to generate a second sampling space corresponding to the second encoding vector includes: Fitting the mean of the latent features of the second encoding vector through the second mean sub-network to obtain a second mean; Fitting the variance of the latent features of the second encoding vector through the second variance sub-network to obtain a second variance; A second normal distribution is determined according to the second mean and the second variance by using the second distribution generation subnetwork.

8. The method according to any one of claims 1 to 4, characterized in that The method further comprises: removing duplicate molecules from the candidate molecules; Or, removing molecules that do not conform to chemical rules from the candidate molecules; Alternatively, the molecules whose molecular graphs are not connected are removed from the candidate molecules.

9. The method according to any one of claims 1 to 4, characterized in that The step of screening candidate molecules to obtain screening molecules according to the molecular optimization conditions comprises: Calculating the optimization value of the candidate molecule; using the candidate molecules corresponding to N maximum optimization values as the screening molecules, where N is a positive integer; Alternatively, the optimization value of the candidate molecule is calculated; and the candidate molecule corresponding to the optimization value in a preset range is determined as the screening molecule.

10. The method according to any one of claims 1 to 4, characterized in that In the first iteration, the candidate molecule is a molecule sampled from the second sampling space of the molecular reconstruction model; in the i-th iteration, the candidate molecule is the generated molecule in the i-1-th iteration, where i is an integer greater than 1; The method further comprises: When the training completion condition of the molecule generation model is met in the i-th iteration, the training of the molecule generation model is completed.

11. The method according to claim 10, characterized in that The method further comprises: If the number of molecules screened in this round is less than the quantity threshold in the j-th iteration, the range of the molecule optimization condition is adjusted, where j is an integer less than i; In the iterative process after the j-th iteration, the range of the molecular optimization condition is reduced as the number of iterations increases.

12. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The molecular generation model is trained based on the error between the generated molecules and the reconstructed molecules, and the error between the generated molecules and the screened molecules.

13. The method according to any one of claims 1 to 4, characterized in that The method further comprises: When the training completion condition of the molecule generation model is met, sampling is performed from the first sampling space by the molecule generation model to obtain a target generated molecule that meets the molecule optimization condition.

14. A training device for a molecular generation model, characterized in that: The device comprises: a screening module, configured to screen candidate molecules according to molecular optimization conditions to obtain screening molecules, wherein the molecular optimization conditions are used to evaluate the properties of the candidate molecules; a reconstruction module, configured to reconstruct the screening molecule to generate a reconstructed molecule, wherein the reconstructed molecule satisfies the molecular optimization conditions and the chemical structure of the reconstructed molecule is different from the chemical structure of the screening molecule; wherein the molecular reconstruction is configured to adjust other regions in the chemical structure of the screening molecule while retaining the regions in the screening molecule that satisfy the molecular optimization conditions; an optimization module, configured to adjust the chemical structure of the screened molecule using the molecular generation model to obtain a generated molecule; A training module is used to train the molecule generation model according to the error between the generated molecule and the reconstructed molecule.

15. A computer device, characterized in that: The computer device includes: a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the training method for the molecular generation model according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one program code, and the program code is loaded and executed by a processor to implement the training method for a molecular generation model according to any one of claims 1 to 13.

17. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instructions are executed by a processor, the method for training a molecular generation model according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Transfer learning for molecular structure generation

    US20210374551A1

  • Drug molecule screening method and system

    WO2022047677A1