Training method, device, equipment and storage medium for molecular generation model
By training a molecular generation model, the expression information of linker molecules is generated using the expression information of sample molecules. This solves the problem of the time-consuming process of manually designing linker molecules, achieves efficient generation of macrocyclic small molecules, and reduces drug development costs.
Patent Information
- Application Number
- CN202210039386.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-13
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-01-13
AI Technical Summary
Existing technologies for generating macrocyclic small molecules are inefficient, and the cost and time-consuming design of artificial linker molecules are high, resulting in low efficiency in drug development.
By acquiring the expression information of the first sample molecule and the expression information of the second sample molecule, inputting them into the molecule generation model, training the model to generate the expression information of the linker molecule, and using the predicted expression information and the expression information of the second sample molecule to train the model, it is possible to achieve linker molecule design without manual design.
This significantly improves the efficiency of generating macrocyclic small molecules, reduces drug development costs, and enhances drug development efficiency.
Smart Images

Figure CN114373521B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method for a molecular generation model, a method for determining information of macrocyclic small molecules, an apparatus, a device, and a storage medium. Background Technology
[0002] Currently, in the development of small molecule drug compounds, the process often involves screening existing small molecule drug compounds to identify those that can interact with the corresponding disease target, followed by relevant clinical trials. However, existing small molecule drug compounds often have chain structures, resulting in lower pharmacological activity and selectivity. Therefore, it is necessary to modify the structure of small molecules to create macrocyclic small molecules, thereby improving the efficacy of the developed drugs.
[0003] In related technologies, medicinal chemists will artificially design linker molecules that can bind to small molecules to form macrocyclic structures based on the characteristics of small molecules and related medicinal chemistry knowledge. By binding small molecules to these linker molecules, macrocyclic small molecules are generated.
[0004] The aforementioned technologies require the manual design of linker molecules, which is costly and time-consuming, resulting in low efficiency in generating macrocyclic small molecules. Summary of the Invention
[0005] This application provides a method for training a molecular generation model, a method for determining information about macrocyclic small molecules, an apparatus, a device, and a storage medium. This method can improve the efficiency of generating macrocyclic small molecules. The technical solution is as follows:
[0006] On the one hand, a training method for a molecular generation model is provided, which includes:
[0007] The expression information of the first sample molecule and the expression information of the second sample molecule are obtained. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule. The second sample molecule is a replaceable molecular fragment in the molecule.
[0008] The expression information of the first sample molecule and the expression information of the second sample molecule are input into the molecule generation model to obtain the expression information of the predicted molecule.
[0009] The molecule generation model is trained based on the expression information of the predicted molecule and the expression information of the second sample molecule.
[0010] On the one hand, a method for determining information about macrocyclic small molecules is provided, the method comprising:
[0011] Obtain the expression information of the first molecule;
[0012] The expression information of the first molecule is input into the molecular generation model to obtain the expression information of the second molecule, which is a small molecule that can bind to the first molecule.
[0013] Based on the expression information of the first molecule and the expression information of the second molecule, the expression information of the macrocyclic small molecule is obtained;
[0014] The molecular generation model is trained based on the expression information of the first sample molecule and the expression information of the second sample molecule. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment in the molecule.
[0015] On the one hand, a training device for a molecular generation model is provided, the device comprising:
[0016] The acquisition module is used to acquire the expression information of the first sample molecule and the second sample molecule, wherein the first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment in the molecule.
[0017] The generation module is used to input the expression information of the first sample molecule and the expression information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule.
[0018] The training module is used to train the molecule generation model based on the expression information of the predicted molecule and the expression information of the second sample molecule.
[0019] In some embodiments, the generation module includes:
[0020] The acquisition submodule is used to input the expression information of the first sample molecule into the first embedding vector sub-model of the molecule generation model to obtain the embedding vector of the first sample molecule.
[0021] The acquisition submodule is used to input the embedding vector of the first sample molecule into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first sample molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first sample molecule.
[0022] The acquisition submodule is used to input the expression information of the second sample molecule into the second embedding sub-model of the molecule generation model to obtain the embedding vector of the second sample molecule.
[0023] The generation submodule is used to input the molecular binding information of the first sample molecule and the embedding vector of the second sample molecule into the decoding submodule of the molecule generation model to obtain the expression information of the predicted molecule.
[0024] In some embodiments, the generation submodule is used to input the embedding vector of the first sample molecule into the encoding submodel; based on the embedding vector of the first sample molecule and the multi-head self-attention unit in the encoding submodel, to extract the correlation information between each atom in the first sample molecule and the correlation information between each atom and each chemical bond in the first sample molecule, so as to obtain the initial molecular binding information of the first sample molecule.
[0025] Based on the feedforward neural network unit of this coding sub-model, multiple correlation information in the initial molecular binding information of the first sample molecule is nonlinearly fused to obtain the molecular binding information of the first sample molecule.
[0026] In some embodiments, the generation module is used to input the expression information of the first sample molecule, the expression information of the second sample molecule, and the length and structure information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule, wherein the structure information is used to indicate whether the second sample molecule includes a ring structure.
[0027] In some embodiments, the training module is configured to obtain a loss value based on the expression information of the predicted molecule and the expression information of the second sample molecule, the loss value representing the error between the expression information of the predicted molecule and the expression information of the second sample molecule; if the training process does not reach the training termination condition, the network parameters of the molecule generation model are adjusted based on the loss value; if the training process reaches the training termination condition, the molecule generation model is output.
[0028] In some embodiments, the device further includes a segmentation module for determining replaceable molecular fragments in the molecule based on a matching molecule pair algorithm and the expression information of the molecule; and segmenting the molecule based on the position of the molecular fragment in the molecule to obtain the expression information of the first sample molecule and the expression information of the second sample molecule, wherein the expression information of the first sample molecule and the expression information of the second sample molecule include the segmentation position of the molecule.
[0029] On the one hand, an information determination device for macrocyclic small molecules is provided, the device comprising:
[0030] The acquisition module is used to acquire the expression information of the first molecule;
[0031] The generation module is used to input the expression information of the first molecule into the molecular generation model to obtain the expression information of the second molecule, which is a small molecule that can bind to the first molecule.
[0032] The acquisition module is used to acquire the expression information of macrocyclic small molecules based on the expression information of the first molecule and the expression information of the second molecule;
[0033] The molecular generation model is trained based on the expression information of the first sample molecule and the expression information of the second sample molecule. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment in the molecule.
[0034] In some embodiments, the generation module includes:
[0035] The acquisition submodule is used to input the expression information of the first molecule into the first embedding sub-model of the molecule generation model to obtain the embedding vector of the first molecule;
[0036] The acquisition submodule is used to input the embedding vector of the first molecule into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first molecule.
[0037] The generation submodule is used to input the molecular binding information of the first molecule into the decoding submodule of the molecule generation model to obtain the expression information of the second molecule.
[0038] In some embodiments, the generation submodule is used to input the embedding vector of the first molecule into the encoding submodel; based on the embedding vector of the first molecule and the multi-head self-attention unit of the encoding submodel, to extract the correlation information between each atom in the first molecule and the correlation information between each atom in the first molecule and each chemical bond, so as to obtain the initial molecular binding information of the first molecule.
[0039] Based on the feedforward neural network unit of this coding sub-model, multiple correlation information in the initial molecular binding information of the first molecule is nonlinearly fused to obtain the molecular binding information of the first molecule.
[0040] In some embodiments, the generation module is used to input molecular length, molecular structure information and expression information of the first molecule into the molecular generation model to obtain expression information of the second molecule. The molecular length is used to represent the length of the small molecule bound to the first molecule, and the molecular structure information is used to represent whether the small molecule bound to the first molecule includes a ring structure.
[0041] In some embodiments, the expression information of the first molecule includes a first linking site and a second linking site of the first molecule, the first linking site and the second linking site being atoms in the first molecule used to bind with other molecules; the expression information of the second molecule includes a third linking site and a fourth linking site of the second molecule, the third linking site and the fourth linking site being atoms in the second molecule that bind with the first linking site and the second linking site, respectively.
[0042] The acquisition module is used to determine the binding mode between the first molecule and the second molecule based on the first and second linking sites of the first molecule, and the third and fourth linking sites of the second molecule; to determine the chemical structure of the macrocyclic small molecule based on the binding mode; and to acquire the expression information of the macrocyclic small molecule based on the chemical structure of the macrocyclic small molecule.
[0043] On one hand, a computer device is provided, comprising one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, the at least one computer program being loaded and executed by the one or more processors to implement a training method for the molecular generation model or a method for determining information of macrocyclic small molecules.
[0044] On the one hand, a computer-readable storage medium is provided, which stores at least one computer program, which is loaded and executed by a processor to implement a training method for the molecular generation model or a method for determining information of macrocyclic small molecules.
[0045] On one hand, a computer program product is provided, comprising at least one computer program stored in a computer-readable storage medium. A processor of a computer device reads the at least one computer program from the computer-readable storage medium and executes the at least one computer program, causing the computer device to implement a training method for the molecular generation model or a method for determining information of macrocyclic small molecules.
[0046] The technical solution provided in this application involves inputting the expression information of the first sample molecule and the expression information of the second sample molecule obtained from existing molecular segmentation into a molecular generation model. This yields the expression information of the linker molecule corresponding to the first sample molecule predicted by the model. Furthermore, the model is trained using the predicted expression information and the expression information of the second sample molecule, enabling the model to accurately generate the expression information of the linker molecule. This eliminates the need for drug researchers to manually design linker molecules, significantly improving the efficiency of generating macrocyclic small molecules. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a drug development flowchart provided in an embodiment of this application;
[0049] Figure 2 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0050] Figure 3 This is a flowchart of a training method for a molecular generation model provided in an embodiment of this application;
[0051] Figure 4 This is a flowchart of a method for determining information about macrocyclic small molecules provided in an embodiment of this application;
[0052] Figure 5 This is a flowchart of a training method for a molecular generation model provided in an embodiment of this application;
[0053] Figure 6 This is a schematic diagram of a molecular generation model provided in an embodiment of this application;
[0054] Figure 7 This is a flowchart of a method for determining information about macrocyclic small molecules provided in an embodiment of this application;
[0055] Figure 8 This is a schematic diagram of the structure of a training device for a molecular generation model provided in an embodiment of this application;
[0056] Figure 9 This is a schematic diagram of the structure of an information determination device for macrocyclic small molecules provided in an embodiment of this application;
[0057] Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0059] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.
[0060] To facilitate understanding of the technical processes involved in the embodiments of this application, some terms used in the embodiments of this application are explained below:
[0061] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0062] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0063] The technical solutions provided in this application can also be combined with cloud technology, for example, by deploying the trained molecular generation model on a cloud server. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize the computation, storage, processing, and sharing of data.
[0064] Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and cryptographic algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying platform, a platform product service layer, and an application service layer.
[0065] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learn-by-doing.
[0066] The Simplified Molecular Input Line Entry Specification (SMILES) is a specification that uses ASCII strings to represent the chemical structure of molecules.
[0067] The following section introduces the training method for the molecular generation model proposed in this application and the application scenarios of the method for determining information of macrocyclic small molecules.
[0068] The method provided in this application can be used for the development of small molecule drugs, such as... Figure 1 As shown, small molecule drug development mainly includes four processes: target identification and validation, compound screening and lead discovery, preclinical study, and clinical trials. After target identification and validation, drug researchers screen for small molecules that can interact with the target from a variety of existing small molecule drug compounds. However, the screened small molecules are often chain-like structures, which may have low pharmacological activity and selectivity. Based on the macrocyclic small molecule information determination method proposed in this application, the expression information of the linker molecules corresponding to the chain-like molecules can be obtained through a molecular generation model. Furthermore, based on the expression information of the chain-like molecules and the corresponding linker molecules, the expression information of the macrocyclic small molecules can be obtained. The molecular generation model used in this method can be obtained through the training method of the molecular generation model provided in this application. Therefore, the method provided in this application can assist drug researchers in transforming screened chain molecules into macrocyclic small molecules, which greatly improves research and development efficiency and reduces drug development costs.
[0069] The following describes the implementation environment of this application. Figure 2 This is a schematic diagram of an implementation environment provided in an embodiment of this application. This implementation environment can be used for training methods of molecular generation models or methods for determining information of macrocyclic small molecules, such as... Figure 2 As shown, the implementation environment includes a terminal 201 and a server 202, which are interconnected via wired or wireless networks.
[0070] Terminal 201 may be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. In some embodiments, terminal 201 is used to provide server 202 with relevant data required for training the molecular generative model, such as initializing network parameters, model training hyperparameters, etc.
[0071] Server 202 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In some embodiments, server 202 is used to execute the method for determining information of macrocyclic small molecules provided in the embodiments of this application, and / or to train a molecular generation model based on data provided by terminal 201.
[0072] Optionally, in the process of training the molecular generation model or determining the expression information of macrocyclic small molecules, server 202 undertakes the main computational work and terminal 201 undertakes the secondary computational work; or, server 202 undertakes the secondary computational work and terminal 201 undertakes the main computational work; or, server 202 or terminal 201 can each undertake the computational work independently.
[0073] In some embodiments, the terminal 201 and server 202 described above can function as nodes in a blockchain system.
[0074] To more clearly illustrate the method provided in this application, a brief introduction to the molecular generation model involved in this application is given below. The structure of this molecular generation model is based on the Seq2Seq algorithm for text generation in Natural Language Processing (NLP). The molecular generation model includes a first embedding vector sub-model, an encoding sub-model, a second embedding vector sub-model, and a decoding sub-model. The first embedding vector sub-model is used to transform the molecular expression information required by the encoding sub-model into vectors for processing by the encoding sub-model to obtain encoded information. The second embedding vector sub-model is used to transform the molecular expression information required by the decoding sub-model into vectors, enabling the decoding sub-model to generate the expression information of the linker molecule based on the vectors and the encoded information. Optionally, this molecular generation model includes an encoding sub-model, a decoding sub-model, and an embedding vector sub-model. This embedding vector sub-model can transform the molecular expression information required by the encoding and decoding sub-models into vectors, enabling the encoding and decoding sub-models to generate the expression information of the linker molecule based on the corresponding vectors. Optionally, this molecular generation model is also called LinkerTransformer.
[0075] based on Figure 2 The implementation environment shown and the molecular generation model described above, Figure 3 This is a flowchart illustrating a training method for a molecular generation model provided in an embodiment of this application. This method is executed by a server, such as... Figure 3 As shown, this embodiment includes the following steps.
[0076] 301. The server obtains the expression information of the first sample molecule and the expression information of the second sample molecule. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule. The second sample molecule is a replaceable molecular fragment in the molecule.
[0077] The expression information is used to represent the chemical structure of the corresponding sample molecule. Optionally, this expression information is a string conforming to the Simplified Molecular Linear Input Specification (SMILES).
[0078] The molecules used to obtain the first and second sample molecules are known small molecule drug compounds. Optionally, the molecule can be a chain-like small molecule or a macrocyclic small molecule; this embodiment of the application does not limit this. By segmenting the molecules, sample molecule pairs for training the model are obtained, and these sample molecule pairs include the expression information of the first and second sample molecules. Optionally, the second sample molecule is also referred to as the linker molecule of the first sample molecule.
[0079] In some embodiments, the expression information of the first sample molecule and the expression information of the second sample molecule also include the molecular segmentation position. For example, the expression information of the first sample molecule is [*:1]C1=CC(C2=CC=NC(NC3=CC([*:2])=C3)=N2)=CC=C1, and the expression information of the second sample molecule is CN(CCCCO[*:1])C[*:2], where [*:1] and [*:2] represent the corresponding linking sites in the first and second sample molecules. These linking sites refer to atoms in the molecule that can bind to other molecules, and the corresponding linking sites in the first and second sample molecules are also the molecular segmentation positions.
[0080] 302. The server inputs the expression information of the first sample molecule and the expression information of the second sample molecule into the molecular generation model to obtain the expression information of the predicted molecule.
[0081] The molecule generation model can obtain the expression information of the linker molecule, and the predicted molecule is the linker molecule of the first sample molecule predicted by the model.
[0082] 303. The server trains the molecule generation model based on the expression information of the predicted molecule and the expression information of the second sample molecule.
[0083] In some embodiments, the server trains the molecule generation model by predicting the error between the expression information of the molecule and the expression information of the second sample molecule.
[0084] The technical solution provided in this application involves inputting the expression information of the first sample molecule and the expression information of the second sample molecule obtained from existing molecular segmentation into a molecular generation model. This yields the expression information of the linker molecule predicted by the model for the first sample molecule. Furthermore, the model is trained using the predicted expression information and the expression information of the second sample molecule, enabling the model to accurately generate the expression information of the linker molecule. This eliminates the need for drug researchers to manually design linker molecules, significantly improving the efficiency of generating macrocyclic small molecules.
[0085] based on Figure 2 The implementation environment shown and the molecular generation model described above, Figure 4 This is a flowchart of a method for determining information about macrocyclic small molecules provided in an embodiment of this application. This method is executed by a server, such as... Figure 4 As shown, the method includes the following steps.
[0086] 401. The server obtains the expression information of the first molecule.
[0087] The first molecule is an existing small molecule drug compound that requires structural modification, and the expression information of the first molecule is used to represent the chemical structure of the first molecule.
[0088] In some embodiments, the expression information of the first molecule is also used to represent the first linking site and the second linking site of the first molecule, which are atoms in the first molecule that are bonded to other molecules.
[0089] 402. The server inputs the expression information of the first molecule into the molecular generation model to obtain the expression information of the second molecule, which is a small molecule that can bind to the first molecule.
[0090] The molecular generation model is trained based on the expression information of the first sample molecule and the expression information of the second sample molecule. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment in the molecule.
[0091] The second molecule can form a macrocyclic small molecule by binding with the first molecule, and the expression information of the second molecule is used to represent the chemical structure of the second molecule. Optionally, the second molecule is also referred to as a linker molecule of the first molecule.
[0092] In some embodiments, the expression information of the second molecule further includes a third linking site and a fourth linking site in the second molecule, wherein the third linking site and the fourth linking site are atoms in the second molecule that are bonded to the first linking site and the second linking site, respectively.
[0093] 403. Based on the expression information of the first molecule and the expression information of the second molecule, the server obtains the expression information of the macrocyclic small molecule.
[0094] The technical solution provided in this application, by inputting the expression information of the first molecule into a molecular generation model, can obtain the expression information of a second molecule that can bind to the first molecule without manual design. Furthermore, the expression information of the macrocyclic small molecule can be obtained from the expression information of the first and second molecules, thereby enabling drug researchers to directly obtain the chemical structure of the macrocyclic small molecule and greatly improving the efficiency of generating macrocyclic small molecules.
[0095] The above Figure 3The corresponding embodiment is a brief introduction to the training method of the molecular generation model provided in this application. It should be noted that before model training, the server first obtains a sample dataset for training the model. This sample dataset includes multiple pairs of sample molecules, each pair including expression information of a first sample molecule and expression information of a second sample molecule obtained by segmenting the same molecule. In some embodiments, the server can obtain expression information of multiple existing molecules from an associated molecular database, and segment these molecules based on their expression information to obtain the sample dataset. Optionally, the molecular database is the ChEMBL database.
[0096] For example, the molecular segmentation process is explained using any sample molecule pair as an example. The server obtains the expression information of the molecule from the associated molecular database. Based on the expression information of the molecule and the Matched Molecular Pairs (MMPs) algorithm, it determines the replaceable molecular fragments in the molecule. Based on the position of the molecular fragments in the molecule, the molecule is segmented to obtain the expression information of the first sample molecule and the expression information of the second sample molecule.
[0097] It should be noted that, in this embodiment, the server needs to perform two segments on each molecule to ensure that both the first and second sample molecules include two linker sites, so that the linker molecule obtained by the model can combine with the input molecule to form a macrocyclic structure. Accordingly, when the molecule is a small macrocyclic molecule, the server determines the replaceable molecular fragments in the macrocyclic structure of the molecule based on the molecule's expression information and the MMPs algorithm, and then segments the molecule. When the molecule is a small chain-like molecule, the server determines the replaceable molecular fragments in the molecule based on the molecule's expression information and the MMPs algorithm, and segments the molecule into three molecular fragments based on the position of the replaceable molecular fragments. The replaceable molecular fragments are used as the second sample molecule, and the remaining two molecular fragments are spliced together using chemical calculation tools to obtain the first sample molecule.
[0098] During the training of the molecular generation model, the server inputs multiple sample molecule pairs from the sample dataset into the molecular generation model in batches, and updates the network parameters of the molecular generation model through multiple iterations until the training termination condition is met. The following section combines... Figure 5 This paper provides a detailed description of the training method for the molecular generation model. The embodiments of this application use the server performing the first iteration of training as an example to illustrate the method. Figure 5 This is a flowchart of a training method for a molecular generation model provided in an embodiment of this application, such as... Figure 5 As shown, the method includes the following steps.
[0099] 501. The server obtains the expression information of the first sample molecule and the expression information of the second sample molecule. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule. The second sample molecule is a replaceable molecular fragment in the molecule.
[0100] In some embodiments, the server first randomly obtains a batch of sample molecule pairs from the sample dataset. Each sample molecule pair includes the expression information of a first sample molecule and the expression information of a second sample molecule. Then, the server initializes the molecule generation model by configuring the network parameters of the molecule generation model as initialization parameters. Optionally, the batch size is 8192.
[0101] It should be noted that if this training process is not the first iteration, the server does not need to initialize the molecular generation model and can directly execute the current training process based on the model obtained from the previous training.
[0102] 502. Based on the expression information of the first sample molecule, the first embedding vector sub-model and the encoding sub-model of the molecular generation model, the server obtains the molecular binding information of the first sample molecule. This molecular binding information is used to represent the characteristics of small molecules that can bind to the first sample molecule.
[0103] In this embodiment of the application, the molecular binding information of the first sample molecule in any sample molecule pair is used as an example for illustration.
[0104] In some embodiments, the server first inputs the expression information of the first sample molecule into the first embedding vector sub-model to obtain the embedding vector of the first sample molecule, and then inputs the embedding vector of the first sample molecule into the encoding sub-model to obtain the molecular binding information of the first sample molecule. It should be noted that this molecular binding information is in vector form.
[0105] The process of obtaining molecular binding information of the first sample molecule is illustrated by way of example. Figure 6As shown, the encoding sub-model includes a multi-head attention unit and a feedforward neural network unit. The server inputs the embedding vector of the first sample molecule into the multi-head attention unit of the encoding sub-model. Based on the embedding vector and the multi-head attention unit of the encoding sub-model, the correlation information between atoms in the first sample molecule and the correlation information between each atom and each chemical bond in the first sample molecule are extracted to obtain the initial molecular binding information of the first sample molecule. Specifically, the correlation information between any two atoms represents the degree of association between the two atoms in the first sample molecule; a higher correlation value indicates a higher degree of association between the two atoms. Similarly, the correlation information between any atom and any chemical bond represents the degree of association between the atom and the chemical bond in the first sample molecule; a higher correlation value indicates a higher degree of association between the atom and the chemical bond. Further, the server uses the feedforward neural network unit of the encoding sub-model to nonlinearly fuse the obtained multiple correlation information to obtain the molecular binding information of the first sample molecule.
[0106] In some embodiments, the server inputs the embedding vector of the first sample molecule and the first position code of the first sample molecule into the encoding sub-model to obtain the molecular binding information of the first sample molecule. The first position code is used to represent the position of each character in the expression information of the first sample molecule. Because the first position code is input, the data input into the encoding sub-model can better represent the chemical structure of the first sample molecule, thereby improving the accuracy of the molecular binding information of the first sample molecule.
[0107] In some embodiments, a normalization unit is connected after both the multi-head self-attention unit and the feedforward neural network unit of the encoded sub-model. The normalization unit is used to normalize the output data of the previous unit, thereby avoiding the gradient vanishing problem during model training.
[0108] In some embodiments, the molecular generation model includes N coded sub-models, where N is an integer greater than 0. These N coded sub-models are cascaded. The input data of the first coded sub-model is the embedding vector of the first sample molecule. The input data of the other coded sub-models is the output data of the preceding coded sub-model. The output data of the last coded sub-model is used as the molecular binding information of the first sample molecule. Optionally, N can be 6. Since each coded sub-model contains a multi-head self-attention unit and a feedforward neural network unit, by cascading the N coded sub-models, the correlation information of the first sample molecule can be extracted and fused multiple times, thus obtaining molecular binding information containing richer molecular information.
[0109] 503. The server obtains the expression information of the predicted molecule based on the molecular binding information of the first sample molecule, the expression information of the second sample molecule, the second embedding vector sub-model and the decoding sub-model of the molecular generation model.
[0110] In this embodiment of the application, the example of obtaining the expression information of the predicted molecule of any sample molecule pair is used for illustration.
[0111] In some embodiments, the process by which the server obtains the expression information of the predicted molecule is a multi-iterative process, where each iteration is used to obtain one character from the expression information of the predicted molecule. For example, taking the i-th character from the expression information of the predicted molecule in the i-th iteration as an example, i is an integer greater than 0 and less than or equal to the number of characters in the expression information of the second sample molecule. Figure 6 As shown, the molecular generation model also includes a linear transformation sub-model and a normalization sub-model. The server first inputs the first i-1 characters of the expression information of the second sample molecule into the second embedding vector sub-model to obtain the embedding vector of the first i-1 characters. Then, it inputs the molecular binding information of the first sample molecule and the embedding vector of the first i-1 characters into the decoding sub-model to obtain the decoding information of the i-th character in the expression information of the predicted molecule. Based on the linear transformation sub-model and the normalization sub-model, the decoded information is mapped to obtain the probabilities corresponding to multiple candidate characters. The probability of each candidate character refers to the probability that the i-th character in the expression information of the predicted molecule is the corresponding candidate character. These multiple candidate characters are all characters included in the pre-defined expression information of the molecules. The candidate character with the highest probability is determined as the i-th character in the expression information of the predicted molecule. Optionally, the normalization sub-model is a Softmax function.
[0112] For example, the process of obtaining the decoding information of the i-th character in the expression information of the predicted molecule based on the decoding sub-model is explained. Figure 6As shown, the decoding sub-model includes a multi-head self-attention unit, a multi-head encoder-decoder attention unit, and a feedforward neural network unit. The server inputs the embedding vectors of the first i-1 characters in the expression information of the second sample molecule into the multi-head self-attention unit of the decoding sub-model to extract the correlation information contained in the first i-1 characters in the expression information of the second sample molecule. The obtained correlation information and the molecular combination information of the first sample molecule are input into the multi-head encoder-decoder attention unit to obtain the initial decoding information of the i-th character in the expression information of the predicted molecule. The initial decoding information is input into the feedforward neural network unit of the decoding sub-model to perform a nonlinear transformation on the decoding information to obtain the decoding information of the i-th character in the expression information of the predicted molecule.
[0113] In some embodiments, during the process of obtaining decoding information described above, the server inputs the embedding vectors corresponding to the first i-1 characters in the expression information of the two sample molecules and the second position codes corresponding to the first i-1 characters into the decoding sub-model to obtain the decoding information of the i-th character in the expression information of the predicted molecule. The second position code represents the position of the first i-1 characters in the expression information of the second sample molecule. By inputting this second position code, the input data of the decoding sub-model contains richer information, enabling the decoding sub-model to decode the molecular binding information of the first sample molecule more accurately, thus improving the accuracy of the obtained decoding information.
[0114] In some embodiments, a normalization unit is connected after the multi-head self-attention unit, the multi-head encoder-decoder attention unit, and the feedforward neural network unit of the decoding sub-model. The normalization unit is used to normalize the output data of the previous unit, thereby avoiding the gradient vanishing problem during model training.
[0115] In some embodiments, the molecular generation model includes N decoding sub-models in a cascaded structure. For the process of obtaining the decoding information of the i-th character, the input data of the first decoding sub-model is the embedding vector of the first i-1 characters in the expression information of the second sample molecule and the molecular binding information of the first sample molecule. Excluding the first encoding sub-model, the input data of the other encoding sub-models is the output data of the previous encoding sub-model and the molecular binding information of the first sample molecule. The server uses the output data of the last encoding sub-model as the decoding information of the i-th character in the expression information of the predicted molecule. By cascading the N decoding sub-models, the output data of the previous encoding sub-model can be further decoded based on the molecular binding information of the first sample molecule, thereby making the decoding information obtained by the last encoding sub-model more accurate and improving the accuracy of the obtained expression information.
[0116] It should be noted that steps 502 to 503 above are illustrated using the example of obtaining the expression information of the predicted molecule through the expression information of the first sample molecule and the expression information of the second sample molecule. In some embodiments, the server inputs the expression information of the first sample molecule, the expression information of the second sample molecule, the length and structural information of the second sample molecule into the molecular generation model to obtain the expression information of the predicted molecule. The structural information is used to indicate whether the second sample molecule includes a ring structure. Optionally, the structural information is represented by ring_1 and ring_0, where ring_1 indicates that the second sample molecule includes a ring structure, and ring_0 indicates that the second sample molecule does not include a ring structure. The length of the second sample molecule is represented by L_num, where num is an integer greater than 0. For example, L_8 indicates that the length of the second sample molecule is 8.
[0117] For example, such as Figure 6 As shown, the server concatenates the length and structural information of the second sample molecule before the expression information of the first sample molecule. Based on the concatenated string, the first embedding vector sub-model, and the encoding sub-model, the molecular binding information of the first sample molecule is obtained through a method similar to step 502. Based on the molecular binding information of the first sample molecule, the expression information of the second sample molecule, the second embedding vector sub-model, and the decoding sub-model, the expression information of the predicted molecule is obtained through a method similar to step 503. Since the model's input data includes the length and structural information of the second sample molecule, the accuracy of the molecular binding information of the first sample molecule is improved. Furthermore, based on this molecular binding information, the expression information of the predicted molecule is obtained, thus reducing the error between the obtained expression information of the predicted molecule and the expression information of the second sample molecule. This allows the model to learn the method of generating linker molecules more quickly, improving the model's training speed.
[0118] It should be noted that the process of obtaining the expression information of the predicted molecule in steps 502 to 503 above is also called the forward calculation process.
[0119] It should be noted that if the molecular generation model includes only one embedding vector sub-model, then steps 502 to 503 above can be replaced by: the server obtaining the molecular binding information of the first sample molecule based on the expression information of the first sample molecule, the embedding vector sub-model of the molecular generation model, and the encoding sub-model. This molecular binding information is used to represent the features of small molecules that can bind to the first sample molecule. The server then obtains the expression information of the predicted molecule based on the molecular binding information of the first sample molecule, the expression information of the second sample molecule, the embedding vector sub-model of the molecular generation model, and the decoding sub-model.
[0120] 504. The server trains the molecule generation model based on the expression information of the predicted molecule and the expression information of the second sample molecule.
[0121] In some embodiments, for any sample molecule pair in a batch of sample molecule pairs, the server obtains a loss value based on the expression information of the predicted molecule and the expression information of the second sample molecule. This loss value represents the error between the expression information of the predicted molecule and the expression information of the second sample molecule. The server takes the mean of the loss values of the batch of sample molecule pairs as the target loss value for this training process. If the training process does not reach the training termination condition, the server adjusts the network parameters of the molecule generation model based on the target loss value. If the training process reaches the training termination condition, the server outputs the molecule generation model. Optionally, the training termination condition is that the target loss value is less than a first threshold, or that the number of training iterations reaches a second threshold. Optionally, during training, the learning rate is set to 0.001, the dropout rate is set to 0.1, the weight decay coefficient is set to 0.000001, and the training epochs are set to 100.
[0122] For example, the process of updating the network parameters of a molecular generation model is described. Based on the target loss value, the server obtains the gradient of each network layer in the molecular generation model through a back-forward algorithm. Based on the gradient of each network layer, the network parameters of the molecular generation model are updated through an adaptive momentum estimation (Adam) algorithm.
[0123] The technical solution provided in this application involves inputting the expression information of the first sample molecule and the expression information of the second sample molecule obtained from existing molecular segmentation into a molecular generation model. This yields the expression information of the linker molecule predicted by the model for the first sample molecule. Furthermore, the model is trained using the predicted expression information and the expression information of the second sample molecule, enabling the model to accurately generate the expression information of the linker molecule. This eliminates the need for drug researchers to manually design linker molecules, significantly improving the efficiency of generating macrocyclic small molecules.
[0124] The following combination Figure 7 This application introduces the method for determining information about macrocyclic small molecules. Figure 7 This is a flowchart of a method for determining information about macrocyclic small molecules provided in this application, such as... Figure 7 As shown, this method is executed by the server and includes the following steps.
[0125] 701. The server obtains the expression information of the first molecule.
[0126] In some embodiments, the terminal provides a molecule generation page, on which researchers can input the expression information of a first molecule. In response to the submission operation on the molecule generation page, the terminal obtains the expression information of the first molecule and sends a molecule information retrieval request M1 to the server. The molecule information retrieval request M1 carries the expression information of the first molecule and is used to instruct the return of the expression information of the macrocyclic small molecule corresponding to the first molecule. The server receives the molecule information retrieval request M1 and obtains the expression information of the first molecule carried by the molecule information retrieval request M1.
[0127] Optionally, researchers input the identifier of the first molecule on the molecule generation page, and the terminal sends a molecule information retrieval request M2 to the server. The molecule information retrieval request M2 carries the identifier of the first molecule and is used to instruct the return of the expression information of the macrocyclic small molecule corresponding to the first molecule. The server receives the molecule information retrieval request M2, retrieves the identifier of the first molecule carried in the molecule information retrieval request M2, and retrieves the expression information of the first molecule from the associated molecule database based on the identifier of the first molecule.
[0128] 702. Based on the expression information of the first molecule, the first embedding vector sub-model and the encoding sub-model of the molecular generation model, the server obtains the molecular binding information of the first molecule, which is used to represent the characteristics of small molecules that can bind to the first molecule.
[0129] In some embodiments, the server obtains the molecular binding information of the first molecule using a method similar to step 502 described above.
[0130] 703. The server obtains the expression information of the second molecule based on the molecular binding information of the first molecule, the second embedding vector sub-model of the molecular generation model, and the decoding sub-model.
[0131] In some embodiments, the process by which the server obtains the expression information of the second molecule is a multi-iterative process, where each iteration is used to obtain one character from the expression information of the second molecule. For example, taking the j-th iteration as an example of obtaining the j-th character from the expression information of the second molecule, where j is an integer greater than 0, the server first inputs the first j-1 characters from the expression information of the second molecule obtained in the previous j-1 iterations into a second embedding vector sub-model to obtain the embedding vector of the first j-1 characters. Then, the server inputs the molecule combination information of the first molecule and the embedding vector of the first j-1 characters into a decoding sub-model to obtain the decoding information of the j-th character from the expression information of the second molecule. Based on a linear transformation sub-model and a normalization sub-model, the decoding information is mapped to obtain the probabilities corresponding to multiple candidate characters. The candidate character with the highest probability is determined as the j-th character from the expression information of the second molecule.
[0132] The process of obtaining the decoding information of the j-th character in the expression information of the second molecule based on the decoding sub-model is the same as the process of obtaining the decoding information of the i-th character in the expression information of the predicted molecule in step 503 above, and will not be repeated here.
[0133] It should be noted that steps 702 to 703 above are illustrated using the example of obtaining the expression information of the second molecule from the expression information of the first molecule. In some embodiments, the server inputs the molecule length, molecular structure information, and the expression information of the first molecule into the molecular generation model to obtain the expression information of the second molecule. Here, molecule length represents the length of the small molecule bound to the first molecule, and molecular structure information represents whether the small molecule bound to the first molecule includes a ring structure.
[0134] By inputting molecular length and structure information into a molecular generation model, researchers can control the length and structure of the resulting linker molecules. By inputting different molecular lengths and structures, researchers can obtain linker molecules with different structures corresponding to the first molecule. Furthermore, through linker molecules with different structures, researchers can obtain a variety of macrocyclic small molecules corresponding to the first molecule. Researchers can then select macrocyclic small molecules with better pharmacological activity and selectivity for subsequent drug clinical research, thereby enabling the developed drugs to have better efficacy.
[0135] It should be noted that if the molecular generation model includes only one embedding vector sub-model, then steps 702 to 703 above can be replaced by: the server obtaining the molecular binding information of the first molecule based on the expression information of the first molecule, the embedding vector sub-model of the molecular generation model, and the encoding sub-model. This molecular binding information is used to represent the characteristics of small molecules that can bind to the first molecule. The server then obtains the expression information of the second molecule based on the molecular binding information of the first molecule, the embedding vector sub-model of the molecular generation model, and the decoding sub-model.
[0136] 704. The server obtains the expression information of macrocyclic small molecules based on the expression information of the first molecule and the expression information of the second molecule.
[0137] In some embodiments, the server determines the binding mode of the first molecule and the second molecule based on the first and second linking sites in the expression information of the first molecule, and the third and fourth linking sites in the expression information of the second molecule. Based on the binding mode, the server determines the chemical structure of the macrocyclic small molecule, and obtains the expression information of the macrocyclic small molecule based on its chemical structure. The binding mode refers to binding the first molecule and the second molecule based on the first linking site of the first molecule and the corresponding third linking site of the second molecule, and the second linking site of the first molecule and the corresponding fourth linking site of the second molecule.
[0138] Optionally, the server can use chemical computing tools, such as RDKit, to perform the above-described process of obtaining expression information of macrocyclic small molecules.
[0139] In some embodiments, the method further includes: the server sending the obtained expression information of the macrocyclic small molecule to the terminal, and the terminal displaying the expression information of the macrocyclic small molecule on a molecule generation page. Optionally, the server sending the expression information of a second molecule and the expression information of the macrocyclic small molecule to the terminal, and the terminal displaying the expression information of the second molecule and the expression information of the macrocyclic small molecule on a molecule generation page.
[0140] The technical solution provided in this application, by inputting the expression information of the first molecule into a molecular generation model, can obtain the expression information of a second molecule that can bind to the first molecule without manual design. Furthermore, the expression information of the macrocyclic small molecule can be obtained from the expression information of the first and second molecules, thereby enabling drug researchers to directly obtain the chemical structure of the macrocyclic small molecule and greatly improving the efficiency of generating macrocyclic small molecules.
[0141] Figure 8 This is a schematic diagram of the structure of a training device for a molecular generation model provided in an embodiment of this application, as shown below. Figure 8 As shown, the device includes: an acquisition module 801, a generation module 802, and a training module 803.
[0142] The acquisition module 801 is used to acquire the expression information of the first sample molecule and the expression information of the second sample molecule. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment in the molecule.
[0143] The generation module 802 is used to input the expression information of the first sample molecule and the expression information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule.
[0144] Training module 803 is used to train the molecule generation model based on the expression information of the predicted molecule and the expression information of the second sample molecule.
[0145] In some embodiments, the generation module 802 includes:
[0146] The acquisition submodule is used to input the expression information of the first sample molecule into the first embedding vector sub-model of the molecule generation model to obtain the embedding vector of the first sample molecule.
[0147] The acquisition submodule is used to input the embedding vector of the first sample molecule into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first sample molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first sample molecule.
[0148] The acquisition submodule is used to input the expression information of the second sample molecule into the second embedding sub-model of the molecule generation model to obtain the embedding vector of the second sample molecule.
[0149] The generation submodule is used to input the molecular binding information of the first sample molecule and the embedding vector of the second sample molecule into the decoding submodule of the molecule generation model to obtain the expression information of the predicted molecule.
[0150] In some embodiments, the generation submodule is used to input the embedding vector of the first sample molecule into the encoding submodel; based on the embedding vector of the first sample molecule and the multi-head self-attention unit in the encoding submodel, to extract the correlation information between each atom in the first sample molecule and the correlation information between each atom and each chemical bond in the first sample molecule, so as to obtain the initial molecular binding information of the first sample molecule.
[0151] Based on the feedforward neural network unit of this coding sub-model, multiple correlation information in the initial molecular binding information of the first sample molecule is nonlinearly fused to obtain the molecular binding information of the first sample molecule.
[0152] In some embodiments, the generation module 802 is used to input the expression information of the first sample molecule, the expression information of the second sample molecule, and the length and structure information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule, wherein the structure information is used to indicate whether the second sample molecule includes a ring structure.
[0153] In some embodiments, the training module 803 is used to obtain a loss value based on the expression information of the predicted molecule and the expression information of the second sample molecule, the loss value being used to represent the error between the expression information of the predicted molecule and the expression information of the second sample molecule; if the training process does not reach the training termination condition, the network parameters of the molecule generation model are adjusted based on the loss value; if the training process reaches the training termination condition, the molecule generation model is output.
[0154] In some embodiments, the device further includes a segmentation module for determining replaceable molecular fragments in the molecule based on a matching molecule pair algorithm and the expression information of the molecule; and segmenting the molecule based on the position of the molecular fragment in the molecule to obtain the expression information of the first sample molecule and the expression information of the second sample molecule, wherein the expression information of the first sample molecule and the expression information of the second sample molecule include the segmentation position of the molecule.
[0155] It should be noted that the molecular generation model training device provided in the above embodiments is only illustrated by the division of the above functional modules when training the molecular generation model. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the molecular generation model training device and the molecular generation model training method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0156] Figure 9 This is a schematic diagram of the structure of a device for determining the information of macrocyclic small molecules provided in an embodiment of this application, as shown below. Figure 9 As shown, the device includes an acquisition module 901 and a generation module 902.
[0157] The acquisition module 901 is used to acquire the expression information of the first molecule;
[0158] The generation module 902 is used to input the expression information of the first molecule into the molecular generation model to obtain the expression information of the second molecule, which is a small molecule that can bind to the first molecule;
[0159] The acquisition module 901 is used to acquire the expression information of macrocyclic small molecules based on the expression information of the first molecule and the expression information of the second molecule;
[0160] The molecular generation model is trained based on the expression information of the first sample molecule and the expression information of the second sample molecule. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment in the molecule.
[0161] In some embodiments, the generation module 902 includes:
[0162] The acquisition submodule is used to input the expression information of the first molecule into the first embedding sub-model of the molecule generation model to obtain the embedding vector of the first molecule;
[0163] The acquisition submodule is used to input the embedding vector of the first molecule into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first molecule.
[0164] The generation submodule is used to input the molecular binding information of the first molecule into the decoding submodule of the molecule generation model to obtain the expression information of the second molecule.
[0165] In some embodiments, the generation submodule is used to input the embedding vector of the first molecule into the encoding submodel; based on the embedding vector of the first molecule and the multi-head self-attention unit of the encoding submodel, to extract the correlation information between each atom in the first molecule and the correlation information between each atom in the first molecule and each chemical bond, so as to obtain the initial molecular binding information of the first molecule.
[0166] Based on the feedforward neural network unit of this coding sub-model, multiple correlation information in the initial molecular binding information of the first molecule is nonlinearly fused to obtain the molecular binding information of the first molecule.
[0167] In some embodiments, the generation module 902 is used to input the molecular length, molecular structure information and the expression information of the first molecule into the molecular generation model to obtain the expression information of the second molecule. The molecular length is used to represent the length of the small molecule bound to the first molecule, and the molecular structure information is used to represent whether the small molecule bound to the first molecule includes a ring structure.
[0168] In some embodiments, the expression information of the first molecule includes a first linking site and a second linking site of the first molecule, the first linking site and the second linking site being atoms in the first molecule used to bind with other molecules; the expression information of the second molecule includes a third linking site and a fourth linking site of the second molecule, the third linking site and the fourth linking site being atoms in the second molecule that bind with the first linking site and the second linking site, respectively.
[0169] The acquisition module 901 is used to determine the binding mode between the first molecule and the second molecule based on the first linking site and the second linking site of the first molecule, and the third linking site and the fourth linking site of the second molecule; to determine the chemical structure of the macrocyclic small molecule based on the binding mode; and to acquire the expression information of the macrocyclic small molecule based on the chemical structure of the macrocyclic small molecule.
[0170] It should be noted that the macrocyclic small molecule information determination device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the macrocyclic small molecule information determination device and the macrocyclic small molecule information determination method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0171] This disclosure provides a computer device for performing the training method of the above-described molecular generation model or the information determination method for macrocyclic small molecules. In some embodiments, the computer device is provided as a server. Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application, such as... Figure 10 As shown, the server 1000 can vary considerably due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 1001 and one or more memories 1002, wherein the one or more memories 1002 store at least one line of program code, which is loaded and executed by the one or more processors 1001 to implement the methods provided in the above-described method embodiments. Of course, the server 1000 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1000 may also include other components for implementing device functions, which will not be elaborated here.
[0172] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including at least one line of program code, which can be executed by a processor to complete the training method for the molecular generation model or the method for determining information of macrocyclic small molecules in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0173] In an exemplary embodiment, a computer program product is also provided, comprising at least one computer program stored in a computer-readable storage medium. A processor of a computer device reads the at least one computer program from the computer-readable storage medium and executes the at least one computer program, causing the computer device to perform the operations performed by the aforementioned molecular generation model training method or macrocyclic small molecule information determination method.
[0174] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.
[0175] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0176] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a molecular generation model, characterized in that, The method includes: The expression information of the first sample molecule and the expression information of the second sample molecule are obtained. The first sample molecule and the second sample molecule are obtained by segmenting the same molecule. The second sample molecule is a replaceable molecular fragment in the molecule. The expression information of the first sample molecule and the expression information of the second sample molecule are input into the molecule generation model to obtain the expression information of the predicted molecule. Based on the expression information of the predicted molecule and the expression information of the second sample molecule, a loss value is obtained, which is used to represent the error between the expression information of the predicted molecule and the expression information of the second sample molecule. If the training process does not reach the training termination condition, the network parameters of the molecular generation model are adjusted based on the loss value; if the training process reaches the training termination condition, the molecular generation model is output.
2. The method according to claim 1, characterized in that, The step of inputting the expression information of the first sample molecule and the expression information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule includes: The expression information of the first sample molecule is input into the first embedding vector sub-model of the molecule generation model to obtain the embedding vector of the first sample molecule. The embedding vector of the first sample molecule is input into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first sample molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first sample molecule. The expression information of the second sample molecule is input into the second embedding sub-model of the molecule generation model to obtain the embedding vector of the second sample molecule; The molecular binding information of the first sample molecule and the embedding vector of the second sample molecule are input into the decoding sub-model of the molecular generation model to obtain the expression information of the predicted molecule.
3. The method according to claim 2, characterized in that, The step of inputting the embedding vector of the first sample molecule into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first sample molecule includes: The embedding vector of the first sample molecule is input into the encoding sub-model; Based on the embedding vector of the first sample molecule and the multi-head self-attention unit in the coding sub-model, the correlation information between each atom in the first sample molecule and the correlation information between each atom and each chemical bond in the first sample molecule are extracted to obtain the initial molecular binding information of the first sample molecule. Based on the feedforward neural network unit of the encoding sub-model, multiple correlation information in the initial molecular binding information of the first sample molecule are nonlinearly fused to obtain the molecular binding information of the first sample molecule.
4. The method according to claim 1, characterized in that, The step of inputting the expression information of the first sample molecule and the expression information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule includes: The expression information of the first sample molecule, the expression information of the second sample molecule, and the length and structure information of the second sample molecule are input into the molecule generation model to obtain the expression information of the predicted molecule. The structure information is used to indicate whether the second sample molecule includes a ring structure.
5. The method according to claim 1, characterized in that, The process of splitting the same molecule includes: Based on the matching molecule pair algorithm and the expression information of the molecule, replaceable molecular fragments in the molecule are determined; Based on the position of the molecular fragment within the molecule, the molecule is segmented to obtain the expression information of the first sample molecule and the expression information of the second sample molecule, wherein the expression information of the first sample molecule and the expression information of the second sample molecule include the segmentation position of the molecule.
6. A method for determining information about macrocyclic small molecules, characterized in that, The method includes: Obtain the expression information of the first molecule; The expression information of the first molecule is input into the molecular generation model to obtain the expression information of the second molecule, which is a small molecule that can bind to the first molecule. Based on the expression information of the first molecule and the expression information of the second molecule, the expression information of the macrocyclic small molecule is obtained; The training process of the molecular generation model includes: inputting the expression information of a first sample molecule and the expression information of a second sample molecule into the molecular generation model to obtain the expression information of a predicted molecule, wherein the first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment within the molecule; obtaining a loss value based on the expression information of the predicted molecule and the expression information of the second sample molecule, wherein the loss value is used to represent the error between the expression information of the predicted molecule and the expression information of the second sample molecule; if the training process does not reach the training termination condition, adjusting the network parameters of the molecular generation model based on the loss value; and outputting the molecular generation model if the training process reaches the training termination condition.
7. The method according to claim 6, characterized in that, The step of inputting the expression information of the first molecule into the molecule generation model to obtain the expression information of the second molecule includes: The expression information of the first molecule is input into the first embedding sub-model of the molecule generation model to obtain the embedding vector of the first molecule; The embedding vector of the first molecule is input into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first molecule. The molecular binding information of the first molecule is input into the decoding sub-model of the molecular generation model to obtain the expression information of the second molecule.
8. The method according to claim 6, characterized in that, The step of inputting the expression information of the first molecule into the molecule generation model to obtain the expression information of the second molecule includes: The molecular length, molecular structure information, and expression information of the first molecule are input into the molecular generation model to obtain the expression information of the second molecule. The molecular length is used to represent the length of the small molecule that binds to the first molecule, and the molecular structure information is used to represent whether the small molecule that binds to the first molecule includes a ring structure.
9. The method according to claim 6, characterized in that, The expression information of the first molecule includes the first linking site and the second linking site of the first molecule, wherein the first linking site and the second linking site are atoms in the first molecule that are bound to other molecules. The expression information of the second molecule includes the third linking site and the fourth linking site of the second molecule, wherein the third linking site and the fourth linking site are atoms in the second molecule that are bound to the first linking site and the second linking site, respectively. The step of obtaining the expression information of macrocyclic small molecules based on the expression information of the first molecule and the expression information of the second molecule includes: Based on the first and second linking sites of the first molecule, and the third and fourth linking sites of the second molecule, the binding mode between the first molecule and the second molecule is determined. Based on the aforementioned binding mechanism, the chemical structure of the macrocyclic small molecule was determined; Based on the chemical structure of the macrocyclic small molecule, the expression information of the macrocyclic small molecule is obtained.
10. A training device for a molecular generation model, characterized in that, The device comprises: The acquisition module is used to acquire the expression information of a first sample molecule and a second sample molecule, wherein the first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment in the molecule. The generation module is used to input the expression information of the first sample molecule and the expression information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule. The training module is used to obtain a loss value based on the expression information of the predicted molecule and the expression information of the second sample molecule, the loss value being used to represent the error between the expression information of the predicted molecule and the expression information of the second sample molecule; if the training process does not reach the training termination condition, the network parameters of the molecule generation model are adjusted based on the loss value; if the training process reaches the training termination condition, the molecule generation model is output.
11. The apparatus according to claim 10, characterized in that, The generation module includes: The acquisition submodule is used to input the expression information of the first sample molecule into the first embedding vector sub-model of the molecule generation model to obtain the embedding vector of the first sample molecule. The acquisition submodule is used to input the embedding vector of the first sample molecule into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first sample molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first sample molecule. The acquisition submodule is used to input the expression information of the second sample molecule into the second embedding sub-model of the molecule generation model to obtain the embedding vector of the second sample molecule. The generation submodule is used to input the molecular binding information of the first sample molecule and the embedding vector of the second sample molecule into the decoding submodule of the molecular generation model to obtain the expression information of the predicted molecule.
12. The apparatus according to claim 11, characterized in that, The generation submodule is used to input the embedding vector of the first sample molecule into the encoding submodel; based on the embedding vector of the first sample molecule and the multi-head self-attention unit in the encoding submodel, it extracts the correlation information between each atom in the first sample molecule and the correlation information between each atom and each chemical bond in the first sample molecule to obtain the initial molecular binding information of the first sample molecule. Based on the feedforward neural network unit of the encoding sub-model, multiple correlation information in the initial molecular binding information of the first sample molecule are nonlinearly fused to obtain the molecular binding information of the first sample molecule.
13. The apparatus according to claim 10, characterized in that, The generation module is used to input the expression information of the first sample molecule, the expression information of the second sample molecule, and the length and structure information of the second sample molecule into the molecule generation model to obtain the expression information of the predicted molecule. The structure information is used to indicate whether the second sample molecule includes a ring structure.
14. The apparatus according to claim 10, characterized in that, The device further includes a segmentation module for determining replaceable molecular fragments in the molecule based on a matching molecule pair algorithm and the expression information of the molecule; and segmenting the molecule based on the position of the molecular fragments in the molecule to obtain the expression information of the first sample molecule and the expression information of the second sample molecule, wherein the expression information of the first sample molecule and the expression information of the second sample molecule include the segmentation position of the molecule.
15. An information determination device for macrocyclic small molecules, characterized in that, The device comprises: The acquisition module is used to acquire the expression information of the first molecule; The generation module is used to input the expression information of the first molecule into the molecular generation model to obtain the expression information of the second molecule, wherein the second molecule is a small molecule that can bind to the first molecule; The acquisition module is used to acquire the expression information of macrocyclic small molecules based on the expression information of the first molecule and the expression information of the second molecule; The training process of the molecular generation model includes: inputting the expression information of a first sample molecule and the expression information of a second sample molecule into the molecular generation model to obtain the expression information of a predicted molecule, wherein the first sample molecule and the second sample molecule are obtained by segmenting the same molecule, and the second sample molecule is a replaceable molecular fragment within the molecule; obtaining a loss value based on the expression information of the predicted molecule and the expression information of the second sample molecule, wherein the loss value is used to represent the error between the expression information of the predicted molecule and the expression information of the second sample molecule; if the training process does not reach the training termination condition, adjusting the network parameters of the molecular generation model based on the loss value; and outputting the molecular generation model if the training process reaches the training termination condition.
16. The apparatus according to claim 15, characterized in that, The generation module includes: The acquisition submodule is used to input the expression information of the first molecule into the first embedding sub-model of the molecule generation model to obtain the embedding vector of the first molecule; The acquisition submodule is used to input the embedding vector of the first molecule into the encoding sub-model of the molecule generation model to obtain the molecular binding information of the first molecule. The molecular binding information is used to represent the characteristics of small molecules that can bind to the first molecule. The generation submodule is used to input the molecular binding information of the first molecule into the decoding submodule of the molecular generation model to obtain the expression information of the second molecule.
17. The apparatus according to claim 15, characterized in that, The generation module is used to input the molecular length, molecular structure information, and expression information of the first molecule into the molecular generation model to obtain the expression information of the second molecule. The molecular length is used to represent the length of the small molecule that binds to the first molecule, and the molecular structure information is used to represent whether the small molecule that binds to the first molecule includes a ring structure.
18. The apparatus according to claim 15, characterized in that, The expression information of the first molecule includes the first linking site and the second linking site of the first molecule, wherein the first linking site and the second linking site are atoms in the first molecule that are bound to other molecules. The expression information of the second molecule includes the third linking site and the fourth linking site of the second molecule, wherein the third linking site and the fourth linking site are atoms in the second molecule that are bound to the first linking site and the second linking site, respectively. The acquisition module is used to determine the binding mode between the first molecule and the second molecule based on the first and second linking sites of the first molecule, and the third and fourth linking sites of the second molecule; to determine the chemical structure of the macrocyclic small molecule based on the binding mode; and to acquire the expression information of the macrocyclic small molecule based on the chemical structure of the macrocyclic small molecule.
19. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the training method for the molecular generation model as described in any one of claims 1 to 5, or the information determination method for macrocyclic small molecules as described in any one of claims 6 to 9.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the training method for the molecular generation model as described in any one of claims 1 to 5, or the information determination method for macrocyclic small molecules as described in any one of claims 6 to 9.
21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method for the molecular generation model as described in any one of claims 1 to 5, or the information determination method for macrocyclic small molecules as described in any one of claims 6 to 9.
Citation Information
Patent Citations
Molecule generation method and device, computer readable storage medium and terminal equipment
CN111508568A
Article molecule generation method, device and equipment, and storage medium
CN112199884A