Drug molecule generation method and device, computer device, and storage medium
By extracting the type and structural characteristics of protein molecules, information on drug molecules that are active against protein molecules is directly generated, solving the problem of low efficiency in drug molecule generation in existing technologies and achieving highly efficient drug molecule generation.
Patent Information
- Application Number
- CN202211085588.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-09-06
AI Technical Summary
Existing methods for generating drug molecules are cumbersome and inefficient, making it difficult to efficiently screen for drug molecules that are active against protein molecules.
By identifying information about protein molecules, extracting the type and structural features of protein nodes, and using feature encoding and decoding methods, information about drug molecules that are active against protein molecules is generated, simplifying the generation process.
It improves the efficiency of drug molecule generation, allowing for the direct acquisition of effective drug molecules without the need for cumbersome screening processes.
Smart Images

Figure CN116959614B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, and storage medium for generating drug molecules. Background Technology
[0002] Drug molecules refer to the chemical structure of a drug, and drug development can be achieved by generating drug molecules. When developing new drugs, computer simulation methods are typically used to synthesize candidate drug molecules with certain desired properties based on existing drug molecules. Then, through pharmacological experiments, drug molecules that are active against protein molecules are screened from the generated candidate drug molecules; these are the effective drug molecules. However, this method is relatively cumbersome and has low efficiency in generating drug molecules. Summary of the Invention
[0003] This application provides a method, apparatus, computer device, and storage medium for generating drug molecules, which can improve the efficiency of drug molecule generation. The technical solution is as follows:
[0004] On the one hand, a method for generating drug molecules is provided, the method comprising:
[0005] Determine the protein molecule information corresponding to the protein molecule, wherein the protein molecule includes multiple protein nodes, and the protein molecule information is used to represent the type and structure of the multiple protein nodes;
[0006] The protein molecule information is feature-encoded to obtain protein molecule features;
[0007] The protein molecule features are decoded to obtain drug molecule information, and the drug molecule information indicates that the drug molecule is active against the protein molecule.
[0008] Optionally, the step of extracting the second-order type feature corresponding to the first-order type feature based on the first-order connectivity feature and the first-order type feature includes:
[0009] Second-order reference protein nodes and second-order non-reference protein nodes are identified among the plurality of protein nodes;
[0010] The first-order type feature of the second-order reference protein node is determined as the second-order type feature of the second-order reference protein node;
[0011] Based on the first-order connectivity features, a linear transformation is performed on the first-order type features of the second-order reference protein node to obtain second-order type transformation features. The second-order type transformation features represent the degree of influence of the type of the second-order reference protein node on the type of the second-order non-reference protein node.
[0012] Based on the second-order type transformation features, the first-order type features of the second-order non-reference protein nodes are transformed to obtain the second-order type features of the second-order non-reference protein nodes.
[0013] Optionally, the step of extracting the first-order position feature corresponding to the first position information based on the first connection information and the first position information includes:
[0014] Among the plurality of protein nodes, first-order reference protein nodes and first-order non-reference protein nodes are identified;
[0015] The first position information of the first-order reference protein node is determined as the first-order position feature of the first-order reference protein node.
[0016] Based on the first connection information, graph convolution is performed on the first position information of the first-order reference protein node to obtain the first-order position transformation feature. The first-order position transformation feature represents the degree of influence of the position of the first-order reference protein node on the position of the first-order non-reference protein node.
[0017] Based on the first-order position transformation feature, the first position information of the first-order non-reference protein node is transformed to obtain the first-order position feature of the first-order non-reference protein node.
[0018] Optionally, the step of extracting the second-order position feature corresponding to the first-order position feature based on the first-order connectivity feature and the first-order position feature includes:
[0019] Second-order reference protein nodes and second-order non-reference protein nodes are identified among the plurality of protein nodes;
[0020] The first-order position features of the second-order reference protein node are determined as the second-order position features of the second-order reference protein node.
[0021] Based on the first-order connectivity features, a linear transformation is performed on the first-order position features of the second-order reference protein node to obtain second-order position transformation features. The second-order position transformation features represent the degree of influence of the position of the second-order reference protein node on the position of the second-order non-reference protein node.
[0022] Based on the second-order position transformation features, the first-order position features of the second-order non-reference protein node are transformed to obtain the second-order position features of the second-order non-reference protein node.
[0023] Optionally, the step of extracting the first-order connection feature corresponding to the first connection information based on the first-order type feature, the first-order position feature, and the first connection information includes:
[0024] Among the plurality of protein nodes, first-order reference protein nodes and first-order non-reference protein nodes are identified;
[0025] The first connection information of the first-order reference protein node is determined as the first-order connection feature of the first-order reference protein node.
[0026] After fusing the first-order type feature, the first-order position feature, and the first connection information of the first-order reference protein node, feature extraction is performed to obtain the first-order connection transformation feature, which represents the degree of influence of the connection status of the first-order reference protein node on the connection status of the first-order non-reference protein node.
[0027] Based on the first-order connectivity transformation features, the first connectivity information of the first-order non-reference protein node is transformed to obtain the first-order connectivity features of the first-order non-reference protein node.
[0028] Optionally, the step of decoding the protein molecule features to obtain drug molecule information includes:
[0029] The first connection feature and the first type feature are decoded to obtain the second type information;
[0030] The first connection feature and the first location feature are decoded to obtain the second location information;
[0031] The first type feature, the first location feature, and the first connection feature are decoded to obtain the second connection information;
[0032] The second type information, the second location information, and the second connection information constitute the drug molecule information.
[0033] Optionally, the sample drug molecule is a real drug molecule; the method further includes:
[0034] The first discrimination model is invoked to discriminate the sample drug molecule information to obtain a second probability, which represents the probability that the sample drug molecule information is real drug molecule information.
[0035] Based on the first probability and the second probability, the first discrimination model is trained so that the first probability obtained by the first discrimination model decreases and the second probability obtained increases.
[0036] Optionally, the sample protein molecule is a real protein molecule; the method further includes:
[0037] The second discrimination model is invoked to discriminate the sample protein molecule information and obtain a fourth probability, which represents the probability that the sample protein molecule information is real protein molecule information.
[0038] Based on the third probability and the fourth probability, the second discrimination model is trained so that the third probability obtained by the second discrimination model decreases and the fourth probability obtained increases.
[0039] Optionally, the sample drug molecule is active against the sample protein molecule; the method further includes:
[0040] The drug processing model is invoked to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information;
[0041] The coding model is invoked to encode the sample drug molecule information and the predicted drug molecule information respectively, to obtain sample coding features and predicted coding features;
[0042] The protein processing model and the drug processing model are trained based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, respectively, to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and to increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution, including:
[0043] Based on the distribution information of the sample protein molecular features, the sample coding features, and the predicted coding features, the protein processing model is trained to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and to increase the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0044] Based on the distribution information of the sample drug molecule features, the sample coding features, and the predicted coding features, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0045] Optionally, the method further includes:
[0046] The coding model is trained based on the sample coding features and the predicted coding features to reduce the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0047] Optionally, the method further includes:
[0048] The matching model is invoked to determine a second matching degree based on the sample protein molecule information and the sample drug molecule information. The second matching degree represents the probability that the sample drug molecule is active against the sample protein molecule.
[0049] Based on the first matching degree and the second matching degree, the matching model is trained so that the first matching degree obtained by the matching model decreases and the second matching degree increases.
[0050] On the other hand, a drug molecule generation apparatus is provided, the apparatus comprising:
[0051] An information determination module is used to determine the protein molecule information corresponding to a protein molecule, wherein the protein molecule includes multiple protein nodes, and the protein molecule information is used to represent the type and structure of the multiple protein nodes.
[0052] The feature encoding module is used to encode the protein molecule information to obtain protein molecule features;
[0053] The feature decoding module is used to perform feature decoding on the protein molecule features to obtain drug molecule information, wherein the drug molecule information indicates that the drug molecule is active against the protein molecule.
[0054] Optionally, the protein molecule information includes first type information, first position information, and first connection information of the plurality of protein nodes, wherein the first connection information indicates the connection between the plurality of protein nodes, and the protein molecule features include first type features corresponding to the first type information, first position features corresponding to the first position information, and first connection features corresponding to the first connection information;
[0055] The feature encoding module includes:
[0056] The first extraction unit is used to extract the first type feature based on the first type information and the first connection information;
[0057] The second extraction unit is used to extract the first location feature based on the first location information and the first connection information;
[0058] The third extraction unit is used to extract the first connection feature based on the first connection information, the first type information, and the first location information.
[0059] Optionally, the first connectivity feature includes first-order connectivity features; the first extraction unit is configured to:
[0060] Based on the first connection information and the first type information, extract the first-order type feature corresponding to the first type information;
[0061] Based on the first-order connectivity feature and the first-order type feature, extract the second-order type feature corresponding to the first-order type feature until the current order reaches the target order, and then determine the currently obtained type feature as the first type feature.
[0062] Optionally, the first extraction unit is configured to:
[0063] Among the plurality of protein nodes, first-order reference protein nodes and first-order non-reference protein nodes are identified;
[0064] The first type information of the first-order reference protein node is determined as the first-order type feature of the first-order reference protein node;
[0065] Based on the first connection information, graph convolution is performed on the first type information of the first-order reference protein node to obtain the first-order type transformation feature. The first-order type transformation feature represents the degree of influence of the type of the first-order reference protein node on the type of the first-order non-reference protein node.
[0066] Based on the first-order type transformation feature, the first-order type information of the first-order non-reference protein node is transformed to obtain the first-order type feature of the first-order non-reference protein node.
[0067] Optionally, the first extraction unit is configured to:
[0068] Second-order reference protein nodes and second-order non-reference protein nodes are identified among the plurality of protein nodes;
[0069] The first-order type feature of the second-order reference protein node is determined as the second-order type feature of the second-order reference protein node;
[0070] Based on the first-order connectivity features, a linear transformation is performed on the first-order type features of the second-order reference protein node to obtain second-order type transformation features. The second-order type transformation features represent the degree of influence of the type of the second-order reference protein node on the type of the second-order non-reference protein node.
[0071] Based on the second-order type transformation features, the first-order type features of the second-order non-reference protein nodes are transformed to obtain the second-order type features of the second-order non-reference protein nodes.
[0072] Optionally, the first connectivity feature includes first-order connectivity features; the second extraction unit is used for:
[0073] Based on the first connection information and the first location information, extract the first-order location features corresponding to the first location information;
[0074] Based on the first-order connectivity feature and the first-order position feature, extract the second-order position feature corresponding to the first-order position feature until the current order reaches the target order, and then determine the currently obtained position feature as the first position feature.
[0075] Optionally, the second extraction unit is used for:
[0076] Among the plurality of protein nodes, first-order reference protein nodes and first-order non-reference protein nodes are identified;
[0077] The first position information of the first-order reference protein node is determined as the first-order position feature of the first-order reference protein node.
[0078] Based on the first connection information, graph convolution is performed on the first position information of the first-order reference protein node to obtain the first-order position transformation feature. The first-order position transformation feature represents the degree of influence of the position of the first-order reference protein node on the position of the first-order non-reference protein node.
[0079] Based on the first-order position transformation feature, the first position information of the first-order non-reference protein node is transformed to obtain the first-order position feature of the first-order non-reference protein node.
[0080] Optionally, the second extraction unit is used for:
[0081] Second-order reference protein nodes and second-order non-reference protein nodes are identified among the plurality of protein nodes;
[0082] The first-order position features of the second-order reference protein node are determined as the second-order position features of the second-order reference protein node.
[0083] Based on the first-order connectivity features, a linear transformation is performed on the first-order position features of the second-order reference protein node to obtain second-order position transformation features. The second-order position transformation features represent the degree of influence of the position of the second-order reference protein node on the position of the second-order non-reference protein node.
[0084] Based on the second-order position transformation features, the first-order position features of the second-order non-reference protein node are transformed to obtain the second-order position features of the second-order non-reference protein node.
[0085] Optionally, the first type feature includes first-order type features and second-order type features, and the first position feature includes first-order position features and second-order position features;
[0086] The third extraction unit is used for:
[0087] Based on the first-order type feature, the first-order position feature, and the first connection information, extract the first-order connection feature corresponding to the first connection information;
[0088] Based on the second-order type feature, the second-order position feature, and the first-order connection feature, extract the second-order connection feature corresponding to the first-order connection feature until the current order reaches the target order, and then determine the currently obtained connection feature as the first connection feature.
[0089] Optionally, the third extraction unit is used for:
[0090] Among the plurality of protein nodes, first-order reference protein nodes and first-order non-reference protein nodes are identified;
[0091] The first connection information of the first-order reference protein node is determined as the first-order connection feature of the first-order reference protein node.
[0092] After fusing the first-order type feature, the first-order position feature, and the first connection information of the first-order reference protein node, feature extraction is performed to obtain the first-order connection transformation feature, which represents the degree of influence of the connection status of the first-order reference protein node on the connection status of the first-order non-reference protein node.
[0093] Based on the first-order connectivity transformation features, the first connectivity information of the first-order non-reference protein node is transformed to obtain the first-order connectivity features of the first-order non-reference protein node.
[0094] Optionally, the feature decoding module is used for:
[0095] The first decoding unit is used to perform feature decoding on the first connection feature and the first type feature to obtain the second type information;
[0096] The second decoding unit is used to perform feature decoding on the first connection feature and the first position feature to obtain the second position information;
[0097] The third decoding unit is used to perform feature decoding on the first type feature, the first position feature and the first connection feature to obtain the second connection information;
[0098] The second type information, the second location information, and the second connection information constitute the drug molecule information.
[0099] Optionally, the feature encoding module is used to call a protein processing model to perform feature encoding on the protein molecule information to obtain the protein molecule features;
[0100] The feature decoding module is used to call the drug processing model to perform feature decoding on the protein molecule features to obtain the drug molecule information.
[0101] Optionally, the device further includes:
[0102] The sample acquisition module is used to acquire sample protein molecule information corresponding to sample protein molecules and sample drug molecule information corresponding to sample drug molecules.
[0103] The feature encoding module is also used to call the protein processing model to perform feature encoding on the sample protein molecule information to obtain sample protein molecule features; and to call the drug processing model to perform feature encoding on the sample drug molecule information to obtain sample drug molecule features.
[0104] The model training module is used to train the protein processing model and the drug processing model respectively based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, so as to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution.
[0105] Optionally, the device further includes:
[0106] The feature decoding module is also used to call the drug processing model to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information;
[0107] The discrimination module is used to call the first discrimination model to discriminate the predicted drug molecule information and obtain a first probability, wherein the first probability represents the probability that the predicted drug molecule information is the real drug molecule information;
[0108] The model training module is used for:
[0109] Based on the distribution information of the sample protein molecular features and the first probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first probability obtained by the first discrimination model is increased.
[0110] Based on the distribution information of the sample drug molecule features and the first probability, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the first probability obtained by the first discrimination model.
[0111] Optionally, the sample drug molecule is a real drug molecule; the discrimination module is further configured to call the first discrimination model to discriminate the sample drug molecule information and obtain a second probability, wherein the second probability represents the probability that the sample drug molecule information is real drug molecule information;
[0112] The model training module is further configured to train the first discrimination model based on the first probability and the second probability, so that the first probability obtained by the first discrimination model decreases and the second probability obtained increases.
[0113] Optionally, the device further includes:
[0114] The feature decoding module is also used to call the protein processing model to perform feature decoding on the sample drug molecule features to obtain predicted protein molecule information;
[0115] The discrimination module is used to call the second discrimination model to discriminate the predicted protein molecule information and obtain a third probability, wherein the third probability represents the probability that the predicted protein molecule information is the real protein molecule information;
[0116] The model training module is used for:
[0117] Based on the distribution information of the sample protein molecular features and the third probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the third probability obtained by the second discrimination model is increased.
[0118] Based on the distribution information of the sample drug molecule features and the second probability, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the third probability obtained by the second discrimination model.
[0119] Optionally, the sample protein molecule is a real protein molecule; the discrimination module is further configured to call the second discrimination model to discriminate the sample protein molecule information and obtain a fourth probability, wherein the fourth probability represents the probability that the sample protein molecule information is real protein molecule information;
[0120] The model training module is further configured to train the second discrimination model based on the third probability and the fourth probability, so that the third probability obtained by the second discrimination model decreases and the fourth probability obtained increases.
[0121] Optionally, the sample drug molecule is active against the sample protein molecule; the device further includes:
[0122] The feature decoding module is also used to call the drug processing model to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information;
[0123] The encoding module is used to call the encoding model to encode the sample drug molecule information and the predicted drug molecule information respectively, so as to obtain sample encoding features and predicted encoding features;
[0124] The model training module is used for:
[0125] Based on the distribution information of the sample protein molecular features, the sample coding features, and the predicted coding features, the protein processing model is trained to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and to increase the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0126] Based on the distribution information of the sample drug molecule features, the sample coding features, and the predicted coding features, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0127] Optionally, the model training module is further configured to train the coding model based on the sample coding features and the predicted coding features, so as to reduce the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0128] Optionally, the sample drug molecule is active against the sample protein molecule; the device further includes:
[0129] The feature decoding module is also used to call the drug processing model to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information;
[0130] The matching module is used to call the matching model and determine a first matching degree based on the sample protein molecule information and the predicted drug molecule information. The first matching degree represents the probability that the drug molecule indicated by the predicted drug molecule information is active against the sample protein molecule.
[0131] The model training module is used for:
[0132] Based on the distribution information of the sample protein molecular features and the first matching degree, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first matching degree obtained by the matching model is increased.
[0133] Based on the distribution information of the sample drug molecule features and the first matching degree, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the first matching degree obtained by the matching model.
[0134] Optionally, the matching module is further configured to invoke the matching model to determine a second matching degree based on the sample protein molecule information and the sample drug molecule information, wherein the second matching degree represents the probability that the sample drug molecule is active against the sample protein molecule;
[0135] The model training module is further configured to train the matching model based on the first matching degree and the second matching degree, so that the first matching degree obtained by the matching model decreases and the second matching degree obtained increases.
[0136] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to perform the operations performed by the drug molecule generation method as described above.
[0137] On the other hand, a computer-readable storage medium is provided that stores at least one computer program, which is loaded and executed by a processor to perform the operations performed by the drug molecule generation method as described above.
[0138] On the other hand, a computer program product is provided, comprising a computer program loaded and executed by a processor to perform the operations performed by the drug molecule generation method as described above.
[0139] The solution provided in this application takes into account that when a drug molecule is active against a protein molecule, the type and structure of protein nodes in the protein molecule are related to the type and structure of drug nodes in the drug molecule. Therefore, from the perspective of the protein molecule, based on the protein molecule information that can represent the type and structure of protein nodes, protein molecule features are extracted, and based on these protein molecule features, the drug molecule information corresponding to the drug molecule that is active against the protein molecule is directly decoded, thereby obtaining an effective drug molecule. There is no need to screen for drug molecules that are active against the protein molecule, which simplifies the process of generating drug molecules and improves the efficiency of generating drug molecules. Attached Figure Description
[0140] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0141] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0142] Figure 2 This is a flowchart of a drug molecule generation method provided in an embodiment of this application;
[0143] Figure 3 This is a flowchart of another drug molecule generation method provided in the embodiments of this application;
[0144] Figure 4 This is a schematic diagram of a feature extraction method provided in an embodiment of this application;
[0145] Figure 5 This is a flowchart of another drug molecule generation method provided in the embodiments of this application;
[0146] Figure 6 This is a flowchart illustrating a training method for a protein processing model and a drug processing model provided in an embodiment of this application;
[0147] Figure 7 This is an architecture diagram of a model training method provided in an embodiment of this application;
[0148] Figure 8 This is a schematic diagram of the structure of a drug molecule generation device provided in an embodiment of this application;
[0149] Figure 9 This is a schematic diagram of another drug molecule generation device provided in the embodiments of this application;
[0150] Figure 10 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0151] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0152] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0153] It is understood that the terms "first," "second," etc., used in this application may be used to describe various concepts herein, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, first type information may be referred to as second type information, and similarly, second type information may be referred to as first type information.
[0154] "At least one" refers to one or more protein nodes. For example, at least one protein node can be one protein node, two protein nodes, three protein nodes, or any integer number of protein nodes greater than or equal to one. "Multiple" refers to two or more protein nodes. For example, multiple protein nodes can be two protein nodes, three protein nodes, or any integer number of protein nodes greater than or equal to two. "Each" refers to each of the at least one protein node. For example, each protein node refers to each of the multiple protein nodes. If the multiple protein nodes are three protein nodes, then each protein node refers to each of the three protein nodes.
[0155] It is understood that in the embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0156] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0157] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.
[0158] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0159] The following will describe the drug molecule generation method provided in the embodiments of this application based on artificial intelligence and machine learning technologies.
[0160] The drug molecule generation method provided in this application embodiment can be used in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these.
[0161] In one possible implementation, the computer program involved in the embodiments of this application may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.
[0162] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. See also... Figure 1The implementation environment includes a terminal 101 and a server 102. The terminal 101 and server 102 are connected via a wireless or wired network. Optionally, server 102 is used to train a protein processing model and a drug processing model. The protein processing model is used to extract protein molecular features and generate protein molecular information, and the drug processing model is used to extract drug molecular features and generate drug molecular information. Server 102 sends the trained protein processing model and drug processing model to terminal 101, which can then use the protein processing model to extract the protein molecular features corresponding to the protein molecular information and use the drug processing model to generate the corresponding drug molecular information based on the protein molecular features.
[0163] It should be noted that, Figure 1 The example described here is based on server 102 training a protein processing model and a drug processing model and sending them to terminal 101. In another embodiment, the server can also call the protein processing model to extract the protein molecule features corresponding to the protein molecule information, and call the drug processing model to generate the corresponding drug molecule information based on the protein molecule features, and then send the drug molecule information to terminal 101.
[0164] Figure 2 This is a flowchart of a drug molecule generation method provided in an embodiment of this application. This embodiment is executed by a computer device. See also... Figure 2 The method includes:
[0165] 201. Computer equipment determines the protein molecule information corresponding to the protein molecule.
[0166] A protein molecule is a compound composed of multiple protein nodes. Protein molecule information is used to represent the type and structure of these protein nodes. A protein molecule is a compound formed by the folding of amino acids in space. Transforming a protein molecule into a graph structure, where the amino acids constituting the protein molecule serve as protein nodes, is a valid approach.
[0167] In one possible implementation, the protein molecule information includes first-type information of multiple protein nodes. This first-type information can represent the type of the protein node; for example, the first-type information of the protein node includes the molecular weight and isoelectric point of the protein node.
[0168] In one possible implementation, the protein molecule information further includes first position information of multiple protein nodes, which indicates the position of the protein node within the protein molecule. For example, the first position information of a protein node includes the "C" in the protein node. α "Atomic coordinates, each amino acid includes a unique 'C'." α "atom.
[0169] In one possible implementation, the protein molecule information also includes first connection information for multiple protein nodes, which indicates the connection status between the multiple protein nodes. Optionally, since protein nodes in a protein molecule are amino acids, and a protein molecule is formed by the folding of multiple amino acids in space, if the distance between two protein nodes is greater than a distance threshold, it is considered that there is no connection between the two protein nodes; if the distance between two protein nodes is not greater than the distance threshold, it is considered that there is a connection between the two protein nodes. The distance between two protein nodes can be determined by the "C" symbol in the two protein nodes. α "The distance between atoms is represented."
[0170] 202. Computer equipment performs feature encoding on protein molecule information to obtain protein molecule features.
[0171] After acquiring protein molecule information, the computer device performs feature encoding on the protein molecule information to extract the protein molecule features corresponding to the protein molecule information. Since the protein molecule information can represent the types and structures of multiple protein nodes, the protein molecule features refer to the characteristics of the types and structures of multiple protein nodes in the protein molecule.
[0172] 203. Computer equipment decodes the characteristics of protein molecules to obtain drug molecule information, and the drug molecule information indicates that the drug molecule is active against the protein molecule.
[0173] Considering that when a drug molecule is active against a protein molecule, the type and structure of protein nodes in the protein molecule are related to the type and structure of drug nodes in the drug molecule, after the computer device acquires the protein molecule features, it performs feature decoding on the protein molecule features to obtain the drug molecule information corresponding to the drug molecule that is active against the protein molecule.
[0174] The drug molecule has a graph structure with multiple drug nodes, where the atoms that make up the drug molecule are used as drug nodes.
[0175] In one possible implementation, the drug molecule information includes second-type information of multiple drug nodes, which represents the type of the drug node. For example, the second-type information is a feature vector obtained by one-hot encoding the types of atoms in the drug molecule. In another possible implementation, the drug molecule information also includes second-position information of multiple drug nodes, which represents the position of the drug node in the drug molecule. For example, the second-position information includes the coordinates of the drug node in the drug molecule. In yet another possible implementation, the drug molecule information also includes second-connection information of multiple drug nodes, which represents the connections between the multiple drug nodes. Optionally, since a drug molecule contains multiple atoms and drug nodes are atoms, a connection is considered to exist between two drug nodes if a chemical bond exists between them, and no connection is considered to exist between two drug nodes if no chemical bond exists. Optionally, the second-connection information between multiple drug nodes can be defined in different dimensions depending on the type of chemical bond.
[0176] The method provided in this application takes into account that when a drug molecule is active against a protein molecule, the type and structure of protein nodes in the protein molecule are related to the type and structure of drug nodes in the drug molecule. Therefore, from the perspective of the protein molecule, based on the protein molecule information that can represent the type and structure of protein nodes, protein molecule features are extracted, and based on these protein molecule features, the drug molecule information corresponding to the drug molecule that is active against the protein molecule is directly decoded, thereby obtaining an effective drug molecule. There is no need to screen for drug molecules that are active against the protein molecule, which simplifies the process of generating drug molecules and improves the efficiency of generating drug molecules.
[0177] In the above Figure 2 Based on the illustrated embodiment, the protein molecule information includes first type information, first location information, and first connection information of multiple protein nodes, and the drug molecule information includes second type information, second location information, and second connection information of multiple drug nodes. The process by which the computer device generates drug molecule information based on the protein molecule information is described below. Figure 3 The example shown.
[0178] Figure 3 This is a flowchart of another drug molecule generation method provided in this application embodiment. This application embodiment is executed by a computer device. See also... Figure 3 The method includes:
[0179] 301. Computer equipment determines the protein molecule information corresponding to the protein molecule.
[0180] The protein molecule comprises multiple protein nodes. Protein molecule information is used to represent the type and structure of these nodes, including first type information, first location information, and first connection information. Step 301 is similar to step 201 described above and will not be repeated here.
[0181] 302. The computer device extracts first-type features based on first-type information and first-connection information.
[0182] Since the first connection information can represent the connection between protein nodes, when extracting the first type feature corresponding to the first type information, the first connection information is also combined, so that the extracted first type feature not only includes the feature of the type to which each protein node belongs, but also considers the influence of the connection between different protein nodes on the type to which each protein node belongs, thereby improving the accuracy of the first type feature.
[0183] In one possible implementation, the first connection feature corresponding to the first connection information includes a first-order connection feature, the extraction method of which is detailed in step 304 below. The process of the computer device extracting the first type of feature includes the following steps (1) and (2):
[0184] (1) The computer device extracts the first-order type features corresponding to the first type information based on the first connection information and the first type information.
[0185] Optionally, the computer device identifies first-order reference protein nodes and first-order non-reference protein nodes among multiple protein nodes. The first-type information of the first-order reference protein node is defined as its first-order type feature. Based on the first connection information, a graph convolution is performed on the first-type information of the first-order reference protein node to obtain a first-order type transformation feature. This first-order type transformation feature represents the degree of influence of the type of the first-order reference protein node on the type of the first-order non-reference protein node. Based on this first-order type transformation feature, the computer device transforms the first-type information of the first-order non-reference protein node to obtain its first-order type feature.
[0186] Optionally, the first-order type transformation feature includes a first-order type scaling parameter and a first-order type displacement parameter. The computer device scales the first-order type information of the first-order non-reference protein node based on the first-order type scaling parameter, and then translates the scaled feature based on the first-order type displacement parameter to obtain the first-order type feature of the first-order non-reference protein node.
[0187] Optionally, the computer device uses the following formula to determine the first-order type features of the first-order reference protein node and the first-order type features of the first-order non-reference protein node.
[0188]
[0189]
[0190] in, X1 represents the first-order type feature of the first-order reference protein node. X2 represents the first-order type feature of a first-order non-reference protein node, A represents the first-order connection information, and S represents the first-order type information of a first-order non-reference protein node. G (X1|A) represents the first-order type scaling parameter obtained by performing graph convolution on the first type information based on the first connection information, T G (X1|A) represents the first-order type displacement parameter obtained by graph convolution of the first type information based on the first connection information. Where S G (·) and T G (·) represents graph convolution operation, Sigmoid(·) represents activation function, and ⊙ represents XOR operation.
[0191] (2) The computer device extracts the second-order type features corresponding to the first-order type features based on the first-order connection features and the first-order type features, until the current order reaches the target order, and then determines the currently obtained type features as the first type features.
[0192] After acquiring the first-order type features of each protein node, the computer device extracts the corresponding second-order type features based on the first-order connectivity features and the first-order type features. If the current order reaches the target order, the last acquired type feature is designated as the first-order type feature. If the current order has not yet reached the target order, the extraction of the next-order type features continues. When the target order is greater than two, the extraction method for third-order and higher type features is the same as that for second-order type features.
[0193] Optionally, the computer device identifies second-order reference protein nodes and second-order non-reference protein nodes among multiple protein nodes, and determines the first-order type feature of the second-order reference protein node as its second-order type feature. Based on the first-order connectivity feature, a linear transformation is performed on the first-order type feature of the second-order reference protein node to obtain a second-order type transformation feature, which represents the degree of influence of the type of the second-order reference protein node on the type of the second-order non-reference protein node. Based on the second-order type transformation feature, the computer device transforms the first-order type feature of the second-order non-reference protein node to obtain its second-order type feature. Optionally, the reference protein nodes of different orders are different, and each protein node among the multiple protein nodes serves as a first-order reference protein node.
[0194] Optionally, the second-order type transformation feature includes a second-order type scaling parameter and a second-order type displacement parameter. The computer device scales the first-order type feature of the second-order non-reference protein node based on the second-order type scaling parameter, and then translates the scaled feature based on the second-order type displacement parameter to obtain the second-order type feature of the second-order non-reference protein node.
[0195] Optionally, the computer device uses the following formula to determine the second-order type features of the second-order reference protein node and the second-order type features of the second-order non-reference protein node.
[0196]
[0197]
[0198] Where i represents the order, and i ≥ 2. Taking i = 2 as an example, then... This indicates the second-order type feature of the second-order reference protein node. This represents the first-order type feature of the second-order reference protein node. This represents the second-order type features of the second-order non-reference protein node. This represents the first-order type feature of a second-order non-reference protein node. This represents the second-order type scaling parameter obtained by performing a linear transformation on the first-order type feature based on the first-order connectivity feature. This represents the second-order type translation parameter obtained by performing a linear transformation on the first-order type feature based on the first-order connectivity feature. Where S... L (·) and T L (·) represents a linear transformation operation, Sigmoid(·) represents an activation function, and ⊙ represents the XOR operation.
[0199] Optionally, the above S L (·) and T L (·) is expressed by the following formula:
[0200]
[0201]
[0202] in, This represents the first-order type feature of the second-order reference protein node. The first-order connection characteristic is represented by FC(·), which represents a fully connected operation, and Linear(·) which represents a linear transformation operation. L (·) and T L The structures of (·) are the same, but the weight parameters are different.
[0203] 303. The computer device extracts the first location feature based on the first location information and the first connection information.
[0204] Since the first connection information can represent the connection between protein nodes, when extracting the first position feature corresponding to the first position information, the first connection information is also combined, so that the extracted first position feature not only includes the feature of the position to which each protein node belongs, but also considers the influence of the connection between different protein nodes on the position to which each protein node belongs, thereby improving the accuracy of the first position feature.
[0205] In one possible implementation, the first connection feature corresponding to the first connection information includes a first-order connection feature, the extraction method of which is detailed in step 304 below. The process of the computer device extracting the first location feature includes the following steps (1) and (2):
[0206] (1) The computer device extracts the first-order position features corresponding to the first position information based on the first connection information and the first position information.
[0207] Optionally, the computer device identifies first-order reference protein nodes and first-order non-reference protein nodes among multiple protein nodes. The first position information of the first-order reference protein node is defined as its first-order position feature. Based on the first connectivity information, a graph convolution is performed on the first position information of the first-order reference protein node to obtain a first-order position transformation feature. This first-order position transformation feature represents the degree of influence of the position of the first-order reference protein node on the position of the first-order non-reference protein node. Based on the first-order position transformation feature, the computer device transforms the first position information of the first-order non-reference protein node to obtain its first-order position feature.
[0208] Optionally, the first-order position transformation feature includes a first-order position scaling parameter and a first-order position displacement parameter. The computer device scales the first position information of the first-order non-reference protein node based on the first-order position scaling parameter, and then translates the scaled feature based on the first-order position displacement parameter to obtain the first-order position feature of the first-order non-reference protein node.
[0209] Optionally, the computer device uses the following formula to determine the first-order position features of the first-order reference protein node and the first-order position features of the first-order non-reference protein node.
[0210]
[0211]
[0212] in, P1 represents the first-order positional feature of the first-order reference protein node. P2 represents the first-order positional feature of a first-order non-reference protein node, A represents the first-order connection information, and S represents the first-order connection information. G (P1|A) represents the first-order position scaling parameter obtained by graph convolution of the first position information based on the first connection information, T G (P1|A) represents the first-order position displacement parameter obtained by graph convolution of the first position information based on the first connection information. Where S G (·) and T G (·) represents graph convolution operation, Sigmoid(·) represents activation function, and ⊙ represents XOR operation.
[0213] (2) The computer device extracts the second-order position features corresponding to the first-order position features based on the first-order connection features and the first-order position features, until the current order reaches the target order, and the currently obtained position features are determined as the first position features.
[0214] After acquiring the first-order positional features of each protein node, the computer device extracts the corresponding second-order positional features based on the first-order connectivity features and the first-order positional features. If the current order reaches the target order, the last acquired positional feature is designated as the first positional feature. If the current order has not yet reached the target order, the extraction of the next order positional feature continues. When the target order is greater than two, the extraction method for third-order and higher positional features is the same as that for second-order positional features.
[0215] Optionally, the computer device identifies second-order reference protein nodes and second-order non-reference protein nodes among multiple protein nodes, and determines the first-order positional features of the second-order reference protein nodes as their second-order positional features. Based on the first-order connectivity features, a linear transformation is performed on the first-order positional features of the second-order reference protein nodes to obtain second-order positional transformation features. These second-order positional transformation features represent the degree of influence of the position of the second-order reference protein nodes on the position of the second-order non-reference protein nodes. Based on the second-order positional transformation features, the computer device transforms the first-order positional features of the second-order non-reference protein nodes to obtain their second-order positional features. Optionally, the reference protein nodes of different orders are different, and each protein node among the multiple protein nodes serves as a first-order reference protein node.
[0216] Optionally, the second-order position transformation feature includes a second-order position scaling parameter and a second-order position displacement parameter. The computer device scales the first-order position feature of the second-order non-reference protein node based on the second-order position scaling parameter, and then translates the scaled feature based on the second-order position displacement parameter to obtain the second-order position feature of the second-order non-reference protein node.
[0217] Optionally, the computer device uses the following formula to determine the second-order position features of the second-order reference protein node and the second-order position features of the second-order non-reference protein node.
[0218]
[0219]
[0220] Where i represents the order, and i ≥ 2. Taking i = 2 as an example, then... This represents the second-order positional features of the second-order reference protein node. This represents the first-order positional features of the second-order reference protein node. This represents the second-order positional features of a second-order non-reference protein node. This represents the first-order positional features of a second-order non-reference protein node. This represents the second-order position scaling parameter obtained by performing a linear transformation on the first-order position features based on the first-order connectivity features. This represents the second-order position translation parameter obtained by performing a linear transformation on the first-order position features based on the first-order connectivity features. Where S... L (·) and T L (·) represents a linear transformation operation, Sigmoid(·) represents an activation function, and ⊙ represents the XOR operation.
[0221] 304. The computer device extracts the first connection feature based on the first connection information, the first type information, and the first location information.
[0222] In this embodiment of the application, when extracting the first connection feature corresponding to the first connection information, the first type information and the first position information of the protein node are also combined, so that the extracted first connection feature not only includes the feature of the position to which each protein node belongs, but also considers the type and position of each protein node, thereby improving the accuracy of the first position feature.
[0223] In one possible implementation, the first type feature includes first-order type features and second-order type features, the extraction methods of which are detailed in step 302 above. The first position feature includes first-order position features and second-order position features, the extraction methods of which are detailed in step 303 above. The process of the computer device extracting the first connection feature includes the following steps (1) and (2):
[0224] (1) The computer device extracts the first-order connection features corresponding to the first connection information based on the first-order type features, the first-order position features and the first connection information.
[0225] Optionally, the computer device identifies first-order reference protein nodes and first-order non-reference protein nodes among multiple protein nodes, and determines the first connection information of the first-order reference protein nodes as their first-order connection features. After fusing the first-order type features, first-order position features, and the first connection information of the first-order reference protein nodes, feature extraction is performed to obtain first-order connection transformation features. These features represent the degree to which the connection status of the first-order reference protein nodes affects the connection status of the first-order non-reference protein nodes. Based on these first-order connection transformation features, the computer device transforms the first connection information of the first-order non-reference protein nodes to obtain their first-order connection features.
[0226] Optionally, the first-order connectivity transformation feature includes a first-order connectivity scaling parameter and a first-order connectivity displacement parameter. The computer device scales the first connectivity information of the first-order non-reference protein node based on the first-order connectivity scaling parameter, and then translates the scaled feature based on the first-order connectivity displacement parameter to obtain the first-order connectivity feature of the first-order non-reference protein node.
[0227] (2) The computer device extracts the second-order connection features corresponding to the first-order connection features based on the second-order type features, second-order position features and first-order connection features, until the current order reaches the target order, and the currently obtained connection features are determined as the first connection features.
[0228] After acquiring the first-order connectivity features, the computer device extracts second-order connectivity features based on the second-order type features, second-order position features, and first-order connectivity features. If the current order reaches the target order, the last acquired connectivity feature is designated as the first connectivity feature. If the current order has not yet reached the target order, the next order connectivity features are extracted. The extraction method for the type features of each order is the same.
[0229] Optionally, the computer device identifies second-order reference protein nodes and second-order non-reference protein nodes among multiple protein nodes, and determines the first-order connectivity features of the second-order reference protein nodes as their second-order connectivity features. The second-order type features, second-order position features, and the first-order connectivity features of the second-order reference protein nodes are fused for feature extraction to obtain second-order connectivity transformation features. These transformation features represent the degree to which the connectivity of the second-order reference protein nodes affects the connectivity of the second-order non-reference protein nodes. Based on these transformation features, the computer device transforms the first-order connectivity features of the second-order non-reference protein nodes to obtain their second-order connectivity features. Optionally, the reference protein nodes of different orders are different, and each protein node among the multiple protein nodes serves as a first-order reference protein node.
[0230] Optionally, the second-order connectivity transformation features include second-order connectivity scaling parameters and second-order connectivity displacement parameters. The computer device scales the first-order connectivity features of the second-order non-reference protein node based on the second-order connectivity scaling parameters, and then translates the scaled features based on the second-order connectivity displacement parameters to obtain the second-order connectivity features of the second-order non-reference protein node.
[0231] Optionally, the computer device uses the following formula to determine the type characteristics of second-order reference protein nodes and second-order non-reference protein nodes.
[0232]
[0233]
[0234] Where i represents the order, and i ≥ 2. Taking i = 2 as an example, This represents the second-order connection characteristics of the second-order reference protein node. This represents the first-order connectivity features of the second-order reference protein node. This represents the second-order connection features of the second-order non-reference protein nodes. This represents the first-order connectivity features of second-order non-reference protein nodes. This represents the second-order connectivity scaling parameter obtained by fusing second-order type features, second-order position features, and first-order connectivity features before feature extraction. This represents the second-order connectivity translation parameter obtained by fusing second-order type features, second-order position features, and first-order connectivity features and then extracting the feature. Where S... C (·) and T C (·) represents feature extraction operation, Sigmoid(·) represents activation function, and ⊙ represents XOR operation.
[0235] Optionally, the above S C (·) and T C(·) is expressed by the following formula:
[0236]
[0237]
[0238] in, This represents the first-order type feature of the second-order reference protein node. Representing second-order type features, Let S represent second-order location features, FC(·) denote fully connected operation, and CNN(·) denote convolution operation. Optionally, S C (·) and T C The structures of (·) are the same, but the weight parameters are different.
[0239] Figure 4 This is a schematic diagram of a feature extraction method provided in an embodiment of this application, as shown below. Figure 4 As shown, P represents the first location information of the protein node, A represents the first connection information of the protein node, and X represents the first type information of the protein node. Feature extraction needs to be performed on each of these three types of information. If each type of information is extracted independently, the single type of information will limit the feature extraction process, leading to a mismatch between the second location information, second connection information, and second type information of the drug node generated based on the extracted features. Therefore, this embodiment of the application performs joint feature extraction on the three types of information, such as... Figure 4 As shown, l represents the target order, and in the extraction of multi-order position features of the first position information P. At that time, the first connection information A and the corresponding multi-level connection features were referenced respectively. Extracting multi-level type features of the first type of information X At that time, the first connection information A and the corresponding multi-level connection features were referenced respectively. Extracting the first connection information A and the corresponding multi-level connection features At the same time, the multi-level position features of the first position information P were referenced respectively. And the multi-level type features of the first type of information X
[0240] 305. The computer device performs feature decoding on the first connection feature and the first type feature to obtain the second type information.
[0241] After the computer device acquires the first connection feature and the first type feature, since the first type feature is extracted based on the first connection feature, the first connection feature is also combined when decoding the first type feature, thereby decoding the accurate second type information.
[0242] In this step 305, the feature decoding process is the inverse of the feature encoding process in step 302. However, the weight parameters used in the feature encoding process are different from those used in the feature decoding process. Therefore, the second type of information obtained by decoding is the type information corresponding to the drug molecule, rather than the type information corresponding to the protein molecule.
[0243] 306. The computer device performs feature decoding on the first connection feature and the first location feature to obtain the second location information.
[0244] After the computer device acquires the first connection feature and the first location feature, since the first location feature is extracted based on the first connection feature, the first connection feature is also combined when decoding the first location feature, thereby decoding the accurate second location information.
[0245] In this step 306, the feature decoding process is the inverse of the feature encoding process in step 303. However, the weight parameters used in the feature encoding process are different from those used in the feature decoding process. Therefore, the second position information obtained by decoding is the position information corresponding to the drug molecule, rather than the position information corresponding to the protein molecule.
[0246] 307. The computer device performs feature decoding on the first type feature, the first location feature, and the first connection feature to obtain the second connection information.
[0247] After the computer device acquires the first type feature, the first location feature, and the first connection feature, since the first connection feature is extracted based on the first type feature and the first location feature, when decoding the first connection feature, it also combines the first type feature and the first location feature to decode the accurate second connection information.
[0248] In this step 307, the feature decoding process is the inverse of the feature encoding process in step 304. However, the weight parameters used in the feature encoding process are different from those used in the feature decoding process. Therefore, the second connection information obtained by decoding is the connection information corresponding to the drug molecule, rather than the connection information corresponding to the protein molecule.
[0249] In this embodiment of the application, by performing the above steps 305-307, second type information, second position information and second connection information are obtained. The second type information is the type information of the drug node in the drug molecule, the second position information is the position information of the drug node in the drug molecule, and the second connection information is the connection information of the drug node in the drug molecule. The second type information, the second position information and the second connection information constitute drug molecule information, and the drug molecule information indicates that the drug molecule is active against the protein molecule.
[0250] The method provided in this application takes into account that when a drug molecule is active against a protein molecule, the type and structure of protein nodes in the protein molecule are related to the type and structure of drug nodes in the drug molecule. Therefore, from the perspective of the protein molecule, based on the protein molecule information that can represent the type and structure of protein nodes, protein molecule features are extracted, and based on these protein molecule features, the drug molecule information corresponding to the drug molecule that is active against the protein molecule is directly decoded, thereby obtaining an effective drug molecule. There is no need to screen for drug molecules that are active against the protein molecule, which simplifies the process of generating drug molecules and improves the efficiency of generating drug molecules.
[0251] In the above Figure 2 or Figure 3 Based on the illustrated embodiments, the computer device can also implement drug molecule generation methods by calling protein processing models and drug processing models, as detailed below. Figure 5 The example shown.
[0252] Figure 5 This is a flowchart of another drug molecule generation method provided in this application embodiment. This application embodiment is executed by a computer device. See also... Figure 5 The method includes:
[0253] 501. Computer equipment determines the protein molecule information corresponding to the protein molecule.
[0254] The process of step 501 is the same as that of step 201 above, and will not be repeated here.
[0255] 502. The computer equipment calls the protein processing model to encode the protein molecule information and obtain the protein molecule features.
[0256] After acquiring protein molecule information, the computer equipment inputs this information into a protein processing model to obtain the protein molecule's characteristics.
[0257] In one possible implementation, the protein molecule information includes first type information, first location information, and first connectivity information. The protein processing model includes a type information processing network, a location information processing network, and a connectivity information processing network. A computer device invokes the type information processing network to extract first type features based on the first type information and the first connectivity information; invokes the location information processing network to extract first location features based on the first location information and the first connectivity information; and invokes the connectivity information processing network to extract first connectivity features based on the first connectivity information, the first type information, and the first location information.
[0258] The process of extracting protein molecular features based on the protein processing model is the same as steps 302-304 above, and will not be repeated here.
[0259] It should be noted that in some embodiments, the protein processing model is a reversible flow model, meaning that the protein processing model can extract protein molecular features based on protein molecular information, and can also generate protein molecular information based on protein molecular features. Taking the extraction of protein molecular features as the forward process and the generation of protein molecular information as the reverse process as an example, the forward process of this protein processing model is denoted as G. r (·), the reverse process is denoted as The protein treatment model then satisfies the following formula:
[0260]
[0261] Where, r i G represents protein molecule information. r (r i The ) indicates that the protein processing model extracts protein molecular features based on protein molecular information. This indicates that the protein processing model regenerates protein molecule information based on the extracted protein molecule features.
[0262] 503. The computer equipment calls the drug processing model to perform feature decoding on the protein molecule characteristics and obtain drug molecule information.
[0263] After the computer device acquires the protein molecule characteristics, it inputs these characteristics into the drug treatment model to obtain the drug molecule information. The drug molecule information indicates that the drug molecule is active against the protein molecule.
[0264] In one possible implementation, the protein molecule features include a first type feature, a first positional feature, and a first connectivity feature, while the drug molecule information includes second type information, second positional information, and second connectivity information. The drug processing model includes a type information processing network, a positional information processing network, and a connectivity information processing network. A computer device invokes the type information processing network to decode the first connectivity feature and the first type feature to obtain the second type information; invokes the positional information processing network to decode the first connectivity feature and the first positional feature to obtain the second positional information; and invokes the connectivity information processing network to decode the first type feature, the first positional feature, and the first connectivity feature to obtain the second connectivity information.
[0265] The process of generating drug molecule information based on the drug processing model is the same as steps 305-307 above, and will not be repeated here.
[0266] It should be noted that in some embodiments, the drug processing model is a reversible flow model, meaning that the drug processing model can extract drug molecule features based on drug molecule information, and can also generate drug molecule information based on drug molecule features. Taking the extraction of drug molecule features as the forward process and the generation of drug molecule information as the reverse process as an example, the forward process of this drug processing model is denoted as G. l (·), the reverse process is denoted as The drug treatment model then satisfies the following formula:
[0267]
[0268] Among them, l i G represents drug molecule information. l (l i This indicates that the drug processing model extracts drug molecule features based on drug molecule information. This indicates that the drug processing model regenerates drug molecule information based on the extracted drug molecule features.
[0269] In the process of generating drug molecule information from protein molecule information, a protein processing model extracts protein molecule features based on the protein molecule information, and then inputs these features into the drug processing model to generate drug molecule information. This process can be represented by the following formula:
[0270]
[0271]
[0272] Where, r i Represents protein molecule information, Indicates the characteristics of protein molecules, This indicates the generated drug molecule information.
[0273] In the process of generating protein molecule information from drug molecule information, a drug processing model extracts drug molecule features based on the drug molecule information, and then inputs these features into the protein processing model to generate protein molecule information. This process can be represented by the following formula:
[0274]
[0275]
[0276] Among them, l i Indicates drug molecule information, Indicates the molecular characteristics of the drug. This indicates information about the generated protein molecules.
[0277] The training processes for the protein treatment model and the drug treatment model are detailed below. Figure 6 The embodiments shown are not described here.
[0278] The method provided in this application takes into account that when a drug molecule is active against a protein molecule, the type and structure of protein nodes in the protein molecule are related to the type and structure of drug nodes in the drug molecule. Therefore, from the perspective of the protein molecule, based on the protein molecule information that can represent the type and structure of protein nodes, protein molecule features are extracted, and based on these protein molecule features, the drug molecule information corresponding to the drug molecule that is active against the protein molecule is directly decoded, thereby obtaining an effective drug molecule. There is no need to screen for drug molecules that are active against the protein molecule, which simplifies the process of generating drug molecules and improves the efficiency of generating drug molecules.
[0279] Figure 6 This is a flowchart illustrating a training method for a protein processing model and a drug processing model provided in this application embodiment. This application embodiment is executed by a computer device. See [link / reference]. Figure 6 The method includes:
[0280] 601. Computer equipment acquires sample protein molecule information corresponding to sample protein molecules and sample drug molecule information corresponding to sample drug molecules.
[0281] The sample protein molecule information is the same as that in the above embodiments, and the sample drug molecule information is the same as that in the above embodiments.
[0282] In one possible implementation, the sample protein molecule is a real protein molecule, and the sample drug molecule is a real drug molecule. The sample drug molecule is active against the sample protein molecule.
[0283] 602. The computer equipment calls the protein processing model to encode the feature information of the sample protein molecules and obtain the sample protein molecule features.
[0284] The process of step 602 is the same as that of step 502 above, and will not be repeated here.
[0285] 603. The computer equipment calls the drug processing model to perform feature encoding on the sample drug molecule information to obtain the sample drug molecule features.
[0286] The process of step 603 is the same as that of step 502 above, and will not be repeated here.
[0287] 604. The computer equipment trains a protein processing model and a drug processing model based on the distribution information of sample protein molecular features and the distribution information of sample drug molecular features, respectively, so as to increase the similarity between the probability distribution of sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and increase the similarity between the probability distribution of sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution.
[0288] Since the protein processing model and drug processing model in this embodiment are reversible flow models, the closer the probability distribution of the sample protein molecular features extracted by the protein processing model is to the standard Gaussian distribution, the higher the accuracy of the protein processing model. Similarly, the closer the probability distribution of the sample drug molecular features extracted by the drug processing model is to the standard Gaussian distribution, the higher the accuracy of the drug processing model. Therefore, by increasing the similarity between the probability distribution of the sample protein molecular features and the standard Gaussian distribution, the protein processing model can be trained to improve its accuracy; similarly, by increasing the similarity between the probability distribution of the sample drug molecular features and the standard Gaussian distribution, the drug processing model can be trained to improve its accuracy.
[0289] It should be noted that the number of sample protein molecules in steps 601-603 above is multiple, and the number of sample protein molecule features extracted by the protein processing model is also multiple. By statistically analyzing these multiple sample protein molecule features, the probability distribution of the sample protein molecule features output by the protein processing model is obtained. Similarly, the number of sample drug molecules in steps 601-603 above is multiple, and the number of sample drug molecule features extracted by the drug processing model is also multiple. By statistically analyzing these multiple sample drug molecule features, the probability distribution of the sample drug molecule features output by the drug processing model is obtained.
[0290] Optionally, the computer equipment uses the following formula to train the protein processing model.
[0291]
[0292] Among them, G r This represents a protein processing model. P represents maximization. i P represents the probability distribution of the molecular characteristics of the sample protein. m KL(·) represents the standard Gaussian distribution, KL(·) represents the KL divergence (Kullback-Leibler divergence), and KL(P) represents the standard Gaussian distribution. i ||P m This refers to the similarity between the probability distribution of the sample protein molecular characteristics and the standard Gaussian distribution. This represents the determinant of the Jacobian matrix corresponding to the protein processing model. It is a penalty term used to ensure that the integral of the probability distribution of the sample protein molecular features is 1.
[0293] Optionally, the computer equipment uses the following formula to train the drug treatment model.
[0294]
[0295] Among them, G l This represents a drug treatment model. P represents maximization. l P represents the probability distribution of the molecular characteristics of the sample drug. m KL(·) represents the standard Gaussian distribution, KL(·) represents the KL divergence, and KL(P) represents the standard Gaussian distribution. l ||P m This refers to the similarity between the probability distribution of the molecular characteristics of the sample drug and the standard Gaussian distribution. This represents the determinant of the Jacobian matrix corresponding to the drug treatment model. It is a penalty term used to ensure that the integral of the probability distribution of the sample drug molecule features is 1.
[0296] In some embodiments, in addition to the training methods described above, the computer device can also simultaneously train the protein processing model and the drug processing model using adversarial training. The embodiments of this application provide the following four adversarial training methods.
[0297] The first type of adversarial training:
[0298] (1) The computer device calls the drug processing model to perform feature decoding on the sample protein molecule features to obtain the predicted drug molecule information. It then calls the first discrimination model to discriminate the predicted drug molecule information and obtain the first probability. The first probability represents the probability that the predicted drug molecule information is the real drug molecule information.
[0299] (2) The computer equipment trains a protein processing model based on the distribution information of the sample protein molecular features and the first probability, so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution increases, and the first probability obtained by the first discrimination model increases.
[0300] Since the more closely the predicted drug molecule information generated by the drug processing model is to the real drug molecule information, the stronger the protein processing model's ability to extract protein molecule features, computer equipment also trains the protein processing model by increasing the probability that the predicted drug molecule information is the real drug molecule information, thereby improving the accuracy of the protein processing model.
[0301] (3) The computer equipment trains a drug processing model based on the distribution information of the sample drug molecule features and the first probability, so that the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution increases, and the first probability obtained by the first discrimination model increases.
[0302] Since the more closely the predicted drug molecule information generated by the drug processing model is to the real drug molecule information, the stronger the ability of the drug processing model to generate drug molecule information, computer equipment also trains the drug processing model by increasing the probability that the predicted drug molecule information is the real drug molecule information, thereby improving the accuracy of the drug processing model.
[0303] In one possible implementation, the computer device determines a first loss parameter based on the first probability, the first loss parameter being negatively correlated with the first probability, and the computer device further trains a protein processing model and a drug processing model based on the first loss parameter to reduce the obtained first loss parameter.
[0304] Optionally, the computer device uses the following formula to determine the first loss parameter.
[0305]
[0306] in, Indicates the first loss parameter. Let D represent the first probability. r (·) represents the first discriminant model. Indicates the characteristics of protein molecules, This indicates information about the predicted drug molecules.
[0307] In one possible implementation, the sample drug molecule is a real drug molecule. The computer device invokes a first discriminant model to discriminate the sample drug molecule information and obtains a second probability. The second probability represents the probability that the sample drug molecule information is real drug molecule information. Based on the first and second probabilities, the first discriminant model is trained so that the first probability obtained by the first discriminant model decreases and the second probability obtained increases.
[0308] Since the predicted drug molecule information is model-generated molecular information, while the sample drug molecule information is real drug molecule information, the smaller the first probability predicted by the first discriminant model and the larger the predicted second probability, the more accurate the first discriminant model will be. Therefore, by decreasing the first probability and increasing the second probability, the first discriminant model can be trained to improve its accuracy.
[0309] In one possible implementation, the computer device determines a second loss parameter based on the first probability and the second probability, the second loss parameter being positively correlated with the first probability and negatively correlated with the second probability, and the computer device also trains a first discriminant model based on the second loss parameter to reduce the obtained second loss parameter.
[0310] Optionally, the computer device uses the following formula to determine the second loss parameter.
[0311]
[0312] in, This represents the second loss parameter. Let D represent the first probability. r (r i ) represents the second probability, r i This indicates the molecular information of the sample drug.
[0313] The second type of adversarial training:
[0314] (1) The computer device calls the protein processing model to perform feature decoding on the sample drug molecule features to obtain the predicted protein molecule information. The second discrimination model is then called to discriminate the predicted protein molecule information to obtain the third probability. The third probability represents the probability that the predicted protein molecule information is the real protein molecule information.
[0315] (2) The computer equipment trains the protein processing model based on the distribution information of the sample protein molecular features and the third probability, so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution increases, and the third probability obtained by the second discrimination model increases.
[0316] Since the more closely the predicted protein molecules generated by the protein processing model are to the actual protein molecules, the stronger the ability of the protein processing model to generate protein molecules, computer equipment also trains the protein processing model by increasing the probability that the predicted protein molecules are the actual protein molecules, thereby improving the accuracy of the protein processing model.
[0317] (3) The computer equipment trains the drug processing model based on the distribution information and second probability of the sample drug molecule features, so that the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution increases, and the third probability obtained by the second discrimination model increases.
[0318] Since the more closely the predicted protein molecule information generated by the protein processing model is to the real protein molecule information, the stronger the drug processing model's ability to extract drug molecule features, computer equipment also trains the drug processing model by increasing the probability that the predicted protein molecule information is the real protein molecule information, thereby improving the accuracy of the drug processing model.
[0319] In one possible implementation, the computer device determines a third loss parameter based on the third probability, the third loss parameter being negatively correlated with the third probability, and the computer device also trains a protein treatment model and a drug treatment model based on the third loss parameter to reduce the obtained third loss parameter.
[0320] Optionally, the computer device uses the following formula to determine the third loss parameter.
[0321]
[0322] in, This represents the third loss parameter. D represents the third probability. l (·) represents the second discriminant model. Indicates the molecular characteristics of the drug. This indicates information about the predicted protein molecule.
[0323] In one possible implementation, the sample protein molecule is a real protein molecule. The computer device calls a second discriminant model to discriminate the sample protein molecule information and obtain a fourth probability. The fourth probability represents the probability that the sample protein molecule information is real protein molecule information. Based on the third and fourth probabilities, the second discriminant model is trained so that the third probability obtained by the second discriminant model decreases and the fourth probability increases.
[0324] Since the predicted protein molecule information is model-generated molecular information, while the sample protein molecule information is real protein molecule information, the smaller the predicted third probability and the larger the predicted fourth probability of the second discriminant model, the more accurate the second discriminant model will be. Therefore, by decreasing the third probability and increasing the fourth probability, the second discriminant model can be trained to improve its accuracy.
[0325] In one possible implementation, the computer device determines a fourth loss parameter based on the third and fourth probabilities, the fourth loss parameter being positively correlated with the third probability and negatively correlated with the fourth probability. The computer device also trains a second discriminant model based on the fourth loss parameter to reduce the obtained fourth loss parameter.
[0326] Optionally, the computer device uses the following formula to determine the fourth loss parameter.
[0327]
[0328] in, This represents the fourth loss parameter. D represents the third probability. l (l i ) represents the fourth probability, l i This indicates information about the protein molecules in the sample.
[0329] The third type of adversarial training involves drug molecules that are active against protein molecules.
[0330] (1) The computer device calls the drug processing model to decode the features of the sample protein molecules to obtain the predicted drug molecule information. The coding model is then called to encode the sample drug molecule information and the predicted drug molecule information to obtain the sample coding features and the predicted coding features.
[0331] (2) The computer equipment trains the protein processing model based on the distribution information of the sample protein molecular features, the sample coding features and the predicted coding features, so as to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and increase the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0332] Since the more closely the predicted coding features of the drug molecule information are to the sample coding features of the sample drug molecule information, the stronger the ability of the protein processing model to extract protein molecule features, computer equipment also trains the protein processing model by increasing the similarity between the sample coding features and the predicted coding features, thereby improving the accuracy of the protein processing model.
[0333] (3) The computer equipment trains the drug processing model based on the distribution information of sample drug molecule features, sample coding features and predicted coding features, so as to increase the similarity between the probability distribution of sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and increase the similarity between the sample coding features and predicted coding features obtained by the coding model.
[0334] Since the more closely the predicted coding features of drug molecule information are to the sample coding features of the sample drug molecule information, the stronger the ability of the drug processing model to generate drug molecule information, computer equipment also trains the drug processing model by increasing the similarity between the sample coding features and the predicted coding features, thereby improving the accuracy of the drug processing model.
[0335] In one possible implementation, the computer device determines a fifth loss parameter based on the sample coding features and the predicted coding features, the fifth loss parameter being negatively correlated with the similarity between the sample coding features and the predicted coding features. The computer device also trains a protein processing model and a drug processing model based on the fifth loss parameter to reduce the obtained fifth loss parameter.
[0336] Optionally, the computer device uses the following formula to determine the fifth loss parameter.
[0337]
[0338] in, This represents the fifth loss parameter, d(·) represents the Euclidean distance, and E l (l i ) represents the sample coding feature. This represents the predictive coding feature.
[0339] In one possible implementation, the computer device trains a coding model based on sample coding features and predicted coding features to reduce the similarity between the sample coding features and the predicted coding features obtained by the coding model. Since the lower the similarity between the sample coding features and the predicted coding features, the higher the accuracy of the coding model, the accuracy of the coding model is improved by reducing the similarity between the sample coding features and the predicted coding features during training.
[0340] In one possible implementation, the computer device determines a sixth loss parameter based on the sample coding features and the predicted coding features. The sixth loss parameter is positively correlated with the similarity between the sample coding features and the predicted coding features. The computer device also trains a coding model based on the sixth loss parameter to reduce the obtained sixth loss parameter.
[0341] Optionally, the computer device uses the following formula to determine the sixth loss parameter.
[0342]
[0343] in, The sixth loss parameter is represented by d(·), which represents the Euclidean distance. l (l i ) represents the sample coding feature. This represents the predictive coding feature.
[0344] The fourth type of adversarial training involves sample drug molecules that are active against sample protein molecules.
[0345] (1) The computer device calls the drug processing model to perform feature decoding on the sample protein molecule features to obtain the predicted drug molecule information. It then calls the matching model to determine the first matching degree based on the sample protein molecule information and the predicted drug molecule information. The first matching degree represents the probability that the drug molecule indicated by the predicted drug molecule information is active against the sample protein molecule.
[0346] Among them, the matching of drug molecules and protein molecules means that the drug molecules are active against the protein molecules.
[0347] (2) The computer equipment trains the protein processing model based on the distribution information of the sample protein molecular features and the first matching degree, so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution increases, and the first matching degree obtained by the matching model increases.
[0348] Since the greater the probability that the drug molecule indicated by the predicted drug molecule information is active on the sample protein molecule, the stronger the ability of the protein processing model to extract protein molecule features, the computer equipment also trains the protein processing model by increasing the probability that the drug molecule indicated by the predicted drug molecule information is active on the sample protein molecule, thereby improving the accuracy of the protein processing model.
[0349] (3) The computer equipment trains the drug processing model based on the distribution information of the sample drug molecule features and the first matching degree, so that the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution increases, and the first matching degree obtained by the matching model increases.
[0350] Since the greater the probability that the drug molecule indicated by the predicted drug molecule information is active on the sample protein molecule, the stronger the ability of the drug processing model to generate drug molecule information, the computer equipment also trains the drug processing model by increasing the probability that the drug molecule indicated by the predicted drug molecule information is active on the sample protein molecule, thereby improving the accuracy of the drug processing model.
[0351] In one possible implementation, the computer device invokes a matching model to determine a second matching degree based on the sample protein molecule information and the sample drug molecule information. The second matching degree represents the probability that the sample drug molecule is active against the sample protein molecule. The computer device trains the matching model based on the first and second matching degrees, such that the first matching degree obtained by the matching model decreases while the second matching degree increases. Since a smaller first matching degree and a larger second matching degree result in higher accuracy of the matching model, training the matching model by decreasing the obtained first matching degree and increasing the obtained second matching degree improves the accuracy of the matching model.
[0352] This application provides a supervised learning method for generating drug molecule information from protein molecule information. This method enables protein processing models and drug processing models to learn the characteristics of protein molecule information and drug molecule information, thereby realizing the generation of drug molecule information from protein molecule information and simplifying the process of generating drug molecules.
[0353] Furthermore, the embodiments of this application combine reversible flow models and adversarial training to learn the features of protein molecule information and drug molecule information, and incorporate joint training between various types of information in the protein processing model and the drug processing model to enhance the interaction between information within the model.
[0354] Figure 7 This is an architecture diagram of a model training method provided in an embodiment of this application, such as... Figure 7 As shown, the system architecture includes a protein processing model, a drug processing model, a first discriminant model, a second discriminant model, a coding model, and a matching model.
[0355] The protein processing model and drug processing model are used to generate drug molecule information from protein molecule information, or vice versa. In the process of generating drug molecule information from protein molecule information, the protein processing model extracts protein molecule features based on the protein molecule information, and then inputs these features into the drug processing model to generate drug molecule information. Conversely, in the process of generating protein molecule information from drug molecule information, the drug processing model extracts drug molecule features based on the drug molecule information, and then inputs these features into the protein processing model to generate protein molecule information.
[0356] The system comprises three main components: a first discriminant model to determine whether drug molecule information is genuine, a second discriminant model to determine whether protein molecule information is genuine, an encoding model to encode the drug molecule information, and a matching model to determine whether the drug molecule indicated by the drug molecule information is active against the protein molecule indicated by the protein molecule information. These discriminant, encoding, and matching models are used in a recurring adversarial training process against the protein processing and drug processing models. Through this adversarial training, the protein molecule information generated by the protein processing model and the drug molecule information generated by the drug processing model can be made more realistic.
[0357] Figure 8 This is a schematic diagram of a drug molecule generation device provided in an embodiment of this application. See also... Figure 8 The device includes:
[0358] The information determination module 801 is used to determine the protein molecule information corresponding to the protein molecule. The protein molecule includes multiple protein nodes, and the protein molecule information is used to represent the type and structure of the multiple protein nodes.
[0359] The feature encoding module 802 is used to encode the protein molecule information to obtain the protein molecule features;
[0360] The feature decoding module 803 is used to perform feature decoding on the protein molecule to obtain drug molecule information, which indicates that the drug molecule is active on the protein molecule.
[0361] The drug molecule generation apparatus provided in this application takes into account that when a drug molecule is active against a protein molecule, the type and structure of protein nodes in the protein molecule are related to the type and structure of drug nodes in the drug molecule. Therefore, from the perspective of protein molecules, based on the protein molecule information that can represent the type and structure of protein nodes, protein molecule features are extracted, and based on these protein molecule features, the drug molecule information corresponding to the drug molecule that is active against the protein molecule is directly decoded, thereby obtaining effective drug molecules. There is no need to screen for drug molecules that are active against the protein molecule, which simplifies the process of generating drug molecules and improves the efficiency of generating drug molecules.
[0362] Optionally, see Figure 9 The protein molecule information includes first type information, first location information, and first connection information of the plurality of protein nodes. The first connection information indicates the connection status between the plurality of protein nodes. The protein molecule features include first type features corresponding to the first type information, first location features corresponding to the first location information, and first connection features corresponding to the first connection information. The feature encoding module 802 includes:
[0363] The first extraction unit 812 is used to extract the first type feature based on the first type information and the first connection information;
[0364] The second extraction unit 822 is used to extract the first location feature based on the first location information and the first connection information;
[0365] The third extraction unit 832 is used to extract the first connection feature based on the first connection information, the first type information and the first location information.
[0366] Optionally, see Figure 9 The first connectivity feature includes a first-order connectivity feature; the first extraction unit 812 is used for:
[0367] Based on the first connection information and the first type information, extract the first-order type feature corresponding to the first type information;
[0368] Based on the first-order connectivity feature and the first-order type feature, extract the second-order type feature corresponding to the first-order type feature until the current order reaches the target order, and then determine the currently obtained type feature as the first type feature.
[0369] Optionally, see Figure 9 The first extraction unit 812 is used for:
[0370] Identify first-order reference protein nodes and first-order non-reference protein nodes among these multiple protein nodes;
[0371] The first type information of the first-order reference protein node is determined as the first-order type feature of the first-order reference protein node;
[0372] Based on the first connection information, graph convolution is performed on the first type information of the first-order reference protein node to obtain the first-order type transformation feature. The first-order type transformation feature represents the degree of influence of the type of the first-order reference protein node on the type of the first-order non-reference protein node.
[0373] Based on this first-order type transformation feature, the first-order type information of the first-order non-reference protein node is transformed to obtain the first-order type feature of the first-order non-reference protein node.
[0374] Optionally, see Figure 9 The first extraction unit 812 is used for:
[0375] Identify second-order reference protein nodes and second-order non-reference protein nodes among these multiple protein nodes;
[0376] The first-order type feature of the second-order reference protein node is determined as the second-order type feature of the second-order reference protein node;
[0377] Based on this first-order connectivity feature, a linear transformation is performed on the first-order type feature of the second-order reference protein node to obtain the second-order type transformation feature. This second-order type transformation feature represents the degree of influence of the type of the second-order reference protein node on the type of the second-order non-reference protein node.
[0378] Based on this second-order type transformation feature, the first-order type feature of the second-order non-reference protein node is transformed to obtain the second-order type feature of the second-order non-reference protein node.
[0379] Optionally, see Figure 9 The first connectivity feature includes first-order connectivity features; the second extraction unit 822 is used for:
[0380] Based on the first connection information and the first location information, the first-order location feature corresponding to the first location information is extracted;
[0381] Based on the first-order connectivity feature and the first-order position feature, extract the second-order position feature corresponding to the first-order position feature until the current order reaches the target order, and then determine the currently obtained position feature as the first position feature.
[0382] Optionally, see Figure 9 The second extraction unit 822 is used for:
[0383] Identify first-order reference protein nodes and first-order non-reference protein nodes among these multiple protein nodes;
[0384] The first position information of the first-order reference protein node is determined as the first-order position feature of the first-order reference protein node.
[0385] Based on the first connection information, graph convolution is performed on the first position information of the first-order reference protein node to obtain the first-order position transformation feature. The first-order position transformation feature represents the degree of influence of the position of the first-order reference protein node on the position of the first-order non-reference protein node.
[0386] Based on this first-order position transformation feature, the first position information of the first-order non-reference protein node is transformed to obtain the first-order position feature of the first-order non-reference protein node.
[0387] Optionally, see Figure 9 The second extraction unit 822 is used for:
[0388] Identify second-order reference protein nodes and second-order non-reference protein nodes among these multiple protein nodes;
[0389] The first-order position feature of the second-order reference protein node is determined as the second-order position feature of the second-order reference protein node.
[0390] Based on this first-order connectivity feature, a linear transformation is performed on the first-order position feature of the second-order reference protein node to obtain the second-order position transformation feature. This second-order position transformation feature represents the degree of influence of the position of the second-order reference protein node on the position of the second-order non-reference protein node.
[0391] Based on this second-order position transformation feature, the first-order position feature of the second-order non-reference protein node is transformed to obtain the second-order position feature of the second-order non-reference protein node.
[0392] Optionally, see Figure 9 The first type feature includes first-order type features and second-order type features, and the first position feature includes first-order position features and second-order position features;
[0393] The third extraction unit 832 is used for:
[0394] Based on the first-order type feature, the first-order position feature, and the first connection information, extract the first-order connection feature corresponding to the first connection information;
[0395] Based on the second-order type feature, the second-order position feature, and the first-order connection feature, extract the second-order connection feature corresponding to the first-order connection feature until the current order reaches the target order, and then determine the currently obtained connection feature as the first connection feature.
[0396] Optionally, see Figure 9 The third extraction unit 832 is used for:
[0397] Identify first-order reference protein nodes and first-order non-reference protein nodes among these multiple protein nodes;
[0398] The first connection information of the first-order reference protein node is determined as the first-order connection feature of the first-order reference protein node.
[0399] After fusing the first-order type feature, the first-order position feature, and the first connection information of the first-order reference protein node, feature extraction is performed to obtain the first-order connection transformation feature. The first-order connection transformation feature represents the degree of influence of the connection status of the first-order reference protein node on the connection status of the first-order non-reference protein node.
[0400] Based on this first-order connectivity transformation feature, the first connectivity information of the first-order non-reference protein node is transformed to obtain the first-order connectivity feature of the first-order non-reference protein node.
[0401] Optionally, see Figure 9 The feature decoding module 803 is used for:
[0402] The first decoding unit 813 is used to perform feature decoding on the first connection feature and the first type feature to obtain second type information;
[0403] The second decoding unit 823 is used to perform feature decoding on the first connection feature and the first position feature to obtain the second position information;
[0404] The third decoding unit 833 is used to perform feature decoding on the first type feature, the first position feature and the first connection feature to obtain the second connection information;
[0405] The second type of information, the second location information, and the second connection information constitute the drug molecule information.
[0406] Optionally, see Figure 9 The feature encoding module 802 is used to call the protein processing model to perform feature encoding on the protein molecule information to obtain the protein molecule features;
[0407] The feature decoding module 803 is used to call the drug processing model to perform feature decoding on the protein molecule features and obtain the drug molecule information.
[0408] Optionally, see Figure 9 The device also includes:
[0409] The sample acquisition module 804 is used to acquire sample protein molecule information corresponding to sample protein molecules and sample drug molecule information corresponding to sample drug molecules.
[0410] The feature encoding module 802 is also used to call the protein processing model to perform feature encoding on the sample protein molecule information to obtain sample protein molecule features; and to call the drug processing model to perform feature encoding on the sample drug molecule information to obtain sample drug molecule features.
[0411] The model training module 805 is used to train the protein processing model and the drug processing model based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, respectively, so as to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution.
[0412] Optionally, see Figure 9 The device also includes:
[0413] The feature decoding module 803 is also used to call the drug processing model to perform feature decoding on the protein molecule features of the sample to obtain the predicted drug molecule information;
[0414] The discrimination module 806 is used to call the first discrimination model to discriminate the predicted drug molecule information and obtain a first probability, which represents the probability that the predicted drug molecule information is the real drug molecule information.
[0415] The model training module 805 is used for:
[0416] Based on the distribution information of the sample protein molecular features and the first probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first probability obtained by the first discrimination model is increased.
[0417] Based on the distribution information of the sample drug molecule features and the first probability, the drug processing model is trained so that the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution is increased, and the first probability obtained by the first discriminant model is increased.
[0418] Optionally, see Figure 9 The sample drug molecule is a real drug molecule; the discrimination module 806 is also used to call the first discrimination model to discriminate the sample drug molecule information and obtain a second probability, which represents the probability that the sample drug molecule information is real drug molecule information.
[0419] The model training module 805 is also used to train the first discriminant model based on the first probability and the second probability, so that the first probability obtained by the first discriminant model decreases and the second probability obtained increases.
[0420] Optionally, see Figure 9 The device also includes:
[0421] The feature decoding module 803 is also used to call the protein processing model to perform feature decoding on the drug molecule features of the sample and obtain predicted protein molecule information;
[0422] The discrimination module 806 is used to call the second discrimination model to discriminate the predicted protein molecule information and obtain a third probability, which represents the probability that the predicted protein molecule information is the real protein molecule information.
[0423] The model training module 805 is used for:
[0424] Based on the distribution information of the sample protein molecular features and the third probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the third probability obtained by the second discriminant model is increased.
[0425] Based on the distribution information of the sample drug molecule features and the second probability, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the third probability obtained by the second discrimination model.
[0426] Optionally, see Figure 9 The sample protein molecule is a real protein molecule; the discrimination module 806 is also used to call the second discrimination model to discriminate the sample protein molecule information and obtain a fourth probability, which represents the probability that the sample protein molecule information is real protein molecule information.
[0427] The model training module 805 is also used to train the second discriminant model based on the third probability and the fourth probability, so that the third probability obtained by the second discriminant model decreases and the fourth probability obtained increases.
[0428] Optionally, see Figure 9 The sample drug molecule is active against the sample protein molecule; the device also includes:
[0429] The feature decoding module 803 is also used to call the drug processing model to perform feature decoding on the protein molecule features of the sample to obtain the predicted drug molecule information;
[0430] The encoding module 807 is used to call the encoding model to encode the sample drug molecule information and the predicted drug molecule information respectively, so as to obtain the sample encoding features and the predicted encoding features;
[0431] The model training module 805 is used for:
[0432] Based on the distribution information of the protein molecular features of the sample, the sample coding features, and the predicted coding features, the protein processing model is trained to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and to increase the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0433] Based on the distribution information of the drug molecule features of the sample, the sample coding features, and the predicted coding features, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0434] Optionally, see Figure 9 The model training module 805 is also used to train the coding model based on the sample coding features and the predicted coding features, so as to reduce the similarity between the sample coding features and the predicted coding features obtained by the coding model.
[0435] Optionally, see Figure 9 The sample drug molecule is active against the sample protein molecule; the device also includes:
[0436] The feature decoding module 803 is also used to call the drug processing model to perform feature decoding on the protein molecule features of the sample to obtain the predicted drug molecule information;
[0437] The matching module 808 is used to call the matching model and determine a first matching degree based on the sample protein molecule information and the predicted drug molecule information. The first matching degree represents the probability that the drug molecule indicated by the predicted drug molecule information is active in the sample protein molecule.
[0438] The model training module 805 is used for:
[0439] Based on the distribution information of the sample protein molecular features and the first matching degree, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first matching degree obtained by the matching model is increased.
[0440] Based on the distribution information of the sample drug molecule features and the first matching degree, the drug processing model is trained so that the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution is increased, and the first matching degree obtained by the matching model is increased.
[0441] Optionally, see Figure 9 The matching module 808 is also used to call the matching model to determine a second matching degree based on the sample protein molecule information and the sample drug molecule information. The second matching degree represents the probability that the sample drug molecule is active against the sample protein molecule.
[0442] The model training module 805 is also used to train the matching model based on the first matching degree and the second matching degree, so that the first matching degree obtained by the matching model decreases and the second matching degree increases.
[0443] It should be noted that the drug molecule generation apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the drug molecule generation apparatus and the drug molecule generation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0444] This application also provides a computer device, which includes a processor and a memory. The memory stores at least one computer program, which is loaded and executed by the processor to perform the operations performed in the drug molecule generation method of the above embodiments.
[0445] Optionally, the computer device is provided as a terminal. Figure 10 A schematic diagram of the structure of a terminal 1000 provided in an exemplary embodiment of this application is shown.
[0446] The terminal 1000 includes a processor 1001 and a memory 1002.
[0447] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0448] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one computer program, which is used by the processor 1001 to implement the drug molecule generation method provided in the method embodiments of this application.
[0449] In some embodiments, the terminal 1000 may also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 1004 and a display screen 1005.
[0450] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0451] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks.
[0452] Display screen 1005 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, disposed on the front panel of terminal 1000; in other embodiments, there may be at least two display screens 1005, disposed on different surfaces of terminal 1000 or in a foldable design.
[0453] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on terminal 1000 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0454] Optionally, the computer device is provided as a server. Figure 11This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1100 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1101 and one or more memories 1102. The memories 1102 store at least one computer program, which is loaded and executed by the processor 1101 to implement the methods provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.
[0455] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to perform the operations of the drug molecule generation method described above.
[0456] This application also provides a computer program product, including a computer program loaded and executed by a processor to perform the operations performed by the drug molecule generation method of the above embodiments.
[0457] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0458] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present application should be included within the protection scope of the present application.
Claims
1. A method for generating drug molecules, characterized in that, The method includes: Determine the protein molecule information corresponding to the protein molecule, wherein the protein molecule includes multiple protein nodes, and the protein molecule information includes first type information, first position information, and first connection information, wherein the first connection information indicates the connection between the multiple protein nodes; Based on the first type information and the first connection information, extract first-order type features, extract second-order type features based on the first-order connection features and the first-order type features, until the current order reaches the target order, and determine the currently obtained type features as the first type features. Based on the first location information and the first connection information, extract the first location feature; Based on the first connection information, the first type information, and the first location information, a first connection feature is extracted, wherein the first connection feature includes the first-order connection feature; wherein the protein molecule feature corresponding to the protein molecule information includes the first type feature, the first location feature, and the first connection feature; The protein molecule features are decoded to obtain drug molecule information, and the drug molecule information indicates that the drug molecule is active against the protein molecule.
2. The method according to claim 1, characterized in that, The step of extracting first-order type features based on the first type information and the first connection information includes: Among the plurality of protein nodes, first-order reference protein nodes and first-order non-reference protein nodes are identified; The first type information of the first-order reference protein node is determined as the first-order type feature of the first-order reference protein node; Based on the first connection information, graph convolution is performed on the first type information of the first-order reference protein node to obtain the first-order type transformation feature. The first-order type transformation feature represents the degree of influence of the type of the first-order reference protein node on the type of the first-order non-reference protein node. Based on the first-order type transformation feature, the first-order type information of the first-order non-reference protein node is transformed to obtain the first-order type feature of the first-order non-reference protein node.
3. The method according to claim 1, characterized in that, The step of extracting the first location feature based on the first location information and the first connection information includes: Based on the first connection information and the first location information, extract the first-order location features corresponding to the first location information; Based on the first-order connectivity feature and the first-order position feature, extract the second-order position feature corresponding to the first-order position feature until the current order reaches the target order, and then determine the currently obtained position feature as the first position feature.
4. The method according to claim 1, characterized in that, The first type feature includes first-order type features and second-order type features, and the first position feature includes first-order position features and second-order position features; The step of extracting the first connection feature based on the first connection information, the first type information, and the first location information includes: Based on the first-order type feature, the first-order position feature, and the first connection information, extract the first-order connection feature corresponding to the first connection information; Based on the second-order type feature, the second-order position feature, and the first-order connection feature, extract the second-order connection feature corresponding to the first-order connection feature until the current order reaches the target order, and then determine the currently obtained connection feature as the first connection feature.
5. The method according to any one of claims 1-4, characterized in that, The protein molecule features are obtained by calling a protein processing model and performing feature encoding on the protein molecule information; The process of decoding the protein molecule features to obtain drug molecule information includes: The drug processing model is invoked to perform feature decoding on the protein molecule features to obtain the drug molecule information.
6. The method according to claim 5, characterized in that, The training process of the protein treatment model and the drug treatment model includes: Obtain the sample protein molecule information corresponding to the sample protein molecule and the sample drug molecule information corresponding to the sample drug molecule; The protein processing model is invoked to encode the sample protein molecule information to obtain sample protein molecule features; the drug processing model is invoked to encode the sample drug molecule information to obtain sample drug molecule features. Based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, the protein processing model and the drug processing model are trained respectively, so as to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution.
7. The method according to claim 6, characterized in that, The method further includes: The drug processing model is invoked to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information; The first discrimination model is invoked to discriminate the predicted drug molecule information and obtain a first probability, whereby the first probability represents the probability that the predicted drug molecule information is the actual drug molecule information. The protein processing model and the drug processing model are trained based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, respectively, to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and to increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution, including: Based on the distribution information of the sample protein molecular features and the first probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first probability obtained by the first discrimination model is increased. Based on the distribution information of the sample drug molecule features and the first probability, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the first probability obtained by the first discrimination model.
8. The method according to claim 6, characterized in that, The method further includes: The protein processing model is invoked to perform feature decoding on the sample drug molecule features to obtain predicted protein molecule information; The second discrimination model is invoked to discriminate the predicted protein molecule information, and a third probability is obtained, which represents the probability that the predicted protein molecule information is the real protein molecule information. The protein processing model and the drug processing model are trained based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, respectively, to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and to increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution, including: Based on the distribution information of the sample protein molecular features and the third probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the third probability obtained by the second discrimination model is increased. Based on the distribution information of the sample drug molecule features and the third probability, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the third probability obtained by the second discrimination model.
9. The method according to claim 6, characterized in that, The sample drug molecule is active against the sample protein molecule; the method further includes: The drug processing model is invoked to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information; The matching model is invoked, and a first matching degree is determined based on the sample protein molecule information and the predicted drug molecule information. The first matching degree represents the probability that the drug molecule indicated by the predicted drug molecule information is active against the sample protein molecule. The protein processing model and the drug processing model are trained based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, respectively, to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and to increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution, including: Based on the distribution information of the sample protein molecular features and the first matching degree, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first matching degree obtained by the matching model is increased. Based on the distribution information of the sample drug molecule features and the first matching degree, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the first matching degree obtained by the matching model.
10. A drug molecule generation device, characterized in that, The device includes: An information determination module is used to determine the protein molecule information corresponding to a protein molecule. The protein molecule includes multiple protein nodes. The protein molecule information includes first type information, first position information, and first connection information. The first connection information indicates the connection between the multiple protein nodes. The feature encoding module is used to extract first-order type features based on the first type information and the first connection information, extract second-order type features based on the first-order connection features and the first-order type features, until the current order reaches the target order, and determine the currently obtained type features as the first type features; The feature encoding module is further configured to extract a first position feature based on the first position information and the first connection information; The feature encoding module is further configured to extract a first connection feature based on the first connection information, the first type information, and the first position information, wherein the first connection feature includes the first-order connection feature; wherein the protein molecule feature corresponding to the protein molecule information includes the first type feature, the first position feature, and the first connection feature; The feature decoding module is used to perform feature decoding on the protein molecule features to obtain drug molecule information, wherein the drug molecule information indicates that the drug molecule is active against the protein molecule.
11. The apparatus according to claim 10, characterized in that, The feature encoding module is used for: Among the plurality of protein nodes, first-order reference protein nodes and first-order non-reference protein nodes are identified; The first type information of the first-order reference protein node is determined as the first-order type feature of the first-order reference protein node; Based on the first connection information, graph convolution is performed on the first type information of the first-order reference protein node to obtain the first-order type transformation feature. The first-order type transformation feature represents the degree of influence of the type of the first-order reference protein node on the type of the first-order non-reference protein node. Based on the first-order type transformation feature, the first-order type information of the first-order non-reference protein node is transformed to obtain the first-order type feature of the first-order non-reference protein node.
12. The apparatus according to claim 10, characterized in that, The feature encoding module is used for: Based on the first connection information and the first location information, extract the first-order location features corresponding to the first location information; Based on the first-order connectivity feature and the first-order position feature, extract the second-order position feature corresponding to the first-order position feature until the current order reaches the target order, and then determine the currently obtained position feature as the first position feature.
13. The apparatus according to claim 10, characterized in that, The first type feature includes first-order type features and second-order type features, and the first position feature includes first-order position features and second-order position features; the feature encoding module is used for: Based on the first-order type feature, the first-order position feature, and the first connection information, extract the first-order connection feature corresponding to the first connection information; Based on the second-order type feature, the second-order position feature, and the first-order connection feature, extract the second-order connection feature corresponding to the first-order connection feature until the current order reaches the target order, and then determine the currently obtained connection feature as the first connection feature.
14. The apparatus according to any one of claims 10-13, characterized in that, The protein molecule features are obtained by calling a protein processing model and performing feature encoding on the protein molecule information; The feature decoding module is used for: The drug processing model is invoked to perform feature decoding on the protein molecule features to obtain the drug molecule information.
15. The apparatus according to claim 14, characterized in that, The device further includes: The sample acquisition module is used to acquire sample protein molecule information corresponding to sample protein molecules and sample drug molecule information corresponding to sample drug molecules. The feature encoding module is also used to call the protein processing model to perform feature encoding on the sample protein molecule information to obtain sample protein molecule features, and to call the drug processing model to perform feature encoding on the sample drug molecule information to obtain sample drug molecule features; The model training module is used to train the protein processing model and the drug processing model respectively based on the distribution information of the sample protein molecular features and the distribution information of the sample drug molecular features, so as to increase the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution, and increase the similarity between the probability distribution of the sample drug molecular features obtained by the drug processing model and the standard Gaussian distribution.
16. The apparatus according to claim 15, characterized in that, The device further includes: The feature decoding module is also used to call the drug processing model to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information; The discrimination module is used to call the first discrimination model to discriminate the predicted drug molecule information and obtain a first probability, wherein the first probability represents the probability that the predicted drug molecule information is the real drug molecule information; The model training module is used for: Based on the distribution information of the sample protein molecular features and the first probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first probability obtained by the first discrimination model is increased. Based on the distribution information of the sample drug molecule features and the first probability, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the first probability obtained by the first discrimination model.
17. The apparatus according to claim 15, characterized in that, The device further includes: The feature decoding module is also used to call the protein processing model to perform feature decoding on the sample drug molecule features to obtain predicted protein molecule information; The discrimination module is used to call the second discrimination model to discriminate the predicted protein molecule information and obtain a third probability, wherein the third probability represents the probability that the predicted protein molecule information is the real protein molecule information; The model training module is used for: Based on the distribution information of the sample protein molecular features and the third probability, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the third probability obtained by the second discrimination model is increased. Based on the distribution information of the sample drug molecule features and the third probability, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the third probability obtained by the second discrimination model.
18. The apparatus according to claim 15, characterized in that, The sample drug molecule is active against the sample protein molecule; the device further includes: The feature decoding module is also used to call the drug processing model to perform feature decoding on the sample protein molecule features to obtain predicted drug molecule information; The matching module is used to call the matching model and determine a first matching degree based on the sample protein molecule information and the predicted drug molecule information. The first matching degree represents the probability that the drug molecule indicated by the predicted drug molecule information is active against the sample protein molecule. The model training module is used for: Based on the distribution information of the sample protein molecular features and the first matching degree, the protein processing model is trained so that the similarity between the probability distribution of the sample protein molecular features obtained by the protein processing model and the standard Gaussian distribution is increased, and the first matching degree obtained by the matching model is increased. Based on the distribution information of the sample drug molecule features and the first matching degree, the drug processing model is trained to increase the similarity between the probability distribution of the sample drug molecule features obtained by the drug processing model and the standard Gaussian distribution, and to increase the first matching degree obtained by the matching model.
19. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to perform the operations performed by the drug molecule generation method as described in any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to perform the operations performed by the drug molecule generation method as described in any one of claims 1 to 9.
21. A computer program product, comprising a computer program, characterized in that, The computer program is loaded and executed by a processor to perform the operations performed by the drug molecule generation method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Medicine screening method and device and electronic equipment
CN111816252A
Protein-targeted drug compound identification
US20200392178A1