Molecular property prediction method, molecular property prediction model training method, device and equipment and computer readable storage medium
By acquiring geometric and non-geometric information of compounds for attention encoding and nonlinear mapping, the problem of not considering local information of chemical bonds in existing technologies is solved, thereby improving the accuracy of molecular property prediction.
Patent Information
- Application Number
- CN202410474081.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-18
- Publication Date
- 2025-10-24
AI Technical Summary
Existing molecular property prediction methods based on isovariant graph neural networks do not fully consider the local information brought by the chemical bonds of molecules, which affects the accuracy of molecular property prediction.
By obtaining the geometric and non-geometric information of the compound, dependency-based attention encoding is performed, and non-geometric splicing information is combined for nonlinear mapping to improve the accuracy of molecular property prediction.
By considering the local information provided by the chemical bonds of compounds, the accuracy of molecular property predictions is improved.
Smart Images

Figure CN120833864A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a molecular property prediction method, a molecular property prediction model training method and device, equipment and a computer readable storage medium. BACKGROUND
[0002] Molecular property prediction is a crucial part of artificial intelligence-assisted drug discovery. It helps drug researchers better understand the structure and properties of molecules, thereby designing and optimizing drug molecules with specific properties. In the related art, the molecular property prediction method based on the isometric graph neural network models the geometric transformations of the molecular structure, such as translation, rotation, and flipping. However, it does not further consider the local information brought by the chemical bonds of the molecule, which affects the prediction accuracy of the molecular properties. SUMMARY
[0003] The embodiments of the present application provide a molecular property prediction method, a molecular property prediction model training method, device, equipment and a computer readable storage medium, which can improve the accuracy of molecular property prediction.
[0004] The technical solutions of the embodiments of the present application are implemented as follows:
[0005] The embodiments of the present application provide a molecular property prediction method, which comprises:
[0006] Obtaining geometric information and non-geometric information of a compound;
[0007] Performing first encoding based on the geometric information and the non-geometric information to obtain first geometric encoding information and first non-geometric encoding information;
[0008] Performing attention encoding based on the dependency relationship based on the first geometric encoding information and the first non-geometric encoding information to obtain non-geometric attention encoding information;
[0009] Splicing the non-geometric attention encoding information and the non-geometric information to obtain non-geometric splicing information;
[0010] Performing non-linear mapping based on the non-geometric splicing information to obtain the molecular property of the compound.
[0011] The embodiments of the present application provide a molecular property prediction model training method, which comprises:
[0012] Obtaining a training data set, wherein the training data set comprises a plurality of molecular graph data, and the molecular graph data comprises geometric information and non-geometric information;
[0013] The noise adding module is configured to perform noise adding processing on each of the molecular graph data to obtain noise-added molecular graph data corresponding to each of the molecular graph data.
[0014] The data processing module is configured to perform first encoding based on the geometric information and the non-geometric information to obtain first geometric encoding information and first non-geometric encoding information.
[0015] The data processing module is configured to perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the noise-added molecular graph data to obtain non-geometric attention encoding information corresponding to each of the noise-added molecular graph data.
[0016] The data processing module is configured to perform a first training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the noise-added molecular graph data to obtain a first molecular property prediction model.
[0017] The data processing module is configured to perform a first training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the noise-added molecular graph data to obtain a first molecular property prediction model.
[0018] The data acquisition module is configured to acquire geometric information and non-geometric information of a compound.
[0019] The data processing module is configured to perform first encoding based on the geometric information and the non-geometric information to obtain first geometric encoding information and first non-geometric encoding information.
[0020] The data processing module is configured to perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the noise-added molecular graph data to obtain non-geometric attention encoding information corresponding to each of the noise-added molecular graph data.
[0021] The data processing module is configured to perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the noise-added molecular graph data to obtain non-geometric attention encoding information corresponding to each of the noise-added molecular graph data.
[0022] The property prediction module is configured to perform non-linear mapping based on the non-geometric splicing information to obtain a molecular property of the compound.
[0023] The data acquisition module is configured to acquire geometric information and non-geometric information of a compound.
[0024] The data acquisition module is configured to acquire geometric information and non-geometric information of a compound.
[0025] The noise adding module is configured to perform noise adding processing on each of the molecular graph data to obtain noise-added molecular graph data corresponding to each of the molecular graph data.
[0026] The data processing module is configured to perform first encoding on each of the noisy molecular graph data by the initialized molecular property prediction model to obtain first geometric encoding information and first non-geometric encoding information corresponding to each of the noisy molecular graph data.
[0027] The data processing module is further configured to perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the noisy molecular graph data by the molecular property prediction model to obtain non-geometric attention encoding information corresponding to each of the noisy molecular graph data.
[0028] The training module is configured to perform a first training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the noisy molecular graph data to obtain a first molecular property prediction model.
[0029] An electronic device is provided in an embodiment of the present application, and the electronic device includes:
[0030] A memory is configured to store computer executable instructions.
[0031] A processor is configured to execute the computer executable instructions stored in the memory to implement the prediction method of the molecular property or the training method of the molecular property prediction model provided in the embodiments of the present application.
[0032] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores a computer program or computer executable instructions, and is configured to implement the prediction method of the molecular property or the training method of the molecular property prediction model provided in the embodiments of the present application when executed by a processor.
[0033] A computer program product is provided in an embodiment of the present application, and the computer program product includes a computer program or computer executable instructions, and the computer program or computer executable instructions are configured to implement the prediction method of the molecular property or the training method of the molecular property prediction model provided in the embodiments of the present application when executed by a processor.
[0034] The embodiments of the present application have the following beneficial effects:
[0035] After the geometric information and non-geometric information of the compound are encoded, the first geometric encoding information and the first non-geometric encoding information are jointly fused into an attention mechanism based on a dependency relationship, the generation of the non-geometric attention encoding information is affected by the first geometric encoding information, so that the non-geometric attention encoding information includes the geometric information of the molecule, that is, the local information caused by the chemical bond of the molecule, the non-geometric attention encoding information is spliced with the non-geometric information to obtain non-geometric splicing information, further reducing the loss of information, and finally the molecular property of the compound is obtained through the non-geometric splicing information, so as to achieve the beneficial effect of improving the accuracy of the prediction of the molecular property. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 is a structural schematic diagram of a molecular property prediction system architecture provided by an embodiment of the present application;
[0037] Figure 2A is a first structural schematic diagram of a server provided by an embodiment of the present application;
[0038] Figure 2B is a second structural schematic diagram of a server provided by an embodiment of the present application;
[0039] Figure 3A is a first flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0040] Figure 3B is a second flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0041] Figure 3C is a third flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0042] Figure 3D is a fourth flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0043] Figure 3E is a fifth flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0044] Figure 3F is a sixth flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0045] Figure 3G is a seventh flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0046] Figure 3H is an eighth flow schematic diagram of a molecular property prediction method provided by an embodiment of the present application;
[0047] Figure 4AFIG. 1 is a first flowchart of a training method of a molecular property prediction model according to an embodiment of the present application;
[0048] Figure 4B FIG. 2 is a second flowchart of a training method of a molecular property prediction model according to an embodiment of the present application;
[0049] Figure 4C FIG. 3 is a third flowchart of a training method of a molecular property prediction model according to an embodiment of the present application;
[0050] Figure 4D FIG. 4 is a fourth flowchart of a training method of a molecular property prediction model according to an embodiment of the present application;
[0051] Figure 4E FIG. 5 is a fifth flowchart of a training method of a molecular property prediction model according to an embodiment of the present application;
[0052] Figure 4F FIG. 6 is a sixth flowchart of a training method of a molecular property prediction model according to an embodiment of the present application;
[0053] Figure 4G FIG. 7 is a seventh flowchart of a training method of a molecular property prediction model according to an embodiment of the present application;
[0054] Figure 5 FIG. 8 is a schematic diagram of a training architecture of a molecular property prediction model according to an embodiment of the present application;
[0055] Figure 6 FIG. 9 is a schematic diagram of a principle of an isometric graph transformer according to an embodiment of the present application;
[0056] Figure 7A FIG. 10 is a schematic diagram of a structure of a first training task for molecular property prediction according to an embodiment of the present application;
[0057] Figure 7B FIG. 11 is a schematic diagram of a principle of a first training task for molecular property prediction according to an embodiment of the present application;
[0058] Figure 7C FIG. 12 is a schematic diagram of a structure of a second training task for molecular property prediction according to an embodiment of the present application;
[0059] Figure 7D FIG. 13 is a schematic diagram of a graph structure of a compound according to an embodiment of the present application;
[0060] Figure 8 FIG. 14 is a flowchart of a drug discovery according to an embodiment of the present application.
[0061] It should be noted that the above "first", "second" are only used to distinguish different schemes, and do not represent the advantages or disadvantages of the schemes or the priority in the implementation process. DETAILED DESCRIPTION
[0062] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0063] In the following description, "some embodiments" are referred to, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0064] In the following description, the terms "first", "second", "third" are only used to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first", "second", "third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0065] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0066] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0067] The relevant data collection and processing in the embodiments of the present application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.
[0068] Before the embodiments of the present application are further described in detail, the terms and terms related to the embodiments of the present application are explained, and the terms and terms related to the embodiments of the present application are applicable to the following explanations.
[0069] 1) Deep Learning (DL), a new research direction in the field of Machine Learning (ML), is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal rules and representation levels of sample data. The information obtained in this learning process is very helpful for the interpretation of data such as text, images and sound. The ultimate goal of deep learning is to enable machines to have analysis and learning ability like humans, and to be able to recognize text, images and sound data.
[0070] 2) Self-Attention Mechanism, a method to capture the dependency between different positions in the input sequence in neural networks. Compared with traditional Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN), self-attention mechanism can more effectively handle long-distance dependencies. The core idea of self-attention mechanism is to generate a weighted representation for each element in the input sequence by calculating the mutual relationship between each element and other elements.
[0071] 3) Transformer, a deep learning model based on self-attention mechanism, mainly used for sequence-to-sequence natural language processing tasks such as machine translation, text summarization, etc. The Transformer model mainly consists of two parts: encoder and decoder, which are responsible for processing input sequences and generating output sequences respectively. Self-attention mechanism enables Transformer to capture long-distance dependencies in sequences without using recurrent neural networks or convolutional neural network structures. In addition, it has the advantages of parallel computing and capturing long-distance dependencies, and has become the mainstream method in the field of natural language processing.
[0072] 4) Pre-training, a common method in deep learning, involves training a model on a large-scale unlabeled dataset to learn general features or patterns of the data. The main purpose of pre-training is to improve the performance of the model, especially when labeled data is scarce. The process of pre-training a model usually includes two stages: pre-training stage and fine-tuning stage. In the pre-training stage, the model is trained on a large-scale unlabeled dataset to learn the underlying structure and features of the data. This process usually involves unsupervised learning or self-supervised learning tasks, such as predicting the next word, predicting the missing word, predicting the future frame, etc. After pre-training, the model will get a pre-trained weight, which contains a lot of general features and knowledge. Then, in the fine-tuning stage, these pre-trained weights are used as the initial weights of the model, and then trained on specific labeled tasks (such as classification, regression, etc.) to adjust the weights of the model to adapt to specific tasks. The pre-training method has achieved significant results in many deep learning applications, especially in natural language processing, computer vision, and speech recognition.
[0073] 5) Molecular denoising pre-training method, an unsupervised learning method for molecular structure data, mainly includes two steps: noise addition and denoising. In the noise addition step, first add noise to the original 3D coordinates of the molecule, the noise can be Gaussian noise, salt and pepper noise, etc., or more complex physical-based perturbation. The size and type of noise can be adjusted according to the specific task and data. In the denoising step, a model is trained, whose goal is to recover the original 3D coordinates of the molecule from the noisy 3D coordinates of the molecule. In this process, the model can learn the ability to recover the original molecular structure, thereby obtaining a robust representation of the 3D coordinates of the molecule, which can be used for subsequent supervised learning tasks, such as property prediction, drug molecule screening, etc.
[0074] 6) Multi-Layer Perceptron (MLP), a feedforward artificial neural network model, usually includes an input layer, one or more hidden layers and an output layer, each layer contains a number of neurons, and the neurons between adjacent layers are connected through weights. The main purpose of multi-layer perception is to learn the mapping relationship between input data and output data.
[0075] 7) Graph Neural Network (GNN), refers to the use of neural networks to learn graph structure data, extract and explore features and patterns in graph structure data, meet the needs of clustering, classification, prediction, segmentation, generation, etc. Graph learning task algorithm.
[0076] 8) Equivariant Map, an equivariant map is a function between two sets that commutes with the group action. For example, let G be a group, X and Y be two associated sets that can be acted on by the group G. If f(g·x) = g·f(x) for all g∈G and x∈X, then f:X→Y is called an equivariant map.
[0077] 9) Equivariant Graph Neural Networks (EGNN), is a specially designed graph neural network for processing geometric graph data, the core idea of equivariant graph neural network is to integrate the group structure of geometric transformation in the design of network, so that the network can recognize and utilize the invariance of graph data under transformation, considering the characteristics of geometric transformation such as translation, rotation and flip in the operation of graph structure data, so that equivariant graph neural network can process and learn geometric graph data more effectively.
[0078] 10) Compound, also known as molecule, is a whole composed of atoms that make up molecules, combined together according to a certain bonding order and spatial arrangement, such bonding order and spatial arrangement is called molecular structure. Molecule is the smallest unit that can exist independently in matter, which is relatively stable and maintains the physical and chemical properties of the matter.
[0079] 11) Atom, is the smallest particle that cannot be further divided in chemical reactions.
[0080] 12) Chemical bond, is the strong interaction force between two or more adjacent atoms (or ions) in pure substance molecules or crystals. The force that combines ions or atoms is called chemical bond.
[0081] 13) Geometric information, refers to the 3D coordinates of the atoms that make up the molecule.
[0082] 14) Non-geometric information, refers to some properties of the atoms that make up the molecule, such as the number of charges, the number of protons, the number of neutrons, etc.
[0083] 15) Equivariant graph transformer, refers to an optional network structure provided by the embodiments of the present application for encoding the geometric information and non-geometric information of the compound (molecule).
[0084] 16) Dependency, refers to the mutual influence and association between different elements of input data in attention mechanism (such as geometric information and non-geometric information). Specifically, attention mechanism aims to measure (such as by calculating attention weight) the relative importance between different elements in input data in a way, so as to better capture and process the dependency in data.
[0085] 17) Molecular properties refer to the basic characteristics of the molecules that make up a substance, such as the solubility, activity, melting point, boiling point, polarity, etc. of the molecules.
[0086] In the related art molecular property prediction method, the isometric graph neural network-based molecular property prediction method models the geometric transformations such as translation, rotation, and inversion of the molecular structure, but does not further consider the local information brought by the chemical bonds of the molecules, which affects the prediction accuracy of the molecular properties.
[0087] The embodiments of the present application provide a molecular property prediction method, a molecular property prediction model training method, device, equipment, computer readable storage medium and computer program product, which can improve the accuracy of molecular property prediction.
[0088] The following describes an exemplary application of the device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as a notebook computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart television, a vehicle-mounted terminal, and various types of terminal devices. It can also be implemented as a server.
[0089] Referring to Figure 1 , Figure 1 is a structural diagram of a molecular property prediction system architecture provided by the embodiments of the present application, Figure 1 The server 100, the terminal device 200 and the network 300 are involved. The terminal device 200 connects the server 100 through the network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two.
[0090] In some embodiments, the embodiments of the present application can be implemented by the server and the terminal device. For example, the server 100 obtains a second molecular property prediction model by the molecular property prediction model training method provided by the embodiments of the present application, the terminal device 200 sends the geometric information and non-geometric information of the compound to the server 100, and the server 100 obtains the molecular properties of the compound by the molecular property prediction method provided by the embodiments of the present application, and sends the molecular properties of the compound to the terminal device 200.
[0091] Here, the server 100 can be a single server. For this case, the molecular property prediction model training method and the molecular property prediction method provided by the embodiments of the present application can be implemented by the same server. The server 100 can also be a cluster of servers. For the case where the server 100 is a server cluster, the molecular property prediction model training method and the molecular property prediction method provided by the embodiments of the present application can be implemented by different servers, which are not limited by the embodiments of the present application.
[0092] In other embodiments, the present invention can be implemented solely on a terminal device. Terminal device 200 sends a request to server 100. Server 100 receives the request and sends a second molecular property prediction model for performing the molecular property prediction method provided in the present invention to terminal device 200. Terminal device 200 receives the second molecular property prediction model sent by the server and downloads it locally, then obtains the molecular properties corresponding to the compound using the second molecular property prediction model.
[0093] In some embodiments, the terminal device or server can implement the prediction method of the molecular property provided by the embodiment of the present application or the training method of the molecular property prediction model by running various computer executable instructions or computer programs. For example, the computer executable instruction can be a command, machine instruction or software instruction of the microprogram level. The computer program can be a native program or software module in the operating system. In short, the above-mentioned computer executable instruction can be an instruction in any form, and the above-mentioned computer program can be an application program, module or plug-in in any form, and the terminal device includes but is not limited to a mobile phone, a computer, an intelligent voice interaction device, an intelligent home appliance, a vehicle terminal, an aircraft, etc.
[0094] In some embodiments, multiple servers may be organized into a blockchain network, with server 100 being a node on the blockchain network. Information connections may exist between each node in the blockchain network, and information may be transmitted between nodes via the information connections. Data related to the molecular property prediction method or molecular property prediction model training method provided in the embodiments of the present application may be stored on the blockchain.
[0095] In some embodiments, the server 100 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal device and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.
[0096] The embodiments of the present application can be implemented by means of artificial intelligence (AI) technology. Artificial intelligence is a theory, method, technology and application system for perceiving environment, acquiring knowledge and using knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0097] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model, basic model, which can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.
[0098] For example, the server is used for predicting molecular properties, see Figure 2A , Figure 2A is a first structural schematic diagram of the server provided by the embodiments of the present application, Figure 2A The server 100-1 shown in the server 100-1 includes at least one processor 110-1, a memory 130-1 and at least one network interface 120-1. Each component in the server 100-1 is coupled together through a bus system 140-1. It can be understood that the bus system 140-1 is used to realize the connection communication between the components. The bus system 140-1 includes a data bus, a power bus, a control bus and a status signal bus. However, in order to clearly illustrate, all kinds of buses are marked as bus system 140-1 in the Figure 2A
[0099] The processor 110-1 can be an integrated circuit chip with signal processing capability, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor.
[0100] The memory 130-1 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. The memory 130-1 optionally includes one or more storage devices remotely located from the processor 110-1 in a physical location.
[0101] The memory 130-1 includes volatile memory or nonvolatile memory, and can also include both volatile and nonvolatile memory. Nonvolatile memory can be read only memory (ROM), and volatile memory can be random access memory (RAM). The memory 130-1 described in embodiments of the present application is intended to include any suitable type of memory.
[0102] In some embodiments, the memory 130-1 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are exemplarily illustrated below.
[0103] The operating system 131-1 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks;
[0104] The network communication module 132-1 is used to communicate with other electronic devices via one or more (wired or wireless) network interfaces 120-1, examples of which include Bluetooth, wireless compatibility authentication (WiFi), and universal serial bus (USB), and the like;
[0105] In some embodiments, the device provided by the embodiments of the present application can be realized in software, Figure 2A A prediction device 133 of molecular properties stored in the memory 130-1 is shown, which can be software in the form of programs and plug-ins, including the following software modules: a data acquisition module 1331, a data processing module 1332, and a property prediction module 1333, which are logical, and thus can be combined or further split according to the functions implemented. The functions of each module will be described below.
[0106] Taking a server for training a molecular property prediction model as an example, referring to Figure 2B , Figure 2B is a second structural schematic diagram of the server provided by the embodiments of the present application, Figure 2BThe illustrated server 100-2 includes at least one processor 110-2, a memory 130-2, and at least one network interface 120-2. The various components in the server 100-2 are coupled together by a bus system 140-2. As will be appreciated by those skilled in the art, the bus system 140-2 is used to facilitate communication between the various components. The bus system 140-2 is illustrated as including a data bus to facilitate the transfer of data between the components. However, one skilled in the art will recognize the techniques described herein can be implemented using a variety of bus architectures including a System-on-a-Chip (SoC) architecture. Figure 2B For a detailed description of the processor 110-2 and the memory 130-2, see the foregoing, which is incorporated by reference herein.
[0107] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software, Figure 2B A training apparatus 134 of the molecular property prediction model stored in the memory 130-2 is shown, which can be software in the form of programs and plug-ins, including the following software modules: a data acquisition module 1341, a noise adding module 1342, a data processing module 1343, and a training module 1344. These modules are logical, and thus can be combined or further split according to the functions implemented. The functions of the various modules will be described below.
[0108] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software,
[0109] The prediction method of the molecular property provided by the embodiments of the present application will be described below with reference to an exemplary application and implementation of the server provided by the embodiments of the present application, taking the server as the execution subject. See Figure 3A , Figure 3A is the first flowchart of the prediction method of the molecular property provided by the embodiments of the present application, which will be described below with reference to Figure 3AThe steps shown are illustrative.
[0110] In step 101, geometric information and non-geometric information of the compound are obtained.
[0111] In some embodiments, the compound (M) consists of atoms (i), which include two types of information: atomic geometric information (G) and atomic non-geometric information (N), where the atomic geometric information is the 3D coordinates of the atom, i.e. The atomic non-geometric information is some properties of the atom itself, such as the number of charges, the number of protons, the number of neutrons, etc., i.e. d M is the feature dimension of the atom. The compound is represented as a set of atomic geometric information and atomic non-geometric information, i.e.
[0112] In step 102, first encoding is performed based on the geometric information and the non-geometric information, to obtain first geometric encoding information and first non-geometric encoding information.
[0113] In some embodiments, the first encoding is performed multiple times, and the number of times of the first encoding is the same as the input dimension of the attention encoding. The multiple first encodings are respectively implemented by multiple equivariant graph neural networks (the multiple equivariant graph neural networks are trained by not sharing parameters), which include multiple network layers, as shown in Figure 6 , Figure 6 is a schematic diagram of the principle of the equivariant graph transformer provided in the embodiments of the present application. The geometric information and the non-geometric information of the molecular graph can be respectively encoded by three equivariant graph neural networks (corresponding to the equivariant graph neural network-1, the equivariant graph neural network-2 and the equivariant graph neural network-3 in Figure 6 , the four equivariant graph neural networks shown in Figure 6 are not shared in network parameters), and the obtained first geometric encoding information and first non-geometric encoding information output by the three equivariant graph neural networks are taken as the input of the equivariant attention layer (corresponding to the attention encoding).
[0114] In some embodiments, the first geometric encoding information includes the geometric encoding information of each node in the graph structure of the compound, and the first non-geometric encoding information includes the non-geometric encoding information of each node in the graph structure of the compound, where the graph structure includes multiple nodes, and the nodes connected to each node are the neighbor nodes corresponding to the nodes.
[0115] In some embodiments, referring to Figure 3B , Figure 3A the step 102 shown can be implemented by sequentially traversing the multiple network layers, and performing the following steps 1021 to 1024 for the current network layer traversed, which will be specifically explained below.
[0116] In step 1021, in response to the current network layer being the first layer, the geometric information and the non-geometric information of each node in the graph structure of the compound are input into the current network layer.
[0117] In some embodiments, referring to Figure 7D , Figure 7D is a schematic diagram of the graph structure of the compound provided by the embodiments of the present application, for Figure 7D , taking a benzene molecule as an example, atoms are taken as nodes, and chemical bonds are taken as edges, the graph structure of the benzene molecule is abstracted, the geometric information of each node corresponds to the 3D coordinates of the atom, and the non-geometric information of each node corresponds to the properties of the atom itself, such as the number of charges, the number of protons, the number of neutrons, and the like, and the geometric information and the non-geometric information of each node in the graph structure of the compound are input into the current network layer.
[0118] In step 1022, in response to the current network layer being the second layer and the network layers thereafter, the message encoding between each node and the corresponding neighbor node in the current network layer is obtained.
[0119] In some embodiments, the nodes connected to each node are the neighbor nodes corresponding to the node, and the message encoding between each node and the corresponding neighbor node in the current network layer, i.e., the lth layer, can be represented by the following formula:
[0120]
[0121] wherein m ij denotes the message encoding between the node i and the corresponding neighbor node j, denotes the non-geometric information output by the node i in the l-1 network layer, denotes the non-geometric information output by the neighbor node j in the l-1 network layer, denotes the geometric information output by the node i in the l-1 network layer, denotes the geometric information output by the neighbor node j in the l-1 network layer, e ij denotes the edge between the node i and the neighbor node j, and φ m denotes a multi-layer perception (MLP), and formula (1) can be further represented as: wherein Concat denotes a splicing operation.
[0122] In step 1023, the geometric encoding information corresponding to each node output by the current network layer is determined through the message encoding and the geometric encoding information output by the previous network layer.
[0123] In some embodiments, the geometric encoding information corresponding to each node output by the current network layer can be represented by the following formula:
[0124]
[0125] wherein, denotes the geometric encoding information of node i output by the current network layer, i.e., the l-th layer, φ x denotes a multi-layer perception (MLP), m ij and The description of formula (1) is given above.
[0126] In step 1024, the non-geometric encoding information corresponding to each node output by the current network layer is determined by the message encoding and the non-geometric encoding information output by the previous network layer.
[0127] In some embodiments, the non-geometric encoding information corresponding to each node output by the current network layer can be represented by the following formula:
[0128]
[0129] wherein, denotes the non-geometric encoding information of node i output by the current network layer, i.e., the l-th layer, φ h denotes a multi-layer perception (MLP), m ij and The description of formula (1) is given above.
[0130] In other embodiments, an attention mechanism can be introduced in the isometric graph neural network, so as to better model the dependency between adjacent atoms (nodes), i.e., to introduce a weight matrix W Q and W K , to calculate the attention weight β ij between adjacent atoms (nodes).
[0131] For example, formula (2) can be updated as: may be updated as: wherein,
[0132] Continuing to refer to Figure 3A , in step 103, based on the first geometric encoding information and the first non-geometric encoding information, a dependency-based attention encoding is performed to obtain non-geometric attention encoding information.
[0133] In some embodiments, the first geometric encoding information includes the geometric encoding information of each node in the graph structure of the compound, and the first non-geometric encoding information includes the non-geometric encoding information of each node in the graph structure of the compound, as described in Figure 3C , Figure 3A The step 103 shown in the figure can be implemented by the following steps 1031 to 1033, which are specifically described as follows.
[0134] In step 1031, the difference information between the geometric coding information of each node and the geometric coding information of the neighboring nodes is determined, and based on the difference information and the non-geometric information of the neighboring nodes, the first dependency relationship and the second dependency relationship between each node and the neighboring nodes are determined.
[0135] In some embodiments, the number of first encodings and the number of determining difference information are multiple and the number of times is the same, see Figure 3D , Figure 3C Step 1031 shown can be implemented by following steps 10311 to 10314, which are described in detail below.
[0136] In step 10311, after the first encoding, the difference information between the geometric coding information of each node and the geometric coding information of the corresponding neighboring node is determined, and the difference information and the non-geometric information of the neighboring node are spliced to obtain the first splicing information.
[0137] In some embodiments, after the first encoding (corresponding to Figure 6 The geometric coding information and non-geometric coding information of each node can be expressed as in, represents the geometric encoding information of node i, Represents the non-geometric coding information of node i. The difference between the geometric coding information of each node and the geometric coding information of the corresponding neighboring node can be expressed as The difference information and the non-geometric information of the neighboring nodes are spliced together to obtain the first spliced information, which can be expressed as
[0138] In step 10312, nonlinear mapping is performed on the first splicing information to obtain a first dependency relationship.
[0139] In some embodiments, nonlinear mapping is performed on the first splicing information to obtain a first dependency relationship, which can be expressed by the following formula:
[0140]
[0141] Among them, k ij represents the first dependency, φ k Denotes a multi-layer perceptron (MLP), and nonlinear mapping is performed on the first splicing information by the multi-layer perceptron to obtain a first dependency relationship. For example, nonlinear mapping is performed on the first splicing information by the multi-layer perceptron, which can be further expressed by the following formula:
[0142]
[0143] where Concat denotes a concatenation operation, f() denotes an activation function (such as Sigmoid, tanh, and ReLU, etc.), W1 and W2 are weight parameters of two linear layers of the MLP, and b1 and b2 are bias parameters of the two linear layers of the MLP.
[0144] Nonlinear mapping refers to a mapping relationship that cannot be described by a straight line. In nonlinear mapping, the relationship between input and output is not a simple proportional relationship, but is derived through complex calculations. For example, a common nonlinear mapping is the sine function. The sine function maps the input angle to a value between -1 and 1, and the relationship is not a simple proportional relationship, but a periodic curve. Another example is the exponential function, which maps the input value to an exponentially growing output value, which is also a nonlinear relationship. In general, nonlinear mapping is widely used in practical applications and can describe and analyze various complex relationships and phenomena.
[0145] In some embodiments, nonlinear mapping can be achieved by adding a nonlinear activation function in a neural network. In a neural network, the linear combination of multiple neurons can be represented as a linear transformation (such as formula 5) , and by adding a nonlinear activation function, the representation capability of the network can be extended from linear mapping to more complex nonlinear mapping (such as formula 5) Nonlinear mapping can be achieved by manually designing different types of nonlinear activation functions, designing the depth and width of the network, and adjusting the structure of the network. The specific network structure or nonlinear activation function used for nonlinear mapping in the embodiments of the present application is not limited.
[0146] In step 10313, after the second first encoding, the difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node is determined, and the difference information and the non-geometric information of the neighbor node are concatenated to obtain the second concatenation information.
[0147] In some embodiments, after the second first encoding (corresponding to the encoding process of the medium variational graph neural network-2), the geometric encoding information and the non-geometric encoding information of each node can be represented as Figure 6 . where represents the geometric encoding information of node i, represents the non-geometric encoding information of node i, and the difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node can be represented as The difference information and the non-geometric information of the neighbor node are concatenated to obtain the second concatenation information, which can be represented as
[0148] In step 10314, the second splicing information is non-linearly mapped to obtain a second dependency relationship.
[0149] In some embodiments, the non-linear mapping of the second splicing information to obtain the second dependency relationship can be represented by the following formula:
[0150]
[0151] wherein v ij represents the second dependency relationship, and φ V represents a multi-layer perception (MLP) through which the second splicing information is non-linearly mapped to obtain the second dependency relationship, and the specific implementation is described above in step 10312.
[0152] Continuing to refer to Figure 3C In step 1032, the attention weight is determined based on the first dependency relationship and the non-geometric encoding information of each node.
[0153] In some embodiments, referring to Figure 3E , Figure 3C The step 1032 shown can be implemented by the following steps 10321 to 10322, which are described in detail below.
[0154] In step 10321, the non-geometric encoding information of each node is non-linearly mapped to obtain a query vector code corresponding to each node.
[0155] In some embodiments, after the third encoding of the first encoding (corresponding to the encoding processing of the isometric graph neural network-3 in Figure 6 ), the geometric encoding information and the non-geometric encoding information of each node can be represented as wherein represents the geometric encoding information of node i, represents the non-geometric encoding information of node i, and the non-geometric encoding information of each node is non-linearly mapped to obtain a query vector code corresponding to each node, which can be represented by the following formula:
[0156]
[0157] wherein q i represents the query vector code corresponding to node i, and φ Q represents a multi-layer perception (MLP).
[0158] In step 10322, the dot product result of the query vector code and the first dependency relationship is obtained, and the dot product result is normalized to obtain the attention weight.
[0159] In some embodiments, the dot product result of the query vector encoding and the first dependency can be expressed as: Normalizing the dot product result, the attention weight can be expressed as follows:
[0160]
[0161] Among them, α ij represents the attention weight, and exp represents the exponential function.
[0162] Continue to see Figure 3C In step 1033, the second dependency is weighted and summed based on the attention weight to obtain the non-geometric attention encoding information corresponding to each node.
[0163] In some embodiments, the second dependency is weighted and summed by the attention weight to obtain the non-geometric attention encoding information corresponding to each node, which can be expressed by the following formula:
[0164]
[0165] in, Represents the non-geometric attention encoding information corresponding to node i.
[0166] Continue to see Figure 3A ,In step 104, the non-geometric attention encoding information is spliced with the non-geometric information to obtain non-geometric splicing information.
[0167] In some embodiments, see Figure 3F , Figure 3A The step 104 shown can be implemented by following steps 1041 to 1043, which are described in detail below.
[0168] In step 1041, dependency-based attention encoding is performed based on the first geometric encoding information and the first non-geometric encoding information to obtain geometric attention encoding information.
[0169] In some embodiments, attention encoding based on dependency relationship is performed based on multiple first geometric coding information and first non-geometric coding information to obtain geometric attention coding information corresponding to the first geometric coding information, see Figure 3G , Figure 3F Step 1041 shown can be implemented through steps 10411 to 10414, which are described in detail below.
[0170] In step 10411, the difference information between the geometric coding information of each node and the geometric coding information of the corresponding neighbor node is determined, and based on the difference information and the non-geometric information of the neighbor node, the first dependency relationship and the second dependency relationship between each node and the neighbor node are determined.
[0171] In some embodiments, the implementation of determining the first dependency relationship and the second dependency relationship between each node and the neighbor node, see the description of step 1031 above, will not be repeated here.
[0172] In step 10412, the attention weight is determined based on the first dependency relationship and the non-geometric encoding information of each node.
[0173] In some embodiments, the implementation of determining the attention weight based on the first dependency relationship and the non-geometric encoding information of each node, see the description of step 1032 above, will not be repeated here.
[0174] In step 10413, the weighted sum of the product of the second dependency relationship and the difference is weighted based on the attention weight, and the weighted sum result corresponding to each node is obtained.
[0175] In some embodiments, the weighted sum of the product of the second dependency relationship and the difference information based on the attention weight, and the weighted sum result corresponding to each node can be represented as Wherein, α ij represents the attention weight, Ψ x represents the multi-layer perception (MLP), v ij represents the second dependency relationship, represents the difference information after the third first encoding, see the description of step 10321 above.
[0176] In step 10414, the weighted sum result corresponding to each node is added to the geometric encoding information of each node, and the geometric attention encoding information corresponding to each node is obtained.
[0177] In some embodiments, the weighted sum result corresponding to each node is added to the geometric encoding information of each node, and the geometric attention encoding information corresponding to each node can be represented by the following formula:
[0178]
[0179] Wherein, represents the geometric attention encoding information corresponding to node i, represents the geometric encoding information corresponding to node i.
[0180] Continue to see Figure 3F In step 1042, the second non-geometric encoding information is obtained by performing the second encoding based on the dependency relationship on the non-geometric attention encoding information and the geometric attention encoding.
[0181] In some embodiments, the second encoding based on the dependency relationship is performed on the non-geometric attention encoded information and the geometric attention encoding (corresponding to the encoding process of the equivariant graph neural network-4 in Figure 6 , to obtain second geometric encoding information and second non-geometric encoding information, wherein the second encoding is implemented by an equivariant graph neural network, the equivariant graph neural network includes a plurality of network layers, and the specific implementation of the second encoding can be referred to the description of the first encoding above, and here, the second encoding is performed once.
[0182] In step 1043, the second non-geometric encoding information is spliced with the non-geometric information to obtain non-geometric spliced information.
[0183] In some embodiments, the non-geometric spliced information includes the second non-geometric encoding information corresponding to each node and the non-geometric information.
[0184] Continuing to refer to Figure 3A , in step 105, a non-linear mapping is performed based on the non-geometric spliced information to obtain the molecular property of the compound.
[0185] In some embodiments, referring to Figure 3H , Figure 3A Step 105 shown in the figure can be implemented by the following steps 1051 to 1053, which are specifically described below.
[0186] In step 1051, a first non-linear mapping process is performed on the non-geometric spliced information to obtain non-geometric mapping information.
[0187] In some embodiments, the first non-linear mapping process can be performed on the non-geometric spliced information by a multi-layer perception to obtain the non-geometric mapping information, and the specific implementation of the non-linear mapping can be referred to the description in formula (5) in step 10312 above.
[0188] In step 1052, a second non-linear mapping process is performed on the non-geometric mapping information to obtain the probability value of each molecular property corresponding to the compound.
[0189] In some embodiments, the second non-linear mapping process can be performed on the non-geometric mapping information by an MLP to obtain the probability value of each molecular property corresponding to the compound, the MLP model is composed of an input layer, a hidden layer and an output layer, the input layer is responsible for receiving input data (non-geometric mapping information), the hidden layer performs non-linear transformation on the input data through a series of calculations between neurons and activation functions (see the description of formula (5) above), and the output layer converts the result of the hidden layer into an output value (i.e. the probability value of the molecular property), such as mapping the output of the hidden layer to a probability distribution by a softmax activation function, indicating the probability value of each molecular property.
[0190] In step 1053, the molecular property with the highest probability value is taken as the molecular property of the compound.
[0191] Taking the above example, assuming that a compound has three possible molecular properties (such as solubility, activity, melting point, etc.), the output layer of the MLP will have three output nodes, and the output of each node is a probability value. By comparing the three probability values, it can be determined which molecular property has the highest probability, and that property is taken as the predicted molecular property of the compound.
[0192] Through steps 101 to 105, after encoding the geometric information and non-geometric information of the compound, the first geometric encoding information and the first non-geometric encoding information are combined into the attention mechanism based on the dependency relationship, the generation of the non-geometric attention encoding information is affected by the first geometric encoding information, so that the non-geometric attention encoding information includes the geometric information of the molecule, that is, it includes the local information brought by the chemical bond of the molecule. By splicing the non-geometric attention encoding information with the non-geometric information, non-geometric splicing information is obtained, which further reduces the loss of information. Finally, the molecular property of the compound is obtained through the non-geometric splicing information, thereby achieving the beneficial effect of improving the accuracy of molecular property prediction.
[0193] The training method of the molecular property prediction model provided in the embodiments of the present application will be described below with reference to an exemplary application and implementation of a server provided by the embodiments of the present application, taking the server as the execution subject. Figure 4A , Figure 4A is a first flowchart of the training method of the molecular property prediction model provided by the embodiments of the present application, which will be described with reference to the steps shown in Figure 4A .
[0194] In some embodiments, referring to Figure 5 , Figure 5 is a schematic diagram of the training architecture of the molecular property prediction model provided by the embodiments of the present application. The training of the molecular property prediction model can be divided into two stages (corresponding to the first training task and the second training task in Figure 5 ), wherein the first training stage encodes the noisy molecular graph data through the isometric graph transformer (corresponding to steps 203 to 204 below), confirms the first loss value through the noise prediction module (corresponding to step 205 below), thereby training the initialized molecular property model in an unsupervised manner to obtain a first molecular property prediction model; the second stage further trains the first molecular property prediction model in a supervised training manner (the second training task) to obtain a second molecular property prediction model (corresponding to steps 206 to 208 below).
[0195] In step 201, a training data set is obtained, where the training data set includes a plurality of molecular graph data, and the molecular graph data includes geometric information and non-geometric information.
[0196] In some embodiments, referring to Figure 7D , the geometric information of each node corresponds to the 3D coordinates of the atom, and the non-geometric information of each node corresponds to the properties of the atom itself, such as the number of charges, the number of protons, and the number of neutrons. Figure 7D
[0197] In step 202, each molecular graph data is subjected to noise addition processing to obtain noise-added molecular graph data of each molecular graph data.
[0198] In some embodiments, referring to Figure 4B , Figure 4A The step 202 shown in the figure can be implemented by the following steps 2021 to 2023, which are described in detail below.
[0199] In step 2021, each molecular graph data is split into a plurality of sub-graph data.
[0200] In some embodiments, the molecular graph data includes a plurality of nodes, and the node connected to each node is the neighbor node corresponding to the node, referring to Figure 4C , Figure 4B The step 2021 shown in the figure can be implemented by performing the following steps 20211 to 20213 for each molecular graph data, which are described in detail below.
[0201] In step 20211, a node is selected from the plurality of nodes included in the molecular graph as a starting node, and the starting node is added to the sub-graph node set.
[0202] In some embodiments, a node is randomly selected from the plurality of nodes included in the molecular graph as a starting node, and the starting node is added to the initialized sub-graph node set. For example, given a molecular graph G(V, E), where V represents the node set, and E represents the edge (connection between nodes) set, an initialized empty sub-graph G'(V', E') is created, where V' represents the node set of the sub-graph, and E' represents the edge set of the sub-graph. A node v is randomly selected from the node set V of the original graph G and added to the sub-graph node set V'. sub sub sub sub sub sub
[0203] In step 20212, the following process is iteratively performed: selecting one neighbor node from the neighbor nodes of the starting node, and adding the selected neighbor node to the subgraph node set.
[0204] In some embodiments, the following process is iteratively performed: randomly selecting one neighbor node from the neighbor nodes of the starting node, and adding the selected neighbor node and the edge connecting the starting node and the neighbor node to the subgraph node set. For example, for a node v in V sub , a node u is randomly selected from its neighbor nodes, and when u is not in V sub , u and the edge connecting v and u are added to V sub and E sub .
[0205] In step 20213, in response to the number of nodes in the subgraph node set reaching a preset node value, the subgraph node set is taken as the subgraph data corresponding to the molecule graph data.
[0206] In some embodiments, the process of step 20212 is repeated until the number of vertices in V sub reaches a preset node value n.
[0207] Continuing to refer to Figure 4B , in step 2022, random noise data of a preset intensity is obtained.
[0208] In some embodiments, the random noise data of the preset intensity can be random noise subject to Gaussian distribution, Poisson noise, uniform noise, etc. For example, the random noise subject to Gaussian distribution can be expressed as 0 is the mean value, and σ 2I is the variance (corresponding to the preset intensity), where I is a unit matrix, indicating that the covariance matrix of the noise is a unit matrix, i.e., the noise is isotropic and has no directionality.
[0209] In step 2023, the random noise data is superimposed on the geometric information of each subgraph data to obtain the noisy molecule graph data of each molecule graph data.
[0210] In some embodiments, the random noise data is superimposed on the geometric information of each subgraph data (e.g., Gaussian noise is added to the 3D coordinates of each node representing an atom of the subgraph data) to obtain the noisy molecule graph data of each molecule graph data.
[0211] Continuing to refer to Figure 4A , in step 203, the first encoding of each noisy molecule graph data is performed by the initialized molecular property prediction model to obtain the first geometric encoding information and the first non-geometric encoding information corresponding to each noisy molecule graph data.
[0212] In some embodiments, the first encoding is performed multiple times, and the number of times of the first encoding is the same as the input dimension of the attention encoding, and the multiple times of the first encoding are respectively implemented by multiple isometric graph neural networks (the multiple isometric graph neural networks are trained by not sharing parameters), the isometric graph neural network includes multiple network layers, see Figure 6 The three isometric graph neural networks (corresponding to Figure 6 Isometric graph neural network-1, isometric graph neural network-2 and isometric graph neural network-3 in the above embodiment, Figure 6 The four isometric graph neural networks shown in the above embodiment are not shared in network parameters) respectively encode the geometric information and the non-geometric information of each noisy molecular graph data, and the first geometric encoding information and the first non-geometric encoding information output by the three isometric graph neural networks are respectively taken as the input of the isometric attention layer (corresponding to the attention encoding).
[0213] In some embodiments, the first geometric encoding information includes the geometric encoding information of each node in the noisy molecular graph data, and the first non-geometric encoding information includes the non-geometric encoding information of each node in the noisy molecular graph data.
[0214] In some embodiments, in response to the current network layer being the first layer, the geometric information and the non-geometric information of each node in the noisy molecular graph data are input into the current network layer; in response to the current network layer being the second layer and the subsequent network layers, the message encoding between each node and the corresponding neighbor node in the current network layer is obtained; the geometric encoding information corresponding to each node output by the current network layer is determined by the message encoding and the geometric encoding information output by the previous network layer; the non-geometric encoding information corresponding to each node output by the current network layer is determined by the message encoding and the non-geometric encoding information output by the previous network layer, and the specific implementation of obtaining the first geometric encoding information and the first non-geometric encoding information corresponding to each noisy molecular graph data can be referred to the description of step 102 above, which will not be repeated here.
[0215] In step 204, the attention encoding based on the dependency relationship is performed based on the first geometric encoding information and the first non-geometric encoding information corresponding to each noisy molecular graph data by the molecular property prediction model, to obtain the non-geometric attention encoding information corresponding to each noisy molecular graph data.
[0216] In some embodiments, the first geometric encoding information includes the geometric encoding information of each node in the noisy molecular graph data, and the first non-geometric encoding information includes the non-geometric encoding information of each node in the noisy molecular graph data; the attention encoding based on the dependency relationship is performed based on the first geometric encoding information and the first non-geometric encoding information corresponding to each noisy molecular graph data, to obtain the non-geometric attention encoding information corresponding to each noisy molecular graph data, which can be implemented by the following implementation.
[0217] determining difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node, determining a first dependency relationship and a second dependency relationship between each node and the neighbor node based on the difference information and the non-geometric information of the neighbor node; determining an attention weight based on the first dependency relationship and the non-geometric encoding information of each node; and performing weighted summation on the second dependency relationship based on the attention weight to obtain the non-geometric attention encoding information corresponding to each node in the noisy subgraph data.
[0218] In the embodiment, the number of times of the first encoding and the number of times of determining the difference information are both multiple and the same, and the first dependency relationship and the second dependency relationship between each node and the neighbor node can be determined by the following implementation: after the first time of the first encoding, the difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node is determined, and the difference information and the non-geometric information of the neighbor node are spliced to obtain first spliced information; the first spliced information is subjected to nonlinear mapping to obtain the first dependency relationship; after the second time of the first encoding, the difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node is determined, and the difference information and the non-geometric information of the neighbor node are spliced to obtain second spliced information; the second spliced information is subjected to nonlinear mapping to obtain the second dependency relationship. For specific determination of the first dependency relationship and the second dependency relationship, refer to the description of step 1031 above, which will not be repeated here.
[0219] In the embodiment, the determination of the attention weight based on the first dependency relationship and the non-geometric encoding information of each node can be implemented by the following implementation: the non-geometric encoding information of each node is subjected to nonlinear mapping to obtain a query vector encoding corresponding to each node; a dot product result of the query vector encoding and the first dependency relationship is obtained, and the dot product result is normalized to obtain the attention weight. For specific acquisition of the attention weight, refer to the description of step 1032 above, which will not be repeated here.
[0220] In the embodiment, the weighted summation of the second dependency relationship based on the attention weight to obtain the non-geometric attention encoding information corresponding to each node in the noisy subgraph data can be implemented by the following implementation: the weighted summation of the second dependency relationship based on the attention weight is performed to obtain the non-geometric attention encoding information corresponding to each node. For specific implementation, refer to the description of step 1033 above, which will not be repeated here.
[0221] In step 205, based on the non-geometric attention encoding information corresponding to each noisy subgraph data, a first training task for molecular property prediction is performed to obtain a first molecular property prediction model.
[0222] In some embodiments, referring to Figure 7A , Figure 7A is a schematic diagram of the structure of the first training task for molecular property prediction provided by the embodiments of the present application. The noise-added molecular graph data is encoded by the isometric map transformer (corresponding to steps 203 to 204), and the first loss value is confirmed by the noise prediction module (corresponding to step 205), so as to train the initialized molecular property model in an unsupervised manner, and obtain the first molecular property prediction model.
[0223] In some embodiments, referring to Figure 4D , Figure 4A The step 205 shown can be implemented by the following steps 2051 to 2054, which are specifically explained below.
[0224] In some embodiments, referring to Figure 7B , Figure 7B is a schematic diagram of the principle of the first training task for molecular property prediction provided by the embodiments of the present application. The energy data E corresponding to the noise-added molecular graph data is obtained by energy prediction on the non-geometric attention encoded information. Next, the force data F is further obtained on the basis of the energy E by the partial derivative operation. Finally, the first loss value is determined by the loss calculation between the force F and the noise gradient, so as to inversely update the parameters of the initialized molecular property prediction model according to the first loss value, and obtain the first molecular property prediction model.
[0225] In step 2051, the energy data corresponding to each noise-added molecular graph data is obtained by energy prediction on the non-geometric attention encoded information corresponding to each noise-added molecular graph data.
[0226] In some embodiments, the non-geometric attention encoded information includes the non-geometric attention encoded information of each node in the noise-added molecular graph data. The energy data corresponding to each noise-added molecular graph data is obtained by energy prediction on the non-geometric attention encoded information corresponding to each noise-added molecular graph data, which can be represented by the following formula:
[0227]
[0228] Wherein, E represents the energy data corresponding to the noise-added molecular graph data, f() represents an activation function (such as Sigmoid, tanh and ReLU activation function), W3 and W4 are weight parameters of two linear layers of MLP respectively, and b3 and b4 are bias parameters of two linear layers of MLP respectively.
[0229] In step 2052, the force data corresponding to each noise-added molecular graph data is obtained by partial derivation of the geometric information of the molecular graph data on the energy data.
[0230] In some embodiments, the partial derivative of the geometric information (X) of the molecular graph data with respect to the energy data is taken to obtain the force data corresponding to each noisy molecular graph data, which can be expressed as follows:
[0231]
[0232] In step 2053, a first loss value is obtained based on the force data corresponding to the noisy molecular graph data and the noise gradient.
[0233] In some embodiments, the first loss value can be expressed by the following formula:
[0234] loss=||F-∈ / σ 2 || 2 (13)
[0235] wherein, ∈ / σ 2 represents the noise gradient.
[0236] In step 2054, first gradient information is obtained by the first loss value, and the parameters of the initialized molecular property prediction model are updated according to the first gradient information to obtain a first molecular property prediction model.
[0237] In some embodiments, the first gradient information of the first loss value with respect to each parameter of the molecular property prediction model is obtained by a back propagation algorithm, and the parameters of the molecular property prediction model are updated using the obtained first gradient information according to a gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.). The above process is repeated until a certain number of iterations is reached or the molecular property prediction model converges, thereby completing the first training task of the molecular property prediction model.
[0238] By steps 201 to 205, noise is added at the subgraph level of the molecular graph data, and the noise is introduced at the subgraph level of the molecular structure, thereby simulating or simulating the possible small perturbations or defects in the actual molecular structure. The molecular property prediction model is trained to learn how to remove noise and interference from the noisy data and restore the original clean data, learn more distinctive and representative features that can better capture the essential information of the molecular data and represent more stable molecular structures, thereby achieving the beneficial effect of improving the performance and generalization ability of the molecular property prediction model.
[0239] In some embodiments, referring to Figure 4E , in Figure 4A the steps 205 shown, the following steps 206 to 208 can also be performed, which are described in detail below.
[0240] In some embodiments, referring to Figure 7C ,Figure 7C is a schematic diagram of a structure of a second training task for molecular property prediction provided by an embodiment of the present application. The molecular graph data is encoded by a first molecular property prediction model to obtain non-geometric attention encoding information corresponding to the molecular graph data (corresponding to steps 206 to 207 below), the non-geometric attention encoding information is mapped to corresponding predicted molecular properties by a molecular property prediction MLP, and then a second loss value is confirmed according to the predicted molecular properties and the molecular property labels, so as to update the parameters of the first molecular property prediction model to obtain a second molecular property prediction model (corresponding to step 208 below).
[0241] In step 206, the first encoding is performed on each molecular graph data by the first molecular property prediction model to obtain first geometric encoding information and first non-geometric encoding information corresponding to each molecular graph data.
[0242] In some embodiments, the specific implementation of the first encoding performed on each molecular graph data to obtain the first geometric encoding information and the first non-geometric encoding information corresponding to each molecular graph data can be referred to the description of step 203 above, which will not be repeated here.
[0243] In step 207, the attention encoding based on the dependency relationship is performed on each molecular graph data by the first molecular property prediction model based on the first geometric encoding information and the first non-geometric encoding information corresponding to each molecular graph data to obtain non-geometric attention encoding information corresponding to each molecular graph data.
[0244] In some embodiments, the specific implementation of obtaining the non-geometric attention encoding information corresponding to each molecular graph data can be referred to the description of step 204 above, which will not be repeated here.
[0245] In step 208, the second training task for molecular property prediction is performed based on the non-geometric attention encoding information corresponding to each molecular graph data to obtain a second molecular property prediction model.
[0246] In some embodiments, a preset number of molecular graph data in the plurality of molecular graph data carries a molecular property label (such as solubility, activity, melting point, boiling point, polarity, etc.), and the molecular graph data carrying the molecular property label is a labeled molecular graph data, and the molecular graph data not carrying the molecular property label is an unlabeled molecular graph data.
[0247] In some embodiments, referring to Figure 4F , Figure 4E Step 208 shown in the figure can be implemented by the following steps 2081 to 2083, which will be described in detail below.
[0248] In step 2081, nonlinear mapping is performed on the non-geometric attention encoding information of each molecular graph data with a molecular property label to obtain the first predicted molecular property of each molecular graph data with a molecular property label.
[0249] In some embodiments, the non-geometric attention encoding information can be non-linearly mapped using a multi-layer perceptron to obtain non-geometric mapping information; the non-geometric mapping information can be non-linearly mapped again to obtain the probability value of each molecular property corresponding to the molecular graph data; the molecular property with the highest probability value is used as the first predicted molecular property of the molecular graph data, wherein the specific implementation method for obtaining the predicted molecular property can be found in the description of step 105 above and will not be repeated here.
[0250] In step 2082 , a second loss value is determined using the first predicted molecular property and the molecular property label.
[0251] In some embodiments, the error between the first predicted molecular property and the molecular property label can be obtained through a preset loss function (such as a mean absolute error loss function, an L1 regularization loss function, etc.) to determine the second loss value.
[0252] In step 2083, second gradient information is obtained through the second loss value, and the parameters of the first molecular property prediction model are updated according to the second gradient information to obtain a second molecular property prediction model.
[0253] In some embodiments, the second gradient information of the second loss value for each parameter of the first molecular property prediction model is obtained through the back propagation algorithm, and the parameters of the first molecular property prediction model are updated using the obtained second gradient information according to the gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.). The above process is repeated until a certain number of iterations is reached or the first molecular property prediction model converges, thereby completing the second training task of the first molecular property prediction model and obtaining the second molecular property prediction model.
[0254] In some embodiments, see Figure 4G ,exist Figure 4F After step 2083 shown, the following steps 2084 to 2086 may also be executed, as described in detail below.
[0255] In step 2084 , a second predicted molecular property of the molecular graph data without molecular property labels in the plurality of molecular graph data is obtained using the second molecular property prediction model.
[0256] In some embodiments, the second predicted molecular property of each molecular graph data without a molecular property label in the plurality of molecular graph data is obtained. Please refer to the description of step 2081 above, which will not be repeated here.
[0257] In step 2085, the second predicted molecular property exceeding the preset prediction confidence threshold is taken as the molecular property label of the molecular graph data without the molecular property label to form new molecular graph data.
[0258] In some embodiments, the confidence values in the prediction results of the molecular property prediction model can be statistically analyzed by statistical analysis, the distribution of different confidence values is viewed, the value range and distribution of the confidence threshold are understood, and the prediction confidence threshold is set. For each molecular graph data, if the prediction confidence thereof is higher than the threshold, the second predicted molecular property thereof is taken as the new molecular property label, and the labeled molecular graph data is integrated into the new data set for subsequent analysis and processing.
[0259] In step 2086, the new molecular graph data with the molecular property label is obtained, wherein the new molecular graph data with the molecular property label is used to re-execute the second training task for molecular property prediction.
[0260] In some embodiments, in each iteration of the second training task, the molecular property prediction is first trained using the labeled data set, then the unlabeled data set is predicted using the learned molecular property prediction, and the prediction results with the prediction confidence higher than the threshold, i.e., the new molecular graph data with the molecular property label, are added to the labeled data set, and the molecular property prediction model training is performed again.
[0261] Through steps 2084 to 2086, the label information of the molecular graph data set with the molecular property label is combined with the molecular graph data set without the molecular property label, a joint training set is created, a semi-supervised learning algorithm is used to train the molecular property prediction model using the joint training set, and the beneficial effect that the molecular property prediction model can gradually use the information in the unlabeled data set to improve the performance is achieved.
[0262] The molecular property prediction method or the training method of the molecular property prediction model provided by the embodiments of the present application can be applied to various scenarios requiring molecular property prediction, some of which include: (1) drug discovery: such as predicting the molecular properties of compounds by molecular property prediction to help screen potential candidate drug molecules, etc.; (2) synthetic route planning: such as predicting the molecular properties of reaction products of compounds under different reaction conditions, thereby screening and optimizing the synthetic route through the molecular properties of the products, reducing unnecessary steps and side reactions, and improving synthesis efficiency and yield, etc.; (3) new compound discovery: such as predicting the molecular properties of new compounds, which can inspire experimenters to discover new compound types and new synthesis methods, and provide new candidate substances for drug discovery, material science, etc.
[0263] In the following, an exemplary application of the embodiments of the present application in the application scenario of drug discovery will be described. Referring to Figure 8 , Figure 8 is a flowchart of drug discovery provided by the embodiments of the present application, which will be described in detail below.
[0264] In step 301, geometric information and non-geometric information of a to-be-predicted compound are obtained.
[0265] In some embodiments, referring to Figure 7D , taking a benzene molecule (corresponding to the to-be-predicted compound) in Figure 7D as an example, the graph structure (i.e., the molecular graph data of the benzene molecule) of the benzene molecule is abstracted by taking atoms as nodes and chemical bonds as edges, the geometric information of each node corresponds to the 3D coordinates of the atom, and the non-geometric information of each node corresponds to the properties of the atom itself, such as the number of charges, the number of protons, the number of neutrons, and the like.
[0266] In step 302, a predicted molecular property of the to-be-predicted compound is obtained through the geometric information and the non-geometric information of the to-be-predicted compound.
[0267] In some embodiments, the second molecular property prediction model can be obtained by the training method of the molecular property prediction model provided by the embodiments of the present application, and the predicted molecular property of the to-be-predicted compound can be obtained by the second molecular property prediction model and the prediction method of the molecular property provided by the embodiments of the present application. The specific implementation of obtaining the second molecular property prediction model can be referred to the description of steps 201 to 208 above, and the specific implementation of obtaining the predicted molecular property of the to-be-predicted compound by the prediction method of the molecular property provided by the embodiments of the present application can be referred to the description of steps 101 to 105 above, which will not be described here.
[0268] In step 303, the to-be-predicted compound is screened through the molecular property, and the to-be-predicted compound is taken as a candidate drug molecule in response to the screening passing.
[0269] In some embodiments, according to the prediction result of the molecular property of the to-be-predicted compound, a compound with potential biological activity is selected for further experimental verification, so as to carry out subsequent drug development.
[0270] Through steps 301 to 303, systematic analysis and screening of the to-be-predicted compound are realized, and an efficient and accurate method for drug discovery is provided. Through the training method of the molecular property prediction model and the prediction method of the molecular property provided by the embodiments of the present application, the geometric information and the non-geometric information of the compound can be better utilized for molecular property prediction, which provides strong support for the screening and design of candidate drug molecules, so as to achieve the beneficial effects of accelerating the drug development process and improving the efficiency and success rate of drug discovery.
[0271] The following continues to illustrate an exemplary structure of the implementation of the training device 133 of the language model provided by the embodiments of the present application as a software module. In some embodiments, as shown in the figure, the software module stored in the prediction device 133 of the molecular property of the memory 130-1 can include: Figure 2A
[0272] The data acquisition module 1331 is configured to acquire geometric information and non-geometric information of a compound.
[0273] The data processing module 1332 is configured to perform first encoding based on the geometric information and the non-geometric information to obtain first geometric encoding information and first non-geometric encoding information.
[0274] In some embodiments, the data processing module 1332 is further configured to perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information to obtain non-geometric attention encoding information.
[0275] In some embodiments, the data processing module 1332 is further configured to splice the non-geometric attention encoding information with the non-geometric information to obtain non-geometric spliced information.
[0276] The property prediction module 1333 is configured to perform non-linear mapping based on the non-geometric spliced information to obtain a molecular property of the compound.
[0277] In some embodiments, the first geometric encoding information includes geometric encoding information of each node in a graph structure of the compound, and the first non-geometric encoding information includes non-geometric encoding information of each node in the graph structure of the compound, wherein the graph structure includes a plurality of nodes; the data processing module 1332 is further configured to determine difference information between the geometric encoding information of each node and the geometric encoding information of a neighbor node, determine a first dependency relationship and a second dependency relationship between each node and the neighbor node based on the difference information and the non-geometric information of the neighbor node; determine an attention weight based on the first dependency relationship and the non-geometric encoding information of each node; and perform weighted summation on the second dependency relationship based on the attention weight to obtain non-geometric attention encoding information corresponding to each node.
[0278] In some embodiments, the data processing module 1332 is further configured to perform non-linear mapping on the non-geometric encoding information of each node to obtain a query vector encoding corresponding to each node; obtain a dot product result of the query vector encoding and the first dependency relationship, and normalize the dot product result to obtain the attention weight.
[0279] In some embodiments, the number of times of the first encoding and the number of times of determining the difference information are both multiple times and the number of times are the same; the data processing module 1332 is further configured to, after the first time of the first encoding, determine the difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node, and splice the difference information and the non-geometric information of the neighbor node to obtain first spliced information; perform nonlinear mapping on the first spliced information to obtain the first dependency relationship; after the second time of the first encoding, determine the difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node, and splice the difference information and the non-geometric information of the neighbor node to obtain second spliced information; perform nonlinear mapping on the second spliced information to obtain the second dependency relationship.
[0280] In some embodiments, the data processing module 1332 is further configured to perform attention encoding based on the first geometric encoding information and the first non-geometric encoding information based on the dependency relationship to obtain geometric attention encoding information; perform second encoding based on the dependency relationship on the non-geometric attention encoding information and the geometric attention encoding to obtain second non-geometric encoding information; and splice the second non-geometric encoding information and the non-geometric information to obtain non-geometric spliced information.
[0281] In some embodiments, the first geometric encoding information includes geometric encoding information of each node in a graph structure of the compound, and the first non-geometric encoding information includes non-geometric encoding information of each node in the graph structure of the compound, wherein the graph structure includes a plurality of nodes; the data processing module 1332 is further configured to determine difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node, determine a first dependency relationship and a second dependency relationship between each node and a neighbor node based on the difference information and the non-geometric information of the neighbor node; determine an attention weight based on the first dependency relationship and the non-geometric encoding information of each node; weight and sum the product of the second dependency relationship and the difference information based on the attention weight to obtain a weighted sum result corresponding to each node; and add the weighted sum result corresponding to each node to the geometric encoding information of each node to obtain geometric attention encoding information corresponding to each node.
[0282] In some embodiments, the number of times of the first encoding is multiple times, and the number of times of the first encoding is the same as the input dimension of the attention encoding, and the multiple times of the first encoding are respectively implemented by multiple equivariant graph neural networks, and the equivariant graph neural network includes a plurality of network layers.
[0283] The first geometric coding information includes geometric coding information of each node in a graph structure of the compound, and the first non-geometric coding information includes non-geometric coding information of each node in the graph structure of the compound, wherein the graph structure includes a plurality of nodes.
[0284] The data processing module 1332 is further configured to sequentially traverse the plurality of network layers, and perform the following processing for a current network layer in the traversal: in response to the current network layer being a first layer, input geometric information and non-geometric information of each node in a graph structure of the compound into the current network layer; in response to the current network layer being a second layer and subsequent network layers, obtain message encoding between each node and a corresponding neighbor node in the current network layer; determine, by using the message encoding and geometric coding information output by a previous network layer, geometric coding information corresponding to each node output by the current network layer; and determine, by using the message encoding and non-geometric coding information output by the previous network layer, non-geometric coding information corresponding to each node output by the current network layer.
[0285] In some embodiments, the property prediction module 1333 is further configured to perform first non-linear mapping processing on the non-geometric splicing information to obtain non-geometric mapping information; perform second non-linear mapping processing on the non-geometric mapping information to obtain a probability value of each molecular property corresponding to the compound; and take a molecular property with the highest probability value as the molecular property of the compound.
[0286] The following continues to illustrate an exemplary structure of the implementation of the training device 134 of the molecular property prediction model as a software module, in some embodiments, as shown in Figure 2B The software module stored in the training device 134 of the molecular property prediction model in the memory 130-2 can include:
[0287] The data acquisition module 1341 is configured to obtain a training data set, wherein the training data set includes a plurality of molecular graph data, and the molecular graph data includes geometric information and non-geometric information.
[0288] The noise adding module 1342 is configured to perform noise adding processing on each of the molecular graph data to obtain noise-added molecular graph data of each of the molecular graph data.
[0289] The data processing module 1343 is configured to perform first encoding on each of the noise-added molecular graph data by using the initialized molecular property prediction model to obtain first geometric coding information and first non-geometric coding information corresponding to each of the noise-added molecular graph data.
[0290] In some embodiments, the data processing module 1343 is further configured to perform, by the molecular property prediction model, attention encoding based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the noisy molecular graph data, to obtain non-geometric attention encoding information corresponding to each of the noisy molecular graph data.
[0291] The training module 1344 is configured to perform a first training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the noisy molecular graph data, to obtain a first molecular property prediction model.
[0292] In some embodiments, the noise adding module 1342 is further configured to split each of the molecular graph data into a plurality of sub-graph data, obtain random noise data of a preset intensity, and superimpose the random noise data into geometric information of each of the sub-graph data to obtain noisy molecular graph data of each of the molecular graph data.
[0293] In some embodiments, the molecular graph data includes a plurality of nodes, and a node connected to each of the nodes is a neighbor node corresponding to the node; and the noise adding module 1342 is further configured to perform the following processing for each of the molecular graph data: selecting a node from the plurality of nodes included in the molecular graph as a starting node, and adding the starting node to a sub-graph node set; iteratively performing the following processing: selecting a neighbor node from the neighbor nodes of the starting node, and adding the selected neighbor node to the sub-graph node set; and in response to a number of nodes in the sub-graph node set reaching a preset node value, taking the sub-graph node set as a sub-graph data corresponding to the molecular graph data.
[0294] In some embodiments, the training module 1344 is further configured to perform energy prediction on the non-geometric attention encoding information corresponding to each of the noisy molecular graph data to obtain energy data corresponding to each of the noisy molecular graph data; perform partial derivation on geometric information of the molecular graph data through the energy data to obtain force data corresponding to each of the noisy molecular graph data; obtain a first loss value based on the force data and noise gradient corresponding to the noisy molecular graph data; obtain first gradient information through the first loss value, update parameters of the initialized molecular property prediction model according to the first gradient information, and obtain a first molecular property prediction model.
[0295] In some embodiments, the training module 1344 is further used to perform a first encoding on each of the molecular graph data through the first molecular property prediction model to obtain first geometric encoding information and first non-geometric encoding information corresponding to each of the molecular graph data; perform dependency-based attention encoding based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the molecular graph data through the first molecular property prediction model to obtain non-geometric attention encoding information corresponding to each of the molecular graph data; perform a second training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the molecular graph data to obtain a second molecular property prediction model.
[0296] In some embodiments, a preset number of molecular graph data among the multiple molecular graph data carry molecular property labels, the molecular graph data carrying the molecular property labels are labeled molecular graph data, and the molecular graph data not carrying the molecular property labels are unlabeled molecular graph data; the training module 1344 is also used to perform nonlinear mapping on the non-geometric attention encoding information of each molecular graph data with a molecular property label to obtain a first predicted molecular property of each molecular graph data with a molecular property label; determine a second loss value through the first predicted molecular property and the molecular property label; obtain second gradient information through the second loss value, update the parameters of the first molecular property prediction model according to the second gradient information, and obtain a second molecular property prediction model.
[0297] In some embodiments, the training module 1344 is further used to obtain, through the second molecular property prediction model, a second predicted molecular property of the molecular graph data without a molecular property label among the multiple molecular graph data; and use the second predicted molecular property that exceeds a preset prediction confidence threshold as a molecular property label for the molecular graph data without a molecular property label to form new molecular graph data; wherein the new molecular graph data is used to re-execute the second training task for molecular property prediction.
[0298] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the molecular property prediction method described in the present invention or the molecular property prediction model training method described in the present invention.
[0299] The embodiment of the present application provides a computer readable storage medium storing computer executable instructions, wherein the computer executable instructions or computer programs are stored, and when the computer executable instructions or computer programs are executed by a processor, the processor executes a prediction method of molecular properties provided by the embodiment of the present application or a training method of a molecular property prediction model provided by the embodiment of the present application, for example, as shown in the following table. Figure 3A The prediction method of molecular properties or Figure 4A The training method of the molecular property prediction model.
[0300] In some embodiments, the computer readable storage medium can be RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM memory, etc.; and can also be various devices including one or any combination of the above storage.
[0301] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.
[0302] As an example, the computer executable instructions can but not necessarily correspond to files in a file system, can be stored in a part of a file storing other programs or data, for example, stored in one or more scripts in a Hyper Text Marku p Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperative files (for example, files storing one or more modules, subroutines or code parts).
[0303] As an example, the computer executable instructions can be deployed to be executed on one electronic device, or executed on multiple electronic devices located in one place, or executed on multiple electronic devices distributed in multiple places and interconnected through a communication network.
[0304] In summary, by the embodiment of the present application, the geometric information and non-geometric information of the compound are encoded, the first geometric encoding information and the first non-geometric encoding information are jointly integrated into the attention mechanism based on the dependency relationship, the generation of the non-geometric attention encoding information is affected by the first geometric encoding information, the geometric information of the molecule is included in the non-geometric attention encoding information, that is, the local information brought by the chemical bond of the molecule is included, the non-geometric attention encoding information is spliced with the non-geometric information to obtain non-geometric splicing information, the loss of information is further reduced, and finally the molecular properties of the compound are obtained through the non-geometric splicing information, so as to achieve the beneficial effect of improving the accuracy of the prediction of the molecular properties.
[0305] The above merely describes the embodiments of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and scope of the present application shall be included in the protection scope of the present application.
Claims
1. A method of predicting molecular properties, characterized by, The method comprises: obtaining geometric information and non-geometric information of a compound; performing first encoding based on the geometric information and the non-geometric information to obtain first geometric encoding information and first non-geometric encoding information; performing attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information to obtain non-geometric attention encoding information; splicing the non-geometric attention encoding information and the non-geometric information to obtain non-geometric spliced information; performing non-linear mapping based on the non-geometric spliced information to obtain a molecular property of the compound.
2. The method of claim 1, wherein, The first geometric encoding information comprises geometric encoding information of each node in a graph structure of the compound, and the first non-geometric encoding information comprises non-geometric encoding information of each node in the graph structure of the compound, wherein the graph structure comprises a plurality of nodes; The method comprises: determining difference information between the geometric encoding information of each node and the geometric encoding information of a neighbor node, and determining a first dependency relationship and a second dependency relationship between each node and the neighbor node based on the difference information and the non-geometric information of the neighbor node; determining an attention weight based on the first dependency relationship and the non-geometric encoding information of each node; performing weighted summation on the second dependency relationship based on the attention weight to obtain non-geometric attention encoding information corresponding to each node.
3. The method of claim 2, wherein, The method comprises: performing non-linear mapping on the non-geometric encoding information of each node to obtain a query vector encoding corresponding to each node; obtaining a dot product result of the query vector encoding and the first dependency relationship, and normalizing the dot product result to obtain the attention weight.
4. The method of claim 2, wherein: the number of times of the first encoding and the number of times of determining the difference information are both multiple and the same; The method comprises: after the first encoding, determining difference information between the geometric encoding information of each node and the geometric encoding information of a corresponding neighbor node, and splicing the difference information and the non-geometric information of the neighbor node to obtain first spliced information; performing non-linear mapping on the first spliced information to obtain the first dependency relationship; The method comprises: after the first encoding, determining difference information between the geometric encoding information of each node and the geometric encoding information of a corresponding neighbor node, and splicing the difference information and the non-geometric information of the neighbor node to obtain first spliced information; performing non-linear mapping on the first spliced information to obtain the first dependency relationship; The second splicing information is nonlinearly mapped to obtain the second dependency relationship.
5. The method of claim 1, wherein, The non-geometric attention encoding information and the geometric attention encoding are second encoded based on the dependency relationship to obtain second non-geometric encoding information. The first geometric encoding information and the first non-geometric encoding information are based on the dependency relationship to obtain geometric attention encoding information. The non-geometric attention encoding information and the geometric attention encoding are second encoded based on the dependency relationship to obtain second non-geometric encoding information. The second non-geometric encoding information and the non-geometric information are spliced to obtain non-geometric splicing information.
6. The method of claim 5, wherein, The first geometric encoding information includes geometric encoding information of each node in the graph structure of the compound, and the first non-geometric encoding information includes non-geometric encoding information of each node in the graph structure of the compound, wherein the graph structure includes a plurality of nodes. The first geometric encoding information and the first non-geometric encoding information are based on the dependency relationship to obtain geometric attention encoding information, including: Determine the difference information between the geometric encoding information of each node and the geometric encoding information of the corresponding neighbor node, and determine the first dependency relationship and the second dependency relationship between each node and the neighbor node based on the difference information and the non-geometric information of the neighbor node; Determine the attention weight based on the first dependency relationship and the non-geometric encoding information of each node; The product of the second dependency relationship and the difference information is weighted and summed based on the attention weight to obtain a weighted sum result corresponding to each node; The weighted sum result corresponding to each node is added to the geometric encoding information of each node to obtain geometric attention encoding information corresponding to each node.
7. The method of claim 1, wherein, The number of times of the first encoding is multiple, and the number of times of the first encoding is the same as the input dimension of the attention encoding, and the multiple first encodings are respectively implemented by multiple equivariant graph neural networks, and the equivariant graph neural network includes multiple network layers. The first geometric encoding information includes geometric encoding information of each node in the graph structure of the compound, and the first non-geometric encoding information includes non-geometric encoding information of each node in the graph structure of the compound, wherein the graph structure includes a plurality of nodes. The first geometric encoding information and the first non-geometric encoding information are based on the dependency relationship to obtain geometric attention encoding information, including: Sequentially traverse the plurality of network layers, and perform the following processing for the current network layer traversed: In response to the current network layer being the first layer, input the geometric information and the non-geometric information of each node in the graph structure of the compound into the current network layer; In response to the current network layer being the second layer and the network layer thereafter, obtain the message encoding between each node and the corresponding neighbor node in the current network layer; Determine the geometric encoding information corresponding to each node output by the current network layer through the message encoding and the geometric encoding information output by the previous network layer; Determine non-geometric encoding information corresponding to each of the nodes of the current network layer output based on the message encoding and the non-geometric encoding information output by the previous network layer.
8. The method according to any one of claims 1 to 7, characterized in that, The non-linear mapping based on the non-geometric splicing information obtains the molecular property of the compound, including: Performing first non-linear mapping processing on the non-geometric splicing information to obtain non-geometric mapping information; Performing second non-linear mapping processing on the non-geometric mapping information to obtain a probability value of each molecular property corresponding to the compound; The molecular property with the highest probability value is taken as the molecular property of the compound.
9. A method for training a molecular property prediction model, characterized in that: The method comprises: Obtain a training data set, wherein the training data set comprises a plurality of molecular graph data, and the molecular graph data comprises geometric information and non-geometric information; Perform noise adding processing on each of the molecular graph data to obtain noise-added molecular graph data of each of the molecular graph data; Perform first encoding on each of the noise-added molecular graph data by using the initialized molecular property prediction model to obtain first geometric encoding information and first non-geometric encoding information corresponding to each of the noise-added molecular graph data; Perform attention encoding based on the dependency relationship based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the noise-added molecular graph data by using the molecular property prediction model to obtain non-geometric attention encoding information corresponding to each of the noise-added molecular graph data; Perform a first training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the noise-added molecular graph data to obtain a first molecular property prediction model.
10. The method of claim 9, wherein, The noise adding processing on each of the molecular graph data to obtain noise-added molecular graph data of each of the molecular graph data comprises: Split each of the molecular graph data into a plurality of sub-graph data; Obtain random noise data of a preset intensity; Superimpose the random noise data into the geometric information of each of the sub-graph data to obtain noise-added molecular graph data of each of the molecular graph data.
11. The method of claim 10, wherein: The molecular graph data comprises a plurality of nodes, and a node connected to each of the nodes is a neighbor node corresponding to the node; The splitting of each of the molecular graph data into a plurality of sub-graph data comprises: For each of the molecular graph data, perform the following processing: Select a node from the plurality of nodes included in the molecular graph as a starting node, and add the starting node to a sub-graph node set; Iteratively perform the following processing: select a neighbor node from the neighbor nodes of the starting node, and add the selected neighbor node to the sub-graph node set; In response to the number of nodes in the sub-graph node set reaching a preset node value, take the sub-graph node set as a sub-graph data corresponding to the molecular graph data.
12. The method of claim 9, wherein, The performing of a first training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the noise-added molecular graph data to obtain a first molecular property prediction model comprises: perform energy prediction on the non-geometric attention encoding information corresponding to each of the noisy molecular graph data, to obtain energy data corresponding to each of the noisy molecular graph data; perform partial derivation on the geometric information of the molecular graph data through the energy data, to obtain force data corresponding to each of the noisy molecular graph data; obtain a first loss value based on the force data and noise gradient corresponding to the noisy molecular graph data; obtain first gradient information through the first loss value, and update parameters of the initialized molecular property prediction model according to the first gradient information, to obtain a first molecular property prediction model.
13. The method of claim 9, wherein, After the first molecular property prediction model is obtained, the method further includes: perform first encoding on each of the molecular graph data through the first molecular property prediction model, to obtain first geometric encoding information and first non-geometric encoding information corresponding to each of the molecular graph data; perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the molecular graph data through the first molecular property prediction model, to obtain non-geometric attention encoding information corresponding to each of the molecular graph data; perform a second training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the molecular graph data, to obtain a second molecular property prediction model.
14. The method of claim 13, wherein a preset number of molecular graph data in the plurality of molecular graph data carry molecular property labels, the molecular graph data carrying the molecular property labels are labeled molecular graph data, and the molecular graph data not carrying the molecular property labels are unlabeled molecular graph data; the performing of the second training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the molecular graph data to obtain the second molecular property prediction model includes: performing nonlinear mapping on the non-geometric attention encoding information of each of the molecular graph data carrying the molecular property labels to obtain first predicted molecular properties of each of the molecular graph data carrying the molecular property labels; determining a second loss value through the first predicted molecular properties and the molecular property labels; obtaining second gradient information through the second loss value, and updating parameters of the first molecular property prediction model according to the second gradient information to obtain the second molecular property prediction model.
15. The method of claim 14, wherein, After the second molecular property prediction model is obtained, the method further includes: obtaining second predicted molecular properties of the molecular graph data not carrying the molecular property labels in the plurality of molecular graph data through the second molecular property prediction model; taking the second predicted molecular properties exceeding a preset prediction confidence threshold as the molecular property labels of the molecular graph data not carrying the molecular property labels to form new molecular graph data; wherein the new molecular graph data are used to re-perform the second training task for molecular property prediction.
16. A device for predicting molecular properties, characterized by The device includes: a data acquisition module configured to acquire geometric information and non-geometric information of a compound; a data processing module, configured to perform first encoding based on the geometric information and the non-geometric information to obtain first geometric encoding information and first non-geometric encoding information; the data processing module is further configured to perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information to obtain non-geometric attention encoding information; the data processing module is further configured to splice the non-geometric attention encoding information and the non-geometric information to obtain non-geometric spliced information; a property prediction module, configured to perform non-linear mapping based on the non-geometric spliced information to obtain a molecular property of the compound. 17.A device for training a molecular property prediction model, comprising: The device comprises: a data acquisition module, configured to acquire a training data set, wherein the training data set comprises a plurality of molecular graph data, and the molecular graph data comprises geometric information and non-geometric information; a noise adding module, configured to perform noise adding processing on each of the molecular graph data to obtain noise-added molecular graph data corresponding to each of the molecular graph data; a data processing module, configured to perform first encoding on each of the noise-added molecular graph data by an initialized molecular property prediction model to obtain first geometric encoding information and first non-geometric encoding information corresponding to each of the noise-added molecular graph data; the data processing module is further configured to perform attention encoding based on a dependency relationship based on the first geometric encoding information and the first non-geometric encoding information corresponding to each of the noise-added molecular graph data by the molecular property prediction model to obtain non-geometric attention encoding information corresponding to each of the noise-added molecular graph data; a training module, configured to perform a first training task for molecular property prediction based on the non-geometric attention encoding information corresponding to each of the noise-added molecular graph data to obtain a first molecular property prediction model.
18. An electronic device, comprising: The electronic device comprises: a memory, configured to store computer executable instructions; a processor, configured to execute the computer executable instructions stored in the memory to implement the prediction method of the molecular property in any one of claims 1 to 8 or the training method of the molecular property prediction model in any one of claims 9 to 15.
19. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program comprise the steps of claim 18. The computer executable instructions or computer program are executed by the processor to implement the prediction method of the molecular property in any one of claims 1 to 8 or the training method of the molecular property prediction model in any one of claims 9 to 15.
20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, The computer executable instructions or computer program are executed by the processor to implement the prediction method of the molecular property in any one of claims 1 to 8 or the training method of the molecular property prediction model in any one of claims 9 to 15.