Method, device, storage medium and computer equipment for predicting properties of drug molecules

CN114613450BActive Publication Date: 2025-09-19PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210231663.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-09-19
Estimated Expiration
2042-03-09

Smart Images

  • Figure CN114613450B_ABST
    Figure CN114613450B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, storage medium, and computer equipment for predicting the properties of drug molecules. The method includes: obtaining a drug molecule to be predicted, and performing modal conversion on the molecular structure of the drug molecule to obtain a multimodal drug molecular structure, wherein the multimodal drug molecular structure includes a drug molecular sequence, a drug molecular graph, a drug molecular image, and a drug molecular fingerprint; performing feature extraction on the multimodal drug molecular structure using a pre-trained multimodal feature extraction model to obtain a multimodal drug molecular feature vector; converting the multimodal drug molecular feature vector into a multimodal high-dimensional feature vector, and performing feature fusion on the multimodal high-dimensional feature vector to obtain a fused feature vector of the drug molecule; and inputting the fused feature vector of the drug molecule into a pre-trained drug molecular property prediction model to obtain a property prediction result of the drug molecule. The above method can improve the accuracy of drug molecular property prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and digital medical technology, and in particular to a method, device, storage medium and computer equipment for predicting the properties of drug molecules. Background Art

[0002] Drug discovery is the process of identifying new candidate compounds with potential therapeutic effects. Predicting the various properties of drug molecules is an essential step in the drug discovery process. Poor pharmacokinetic properties (absorption, distribution, metabolism, and excretion, ADME) and toxicity (T) are among the main reasons for drug development failure. Therefore, it is crucial to evaluate the ADMET properties of candidate drug molecules in the early stages of drug research.

[0003] Traditionally, drug properties have been verified through experiments. However, this approach is time-consuming and costly, and it is particularly difficult to achieve comprehensive and accurate predictions. Currently, a more common approach is to use machine learning to learn the data distribution representation of drug molecules and then apply this to unknown data to predict drug properties. However, existing drug prediction models fail to fully represent the characteristics of drug molecules, resulting in low prediction accuracy. Summary of the Invention

[0004] In view of this, the present application provides a method, apparatus, storage medium and computer equipment for predicting the properties of drug molecules, the main purpose of which is to solve the technical problem of inaccurate prediction of the properties of drug molecules.

[0005] According to a first aspect of the present invention, a method for predicting the properties of a drug molecule is provided, the method comprising:

[0006] Obtaining a drug molecule to be predicted and performing modal conversion on the molecular structure of the drug molecule to obtain a multimodal drug molecular structure, wherein the multimodal drug molecular structure includes at least two of a drug molecular sequence, a drug molecular graph, a drug molecular image, and a drug molecular fingerprint;

[0007] Through the pre-trained multimodal feature extraction model, the multimodal drug molecular structure is extracted to obtain the multimodal drug molecular feature vector;

[0008] Converting the multimodal drug molecule feature vectors into multimodal high-dimensional feature vectors, and performing feature fusion on the multimodal high-dimensional feature vectors to obtain a fused feature vector of the drug molecule;

[0009] The fused feature vector of the drug molecule is input into the pre-trained drug molecule property prediction model to obtain the property prediction result of the drug molecule.

[0010] According to a second aspect of the present invention, there is provided a device for predicting the properties of a drug molecule, the device comprising:

[0011] a modality conversion module, configured to obtain a drug molecule to be predicted and perform modality conversion on the molecular structure of the drug molecule to obtain a multimodal drug molecular structure, wherein the multimodal drug molecular structure includes at least two of a drug molecular sequence, a drug molecular graph, a drug molecular image, and a drug molecular fingerprint;

[0012] The feature extraction module is used to extract features of the multimodal drug molecular structure through a pre-trained multimodal feature extraction model to obtain a multimodal drug molecular feature vector;

[0013] A feature fusion module is used to convert the multimodal drug molecule feature vector into a multimodal high-dimensional feature vector, and perform feature fusion on the multimodal high-dimensional feature vector to obtain a fused feature vector of the drug molecule;

[0014] The property prediction module is used to input the fusion feature vector of the drug molecule into the pre-trained drug molecule property prediction model to obtain the property prediction results of the drug molecule.

[0015] According to a third aspect of the present invention, there is provided a storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-mentioned method for predicting the properties of drug molecules.

[0016] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for predicting the properties of drug molecules when executing the program.

[0017] The present invention provides a method, device, storage medium, and computer device for predicting the properties of drug molecules. These methods first convert the molecular structure of a drug molecule into a multimodal structure. Feature extraction is then performed on each modality of the drug molecule using a pretrained multimodal feature extraction model. Feature fusion is then performed on the drug molecule feature vectors from each modality. Finally, a prediction of the drug molecule's properties is obtained based on the fused feature vectors. This method can obtain a more comprehensive representation of drug molecule features, enabling more accurate and efficient prediction of drug molecule properties, effectively accelerating the speed and success rate of drug development and reducing the cost of drug molecule property prediction.

[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0020] Figure 1 A schematic diagram showing a flow chart of a method for predicting the properties of a drug molecule provided by an embodiment of the present invention;

[0021] Figure 2 A schematic diagram showing the operation flow of a method for predicting the properties of a drug molecule provided by an embodiment of the present invention is shown;

[0022] Figure 3 A schematic structural diagram of a device for predicting the properties of drug molecules provided by an embodiment of the present invention is shown;

[0023] Figure 4 A schematic diagram of the internal structure of a computer device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0024] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0025] In one embodiment, Figure 1 and Figure 2 As shown, a method for predicting the properties of drug molecules is provided, which is described by taking the application of the method to a computer device as an example, and includes the following steps:

[0026] 101. Obtain the drug molecule to be predicted, and perform modal conversion on the molecular structure of the drug molecule to obtain a multimodal drug molecular structure.

[0027] The multimodal drug molecular structure includes at least two of the following: a drug molecular sequence, a drug molecular graph, a drug molecular image, and a drug molecular fingerprint. In this embodiment, a drug molecular sequence refers to a drug molecular structure represented by a string, such as a SMILES expression, similar to a language sequence; a drug molecular graph refers to a drug molecular structure represented by a data structure diagram; a drug molecular image refers to a drug molecular structure represented by a planar image; and a drug molecular fingerprint refers to a drug molecular structure represented by a series of bit strings.

[0028] Specifically, the computer device can obtain the drug molecules to be predicted through a data interface or network, and then perform multiple rounds of modal conversion processing on the molecular structure of the drug molecules through the modal conversion method corresponding to each modal drug molecular structure to obtain a multi-modal drug molecular structure.

[0029] 102. Through the pre-trained multimodal feature extraction model, the multimodal drug molecular structure is feature extracted to obtain the multimodal drug molecular feature vector.

[0030] For each modality of drug molecular structure, a feature extraction method corresponding to the modality can be used to extract features from the drug molecular structure of each modality, thereby obtaining a multimodal drug molecular feature vector. In this embodiment, after feature extraction, at least two feature vectors can be obtained: a feature vector of a drug molecular sequence, a feature vector of a drug molecular graph, a feature vector of a drug molecular image, and a feature vector of a drug molecular fingerprint.

[0031] 103. The multimodal drug molecule feature vector is converted into a multimodal high-dimensional feature vector, and the multimodal high-dimensional feature vector is subjected to feature fusion to obtain a fused feature vector of the drug molecule.

[0032] Specifically, after obtaining the multimodal drug molecule feature vectors, the drug molecule feature vectors of different modalities can be first converted into high-dimensional feature expressions, and then fused in the intermediate layer of the model. Among them, the intermediate fusion can use a neural network to convert the multimodal drug molecule feature vectors into high-dimensional feature expressions (for example, 768 dimensions), and then obtain the commonalities of the different modal data in the high-dimensional space, so as to fuse the high-dimensional feature vectors of multiple modalities to obtain a more complete and sufficient fused feature vector of the drug molecule.

[0033] 104. Input the fused feature vector of the drug molecule into the pre-trained drug molecule property prediction model to obtain the property prediction result of the drug molecule.

[0034] Specifically, after obtaining the fused feature vector of the drug molecule, the fused feature vector of the drug molecule can be input into a pre-trained drug molecule property prediction model to obtain a drug molecule property prediction result. The drug molecule property prediction model can be trained using a machine learning model such as a neural network, and this embodiment does not specifically limit this.

[0035] The drug molecule property prediction method provided in this embodiment first converts the molecular structure of the drug molecule into a multimodal drug molecular structure. Feature extraction is then performed on the drug molecular structure of each modality using a pre-trained multimodal feature extraction model. Feature fusion is then performed on the drug molecular feature vectors of each modality. Finally, a drug molecule property prediction result is obtained based on the fused feature vectors of the drug molecule. This method can obtain a more comprehensive representation of drug molecule features, thereby enabling more accurate and efficient prediction of drug molecule properties, effectively accelerating the speed and success rate of drug research and development, and reducing the cost of drug molecule property prediction.

[0036] In one embodiment, the method of performing modal conversion on the molecular structure of the drug molecule in step 101 can be implemented by the following method: first, according to a predetermined molecular structure conversion rule, the molecular structure of the drug molecule is converted into a string format to obtain a drug molecule sequence. For example, the molecular structure of the drug molecule can be converted into a SMILES expression according to the SMILES conversion rule. Secondly, the atoms of the molecular structure of the drug molecule are converted into nodes of the drug molecule graph, and the chemical bonds of the molecular structure of the drug molecule are converted into edges of the drug molecule graph to obtain a drug molecule graph. In the drug molecule graph, a variety of attribute information or feature information of atoms or chemical bonds can also be added to enrich the feature information of the drug molecule graph. Furthermore, the molecular structure of the drug molecule can be converted into a two-dimensional image by taking a photo, screenshot, image conversion, etc. to obtain a drug molecule image. The image conversion method is relatively simple and will not be described in detail here. Finally, the structural features of the drug molecule can be extracted and encoded into a bit vector to obtain the drug molecule fingerprint. The drug molecule fingerprint is an abstract representation of a molecule. It can convert (encode) the drug molecule into a series of bit strings (i.e., bit vectors). Then, drug molecules can be easily compared. The more common method is to extract the structural features of the drug molecule and then hash it to generate a bit vector as the drug molecule fingerprint. It is understandable that there are many ways to convert the modality. Therefore, the conversion method is not limited to the above methods and can be selected according to the actual situation.

[0037] In one embodiment, the drug molecular fingerprint can specifically be an extended connectivity fingerprint. In this case, the drug molecular fingerprint extraction method may include the following steps: first, mark each atom in the molecular structure of the drug molecule with an identifier, and store the hash value of the identifier of each atom in a pre-established identifier set, then create a bond list for each atom, and store the bond order and identifier of the adjacent atoms of the atom in the bond list of each atom, and then use the hash value of the bond list of each atom as the updated identifier of the atom, and store the updated identifier of each atom in the identifier set, and finally extract all identifiers in the identifier set to obtain the drug molecular fingerprint.

[0038] In the above embodiment, the extended connectivity fingerprint is a circular fingerprint, whose definition requires setting a radius n (i.e., the number of iterations) and then calculating each atom identifier. The identifier is similar to the connectivity in the Morgan fingerprint and is ultimately determined by the environment with a radius of n. Among them, the algorithm for extending the connectivity fingerprint is as follows: first, create a set S to store the identifiers of all atoms, and then use a 32-bit integer to mark each atom, such as the Morgan algorithm or the CANGEN algorithm, and then hash them and add them to S. Furthermore, for each atom, create a "bond list" to store the information of the atoms around the atom. The list can be sorted according to the bond order (such as single bond, double bond, triple bond, etc.), and then sorted according to the size of the surrounding atom identifiers. Then fill the above list with the following information: the content is [n, identifier, bo1, aid1, bo2, aid2, ...], where n is the number of iterations, starting at 0, bo1 is the bond order of the first root bond, aid1 is the identifier of the atom connected to the first root bond, and so on. Then calculate the hash value of the feature list as the new identifier of the atom. If the newly calculated identifier is structurally different from that in S, it is added to S. Repeat this iteration until the end of the loop. In this embodiment, the drug molecular fingerprint can serve as a good supplement to the drug molecular structure of the other three modalities, more fully exploring and allowing the advantages of each modality to complement each other, thereby more effectively and accurately realizing the prediction of small molecule drug properties.

[0039] In one embodiment, the method of extracting features of drug molecular structures of each modality in step 102 can be implemented by the following method: extracting language structure features in the drug molecular sequence through the language model in the multimodal feature extraction model to obtain a feature vector of the drug molecular sequence; extracting atomic features and chemical bond features in the drug molecular graph through the graph neural network in the multimodal feature extraction model to obtain a feature vector of the drug molecular graph; extracting image features in the drug molecular image through the convolutional neural network in the multimodal feature extraction model to obtain a feature vector of the drug molecular image; extracting identifier features in the drug molecular fingerprint through the deep neural network in the multimodal feature extraction model to obtain a feature vector of the drug molecular fingerprint.

[0040] In the above embodiment, the language model can extract the hidden structural information and the correlation information between sequences in the drug molecule sequence. By splicing the extracted information together and then reducing the dimension after passing it through a fully connected layer, a low-dimensional dense feature vector expression of the drug molecule sequence can be obtained. The graph neural network can extract the features of the atomic nodes of the drug molecule graph and the chemical bond information between the atoms, thereby extracting the molecular-level features of the entire molecular compound. The convolutional neural network can extract image features of different levels in the drug molecule image, and can proceed layer by layer and extract all image features of the entire drug molecule image. The deep neural network can extract deep features in the drug molecule fingerprint, which can serve as a good supplement to the other three modal features, thereby realizing the complementary advantages between the modal features, and ultimately helping to improve the accuracy of the prediction of drug molecular properties.

[0041] In one embodiment, the method of feature fusion of drug molecule feature vectors of each modality in step 103 can be implemented by the following method: first, the multimodal drug molecule feature vector is converted into a multimodal high-dimensional feature vector of the same dimension, and then the multimodal high-dimensional feature vector is input into a pre-trained feature enhancement model to obtain the attention coefficient of the multimodal high-dimensional feature vector, and finally, according to the attention coefficient of the multimodal high-dimensional feature vector, the multimodal high-dimensional feature vector is weighted summed to obtain the fused feature vector of the drug molecule.

[0042] In the above embodiment, after obtaining the drug molecule feature vectors of multiple different modalities, the feature vectors of different modalities can be integrated through conventional operations, such as integration by splicing and weighted summation. However, conventional integration operations will result in no connection between the parameters. Therefore, this embodiment automatically performs adaptive operations on the fusion operation of the feature vectors through the network layer, and determines the contribution of each modality through a pre-trained feature enhancement model. In this embodiment, the attention mechanism can be used to obtain the attention coefficient of the feature vector of each modality, and thereby realize the fusion of multimodal information. Specifically, the high-dimensional feature vector F of each modality can be i Input into the trained attention network, and the attention weight of modality i is β i , through weighted accumulation, we can get the final fusion total feature F for drug molecular property prediction all , and its calculation expression is:

[0043]

[0044] β i =softmax(P i )

[0045]

[0046] Where: P i is the hidden unit state, and are weights and biases, β i is the normalized weight vector. In this way, the feature expression accuracy of the fused feature vector can be effectively improved, thereby improving the accuracy of drug molecular property prediction.

[0047] In one embodiment, the multimodal feature extraction model and the drug molecular property prediction model can be trained by the following method:

[0048] 201. Acquire multiple drug molecule samples, and perform modal conversion on the molecular structure of each drug molecule sample to obtain a multimodal drug molecular structure of each drug molecule sample.

[0049] The method for modal conversion of the molecular structure of a drug molecule sample is described above and will not be repeated here. In this embodiment, the multimodal drug molecule structure includes a drug molecule sequence, a drug molecule graph, a drug molecule image, and a drug molecule fingerprint. Each drug molecule sample includes a classification label for a predetermined property. That is, if the drug molecule property prediction model needs to predict the toxicity of a drug molecule, the predetermined property is toxicity, and the classification labels are toxic and non-toxic.

[0050] 202. Based on the multimodal drug molecular structures of multiple drug molecule samples, language models, graph neural networks, convolutional neural networks, deep neural networks and neural networks are constructed respectively.

[0051] Among them, the language model is used to extract the features of drug molecule sequences, the graph neural network is used to extract the features of drug molecule graphs, the convolutional neural network is used to extract the features of drug molecule images, the deep neural network is used to extract the features of drug molecule fingerprints, the attention network is used to fuse the high-dimensional features of each modality, and the neural network is used to classify the fused multimodal features, that is, to predict the properties of drug molecules.

[0052] 203. The multimodal drug molecular structures of multiple drug molecule samples are input into the language model, graph neural network, convolutional neural network and deep neural network respectively to obtain the multimodal drug molecular feature vector of each drug molecule sample.

[0053] 204. Convert the multimodal drug molecule feature vector of each drug molecule sample into a multimodal high-dimensional feature vector, and perform feature fusion on the multimodal high-dimensional feature vector of each drug molecule sample to obtain a fused feature vector of each drug molecule sample.

[0054] 205. Taking the fused feature vector of each drug molecule sample as input and the classification label of each drug molecule sample as output, the language model, graph neural network, convolutional neural network, deep neural network and neural network are trained synchronously and iteratively to obtain a multimodal feature extraction model and a drug molecule property prediction model.

[0055] In one embodiment, the above-mentioned model training process may further include the following steps: constructing an attention network, and then inputting the multimodal high-dimensional feature vector of each drug molecule sample into the attention network to obtain the attention coefficient of the multimodal high-dimensional feature vector of each drug molecule sample, and then performing weighted summation on the multimodal high-dimensional feature vector of each drug molecule sample according to the attention coefficient of the multimodal high-dimensional feature vector of each drug molecule sample to obtain a fused feature vector of each drug molecule sample, and finally iteratively training the attention network with the fused feature vector of each drug molecule sample as input and the classification label of each drug molecule sample as output to obtain a feature enhancement model.

[0056] In the above embodiment, the multimodal feature extraction model and the drug molecular property prediction model combine the advantages of multiple models such as language models, graph neural networks, convolutional neural networks, deep neural networks, attention networks and neural networks, and can accurately extract the feature information of each modality of drug molecules, and can accurately fuse and predict the feature vectors of each modality, thereby effectively improving the accuracy and generalization of drug molecular property prediction, increasing the speed and success rate of drug research and development, and reducing the cost of drug molecular property prediction.

[0057] Further, as Figure 1 、 Figure 2 The specific implementation of the method shown in this embodiment provides a device for predicting the properties of drug molecules, such as Figure 3 As shown, the device includes: a modality conversion module 31, a feature extraction module 32, a feature fusion module 33 and a property prediction module 34, wherein:

[0058] A modality conversion module 31 is configured to obtain a drug molecule to be predicted and perform modality conversion on the molecular structure of the drug molecule to obtain a multimodal drug molecular structure, wherein the multimodal drug molecular structure includes at least two of a drug molecular sequence, a drug molecular graph, a drug molecular image, and a drug molecular fingerprint;

[0059] The feature extraction module 32 can be used to extract features of the multimodal drug molecular structure through a pre-trained multimodal feature extraction model to obtain a multimodal drug molecular feature vector;

[0060] The feature fusion module 33 can be used to convert the multimodal drug molecule feature vector into a multimodal high-dimensional feature vector, and perform feature fusion on the multimodal high-dimensional feature vector to obtain a fused feature vector of the drug molecule;

[0061] The property prediction module 34 may be used to input the fusion feature vector of the drug molecule into a pre-trained drug molecule property prediction model to obtain a property prediction result of the drug molecule.

[0062] In a specific application scenario, the modal conversion module 31 can be used to convert the molecular structure of the drug molecule into a string format according to a predetermined molecular structure conversion rule to obtain a drug molecule sequence; convert the atoms of the molecular structure of the drug molecule into nodes of a drug molecule graph, and convert the chemical bonds of the molecular structure of the drug molecule into edges of the drug molecule graph to obtain a drug molecule graph; convert the molecular structure of the drug molecule into a two-dimensional image to obtain a drug molecule image; extract the structural features in the molecular structure of the drug molecule, and encode the structural features into a bit vector to obtain a drug molecule fingerprint.

[0063] In a specific application scenario, the drug molecule fingerprint is an extended connectivity fingerprint; the modal conversion module 31 can also be used to mark an identifier for each atom in the molecular structure of the drug molecule, and store the hash value of the identifier of each atom in a pre-established identifier set; create a bond list for each atom, and store the bond order and identifier of the adjacent atoms of the atom in the bond list of each atom; use the hash value of the bond list of each atom as the updated identifier of the atom, and store the updated identifier of each atom in the identifier set; extract all identifiers in the identifier set to obtain the drug molecule fingerprint.

[0064] In a specific application scenario, the feature extraction module 32 can be used to extract the language structure features in the drug molecule sequence through the language model in the multimodal feature extraction model to obtain the feature vector of the drug molecule sequence; extract the atomic features and chemical bond features in the drug molecule graph through the graph neural network in the multimodal feature extraction model to obtain the feature vector of the drug molecule graph; extract the image features in the drug molecule image through the convolutional neural network in the multimodal feature extraction model to obtain the feature vector of the drug molecule image; extract the identifier features in the drug molecule fingerprint through the deep neural network in the multimodal feature extraction model to obtain the feature vector of the drug molecule fingerprint.

[0065] In a specific application scenario, the feature fusion module 33 can be used to convert the multimodal drug molecule feature vector into a multimodal high-dimensional feature vector of the same dimension; input the multimodal high-dimensional feature vector into a pre-trained feature enhancement model to obtain the attention coefficient of the multimodal high-dimensional feature vector; and perform weighted summation on the multimodal high-dimensional feature vector according to the attention coefficient of the multimodal high-dimensional feature vector to obtain the fused feature vector of the drug molecule.

[0066] In a specific application scenario, the device further includes a model training module 35, which can be specifically used to obtain multiple drug molecule samples and perform modal conversion on the molecular structure of each drug molecule sample to obtain a multimodal drug molecular structure of each drug molecule sample, wherein each drug molecule sample contains a classification label of a predetermined property; according to the multimodal drug molecular structures of the multiple drug molecule samples, a language model, a graph neural network, a convolutional neural network, a deep neural network and a neural network are constructed respectively; the multimodal drug molecular structures of the multiple drug molecule samples are input into the language model, the graph neural network, the convolutional neural network and the deep neural network respectively. The multimodal drug molecule feature vector of each drug molecule sample is obtained by using a neural network and a deep neural network; the multimodal drug molecule feature vector of each drug molecule sample is converted into a multimodal high-dimensional feature vector, and the multimodal high-dimensional feature vector of each drug molecule sample is feature fused to obtain a fused feature vector of each drug molecule sample; with the fused feature vector of each drug molecule sample as input and the classification label of each drug molecule sample as output, the language model, graph neural network, convolutional neural network, deep neural network and neural network are synchronously iteratively trained to obtain a multimodal feature extraction model and a drug molecule property prediction model.

[0067] In a specific application scenario, the model training module 35 can also be used to construct an attention network; the multimodal high-dimensional feature vector of each drug molecule sample is input into the attention network to obtain the attention coefficient of the multimodal high-dimensional feature vector of each drug molecule sample; according to the attention coefficient of the multimodal high-dimensional feature vector of each drug molecule sample, the multimodal high-dimensional feature vector of each drug molecule sample is weighted and summed to obtain a fused feature vector of each drug molecule sample; with the fused feature vector of each drug molecule sample as input and the classification label of each drug molecule sample as output, the attention network is iteratively trained to obtain a feature enhancement model.

[0068] It should be noted that for other corresponding descriptions of the functional units involved in the device for predicting the properties of drug molecules provided in this embodiment, please refer to Figure 1 、 Figure 2 The corresponding description in will not be repeated here.

[0069] Based on the above Figure 1 、 Figure 2 The method shown in FIG. 1 is a method for performing the above-mentioned operation. Accordingly, this embodiment further provides a storage medium on which a computer program is stored. When the program is executed by a processor, the above-mentioned Figure 1 、 Figure 2 The method for predicting the properties of drug molecules is shown.

[0070] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product. The software product to be identified can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each implementation scenario of the present application.

[0071] Based on the above Figure 1 、 Figure 2 The method shown, and Figure 3 In order to achieve the above-mentioned purpose, the device for predicting the properties of drug molecules shown in the embodiment is as follows: Figure 4 As shown, this embodiment also provides a computer device for predicting the properties of drug molecules, which can be a personal computer, server, smart phone, tablet computer, smart watch, or other network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store computer programs and operating system; the processor is used to execute the computer program to achieve the above-mentioned Figure 1 、 Figure 2 The method shown.

[0072] Optionally, the computer device may further include an internal memory, a communication interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, a display, an input device such as a keyboard, etc. Optionally, the communication interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Wi-Fi interface), etc.

[0073] Those skilled in the art will understand that the computer device structure for identifying an operation action provided in this embodiment does not constitute a limitation on the computer device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0074] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the computer device hardware and the software resources to be identified, supporting the execution of the information processing program and other software and / or programs to be identified. The network communication module is used to enable communication between components within the storage medium and with other hardware and software in the information processing computer device.

[0075] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform, or it can be implemented by hardware. By applying the technical solution of the present application, the molecular structure of the drug molecule is first converted into a multimodal drug molecular structure, and then the drug molecular structure of each mode of the drug molecule is feature extracted by a pre-trained multimodal feature extraction model, and then the drug molecular feature vectors of each mode are feature fused, and finally the property prediction results of the drug molecule are obtained based on the fused feature vectors of the drug molecule. Compared with the prior art, the above method can obtain a more comprehensive representation of drug molecular features, so that the properties of the drug molecule can be predicted more accurately and effectively, effectively accelerating the speed and success rate of drug research and development, and reducing the cost of drug molecule property prediction.

[0076] Those skilled in the art will understand that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required to implement the present application. Those skilled in the art will understand that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the implementation scenario description, or can be changed accordingly and located in one or more devices different from the implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0077] The serial numbers of the above application are for descriptive purposes only and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure only discloses several specific implementation scenarios of the present application, but the present application is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present application.

Claims

1. A method for predicting the properties of drug molecules, characterized in that: The method comprises: Obtaining a drug molecule to be predicted, and performing modal conversion on the molecular structure of the drug molecule to obtain a multimodal drug molecular structure, wherein the multimodal drug molecular structure consists of a drug molecular sequence, a drug molecular graph, a drug molecular image, and a drug molecular fingerprint, wherein the drug molecular graph refers to a drug molecular structure represented by a data structure graph, and the drug molecular image refers to a drug molecular structure represented by a planar image; Extracting features from the multimodal drug molecular structure using a pre-trained multimodal feature extraction model to obtain a multimodal drug molecular feature vector; The multimodal drug molecule feature vector is converted into a multimodal high-dimensional feature vector, and the multimodal high-dimensional feature vector is subjected to feature fusion to obtain a fused feature vector of the drug molecule; specifically comprising: converting the multimodal drug molecule feature vector into a multimodal high-dimensional feature vector of the same dimension; inputting the multimodal high-dimensional feature vector into a pre-trained feature enhancement model to obtain an attention coefficient of the multimodal high-dimensional feature vector; and performing weighted summation on the multimodal high-dimensional feature vector according to the attention coefficient of the multimodal high-dimensional feature vector to obtain a fused feature vector of the drug molecule; The fused feature vector of the drug molecule is input into a pre-trained drug molecule property prediction model to obtain a property prediction result of the drug molecule.

2. The method according to claim 1, characterized in that The modal conversion of the molecular structure of the drug molecule to obtain a multimodal drug molecular structure includes: According to a predetermined molecular structure conversion rule, the molecular structure of the drug molecule is converted into a string format to obtain the drug molecule sequence; Converting atoms in the molecular structure of the drug molecule into nodes of a drug molecule graph, and converting chemical bonds in the molecular structure of the drug molecule into edges of the drug molecule graph, to obtain the drug molecule graph; Converting the molecular structure of the drug molecule into a two-dimensional image to obtain the drug molecule image; Structural features in the molecular structure of the drug molecule are extracted, and the structural features are encoded into a bit vector to obtain the drug molecule fingerprint.

3. The method according to claim 2, characterized in that The drug molecule fingerprint is an extended connectivity fingerprint; the extraction of structural features from the molecular structure of the drug molecule and encoding the structural features into a bit vector to obtain the drug molecule fingerprint includes: Marking each atom in the molecular structure of the drug molecule with an identifier, and storing a hash value of the identifier of each atom in a pre-established identifier set; creating a bond list for each of the atoms, and storing the bond orders and identifiers of the neighboring atoms of the atom in the bond list of each atom; Using the hash value of the key list of each atom as the updated identifier of the atom, and storing the updated identifier of each atom in the identifier set; All identifiers in the identifier set are extracted to obtain the drug molecular fingerprint.

4. The method according to claim 1, wherein The multimodal feature extraction model is used to extract features from the multimodal drug molecular structure to obtain a multimodal drug molecular feature vector, including: Extracting language structure features from the drug molecule sequence using the language model in the multimodal feature extraction model to obtain a feature vector of the drug molecule sequence; Extracting atomic features and chemical bond features from the drug molecule graph through the graph neural network in the multimodal feature extraction model to obtain a feature vector of the drug molecule graph; Extracting image features from the drug molecule image using a convolutional neural network in a multimodal feature extraction model to obtain a feature vector of the drug molecule image; The identifier features in the drug molecular fingerprint are extracted through a deep neural network in a multimodal feature extraction model to obtain a feature vector of the drug molecular fingerprint.

5. The method according to any one of claims 1 to 4, characterized in that The training method of the multimodal feature extraction model and the drug molecular property prediction model includes: Acquiring multiple drug molecule samples and performing modal conversion on the molecular structure of each drug molecule sample to obtain a multimodal drug molecular structure of each drug molecule sample, wherein each of the drug molecule samples includes a classification label of a predetermined property; According to the multimodal drug molecular structures of the plurality of drug molecule samples, respectively constructing a language model, a graph neural network, a convolutional neural network, a deep neural network and a neural network; Inputting the multimodal drug molecular structures of the multiple drug molecule samples into the language model, the graph neural network, the convolutional neural network, and the deep neural network respectively to obtain a multimodal drug molecular feature vector for each drug molecule sample; Converting the multimodal drug molecule feature vector of each drug molecule sample into a multimodal high-dimensional feature vector, and performing feature fusion on the multimodal high-dimensional feature vector of each drug molecule sample to obtain a fused feature vector of each drug molecule sample; Taking the fused feature vector of each drug molecule sample as input and the classification label of each drug molecule sample as output, the language model, graph neural network, convolutional neural network, deep neural network and neural network are synchronously iteratively trained to obtain the multimodal feature extraction model and drug molecule property prediction model.

6. The method according to claim 5, characterized in that The step of performing feature fusion on the multimodal high-dimensional feature vector of each drug molecule sample to obtain a fused feature vector of each drug molecule sample includes: Build an attention network; Inputting the multimodal high-dimensional feature vector of each drug molecule sample into the attention network to obtain an attention coefficient of the multimodal high-dimensional feature vector of each drug molecule sample; performing weighted summation on the multimodal high-dimensional feature vectors of each drug molecule sample according to the attention coefficient of the multimodal high-dimensional feature vector of each drug molecule sample to obtain a fused feature vector of each drug molecule sample; The method further comprises: The attention network is iteratively trained with the fused feature vector of each drug molecule sample as input and the classification label of each drug molecule sample as output to obtain a feature enhancement model.

7. A device for predicting the properties of drug molecules, characterized in that: The device comprises: a modality conversion module, configured to obtain a drug molecule to be predicted and perform modality conversion on the molecular structure of the drug molecule to obtain a multimodal drug molecular structure, wherein the multimodal drug molecular structure is composed of a drug molecular sequence, a drug molecular graph, a drug molecular image, and a drug molecular fingerprint. The drug molecular graph refers to a drug molecular structure represented by a data structure graph, and the drug molecular image refers to a drug molecular structure represented by a planar image; A feature extraction module is used to extract features of the multimodal drug molecular structure using a pre-trained multimodal feature extraction model to obtain a multimodal drug molecular feature vector; A feature fusion module is used to convert the multimodal drug molecule feature vector into a multimodal high-dimensional feature vector, and perform feature fusion on the multimodal high-dimensional feature vector to obtain a fused feature vector of the drug molecule; specifically comprising: converting the multimodal drug molecule feature vector into a multimodal high-dimensional feature vector of the same dimension; inputting the multimodal high-dimensional feature vector into a pre-trained feature enhancement model to obtain an attention coefficient of the multimodal high-dimensional feature vector; and performing weighted summation on the multimodal high-dimensional feature vector according to the attention coefficient of the multimodal high-dimensional feature vector to obtain a fused feature vector of the drug molecule; The property prediction module is used to input the fusion feature vector of the drug molecule into a pre-trained drug molecule property prediction model to obtain a property prediction result of the drug molecule.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Molecular graph generation from structural features using an artificial neural network

    WO2020243440A1

  • Drug molecular property determining method and device, and storage medium

    WO2022022173A1