A method and device for predicting the characteristics of molecular drugs

By generating a two-dimensional feature pixel matrix of molecules, the problem of low molecular feature representation efficiency in the prior art is solved, more efficient prediction of molecular drug characteristics and better interpretability are achieved, and calculation speed and prediction accuracy are improved.

CN114708928BActive Publication Date: 2025-07-11ACEMAP BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210298949.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-07-11
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

Existing graph representation methods cannot fully utilize distributed computing resources to efficiently calculate molecular features to obtain molecular representations, resulting in insufficient accuracy in predicting molecular drug characteristics in deep learning.

Method used

By obtaining the high-dimensional feature distance of the training molecule from the molecular database, dimensionality reduction and network processing are performed, a two-dimensional feature pixel matrix is generated, and a two-dimensional matrix of representations of the molecule is generated by combining the characteristic calculation value vector of the prediction molecule, and the matrix is used to predict drug characteristics.

Benefits of technology

The efficiency of molecular feature representation learning is improved, better interpretability and higher prediction accuracy is achieved, parallel computing power is fully utilized, and the unified sorting ability of computing speed and molecular feature representation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708928B_ABST
    Figure CN114708928B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for predicting molecular drug properties. The method includes: obtaining training molecules and basic information of the training molecules from a molecular database, and using the basic information of the training molecules to obtain the high-dimensional feature distance D of the training molecules; after respectively performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecules, obtaining a two-dimensional feature pixel matrix GMap of the training molecules; obtaining a feature calculation value vector m of a prediction molecule, and assigning the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain a representation two-dimensional matrix MolMap of the prediction molecule; using the representation two-dimensional matrix MolMap of the prediction molecule as the molecular representation of the prediction molecule, and using the molecular representation of the prediction molecule to predict the drug properties of the prediction molecule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of molecular characteristics, and particularly relates to a method and device for predicting molecular drug properties. Background Art

[0002] The development of deep learning has gradually matured. With improved methods and increasingly rich data, the prediction of various molecular drug properties by deep learning has reached unprecedented accuracy. Successful deep learning algorithms depend on whether they can effectively learn the representation of the research target. Studying how to represent molecules so that computers can recognize them, and then using the molecular representation to predict the drug properties of molecules through deep learning methods is an objective that the industry has been exploring. In recent years, methods based on graph representation have achieved great success. However, graph representation still cannot fully utilize the capabilities of distributed computing, and there are still many directions for improvement in how to efficiently use distributed computing resources to calculate molecular features to obtain molecular representations. Summary of the Invention

[0003] The technical problem solved by the solution provided in the embodiments of the present invention is how to calculate molecular features to obtain a molecular representation, and then use the molecular representation to predict the drug properties of molecules.

[0004] A method for predicting molecular drug properties according to an embodiment of the present invention includes:

[0005] Obtaining training molecules and basic information of the training molecules from a molecular database, and using the basic information of the training molecules to obtain the high-dimensional feature distance D of the training molecules;

[0006] After respectively performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecules, obtaining the two-dimensional feature pixel matrix GMap of the training molecules;

[0007] Obtaining the feature calculation value vector m of the prediction molecule, and assigning the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain the representation two-dimensional matrix MolMap of the prediction molecule;

[0008] Using the representation two-dimensional matrix MolMap of the prediction molecule as the molecular representation of the prediction molecule, and using the molecular representation of the prediction molecule to predict the drug properties of the prediction molecule.

[0009] A device for predicting molecular drug properties according to an embodiment of the present invention includes:

[0010] A first acquisition module, configured to obtain training molecules and basic information of the training molecules from a molecular database, and use the basic information of the training molecules to obtain the high-dimensional feature distance D of the training molecules;

[0011] A processing module, configured to obtain the two-dimensional feature pixel matrix GMap of the training molecule after respectively performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecule;

[0012] A second acquisition module, configured to acquire the feature calculation value vector m of the prediction molecule and assign the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain the representation two-dimensional matrix MolMap of the prediction molecule;

[0013] A prediction module, configured to use the representation two-dimensional matrix MolMap of the prediction molecule as the molecular representation of the prediction molecule, and use the molecular representation of the prediction molecule to predict the drug properties of the prediction molecule.

[0014] According to the solution provided by the embodiment of the present invention, the following beneficial effects are achieved:

[0015] First, by using features at the molecular layer functional group level such as molecular descriptor features and fingerprint features as the high-dimensional feature space of molecular representation, the problem that long-distance structural features cannot be learned is avoided;

[0016] Second, by using the manifold transformation algorithm, the high-dimensional feature space of the molecule is mapped to a plane, and then linear optimization is performed for feature sorting to fix the order of the features. No matter how many molecules there are, there is a unified sorted feature space representation. The differences between different molecules can be intuitively reflected by the differences in the feature space numerical values, which has better interpretability;

[0017] Third, by making full use of the parallel computing ability, the descriptor features, fingerprint features and distance matrix of features of all molecules are calculated in advance, and only the calculated data needs to be loaded in the subsequent prediction tasks, which greatly improves the efficiency of molecular feature representation learning; Description of the Drawings

[0018] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to understand the present invention, and do not constitute an improper limitation to the present invention. In the drawings:

[0019] Figure 1 is a flowchart of a method for predicting the drug properties of a prediction molecule provided by an embodiment of the present invention;

[0020] Figure 2 is a schematic diagram of a device for predicting the drug properties of a prediction molecule provided by an embodiment of the present invention;

[0021] Figure 3 is a schematic diagram of the division of the dataset M and the dataset M provided by an embodiment of the present invention. Detailed implementation mode

[0022] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments described below are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0023] Figure 1 is a flowchart of a method for predicting the characteristics of molecular drugs provided by an embodiment of the present invention, as Figure 1 shown, including:

[0024] Step S101: Obtain training molecules and basic information of the training molecules from a molecular database, and use the basic information of the training molecules to obtain the high-dimensional feature distance D of the training molecules;

[0025] Step S102: After respectively performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecules, obtain the two-dimensional feature pixel matrix GMap of the training molecules;

[0026] Step S103: Obtain the feature calculation value vector m of the prediction molecule, and assign the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain the representation two-dimensional matrix Mol Map of the prediction molecule;

[0027] Step S104: Use the representation two-dimensional matrix Mol Map of the prediction molecule as the molecular representation of the prediction molecule, and use the molecular representation of the prediction molecule to predict the drug characteristics of the prediction molecule.

[0028] Among them, the obtaining of the high-dimensional feature distance D of the training molecules by using the basic information of the training molecules includes: respectively calculating N descriptor features and M fingerprint features of the training molecules according to the basic information of the training molecules; obtaining the high-dimensional descriptor feature distance D1 and high-dimensional fingerprint feature distance D2 of the training molecules by respectively calculating the descriptor feature distances between pairwise descriptor features among the N descriptor features of the training molecules and the fingerprint feature distances between pairwise fingerprint features among the M fingerprint features; where N and M are both positive integers.

[0029] Among them, after performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecule respectively, obtaining the two-dimensional feature pixel matrix GMap of the training molecule includes: performing dimensionality reduction on the high-dimensional descriptor feature distance D1 and the high-dimensional fingerprint feature distance D2 of the training molecule respectively to obtain the two-dimensional descriptor feature matrix MapD1 and the two-dimensional fingerprint feature matrix Map2 of the training molecule; performing grid processing on the two-dimensional descriptor feature matrix MapD1 and the two-dimensional fingerprint feature matrix Map2 of the training molecule respectively to obtain the two-dimensional descriptor feature pixel matrix GMap1 and the two-dimensional fingerprint feature pixel matrix GMap2 of the training molecule.

[0030] Among them, obtaining the characteristic calculation value vector m of the predicted molecule includes: obtaining the descriptor characteristic calculation value vector m1 according to the N descriptor characteristics of the predicted molecule; and obtaining the fingerprint characteristic calculation value vector m2 according to the M fingerprint characteristics of the predicted molecule.

[0031] Among them, obtaining the characteristic calculation value vector m of the predicted molecule and assigning the characteristic calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain the representation two-dimensional matrix Mol Map of the predicted molecule includes: assigning the descriptor characteristic calculation value vector m1 of the predicted molecule to the two-dimensional descriptor feature pixel matrix GMap1 to obtain the representation two-dimensional descriptor matrix Mol Map1 of the predicted molecule, and at the same time assigning the fingerprint characteristic calculation value vector m2 of the predicted molecule to the two-dimensional fingerprint feature pixel matrix GMap2 to obtain the representation two-dimensional fingerprint matrix Mol Map2 of the predicted molecule.

[0032] Among them, using the representation two-dimensional matrix Mol Map of the predicted molecule as the molecular representation of the predicted molecule includes: using the representation two-dimensional descriptor matrix Mol Map1 and the representation two-dimensional fingerprint matrix Mol Map2 of the predicted molecule as the molecular representation of the predicted molecule.

[0033] Figure 2 is a schematic diagram of a device for predicting the drug properties of a molecule provided by an embodiment of the present invention, as Figure 2As shown in the figure, it includes: a first acquisition module 201, configured to acquire training molecules and basic information of the training molecules from a molecular database, and use the basic information of the training molecules to obtain a high-dimensional feature distance D of the training molecules; a processing module 202, configured to obtain a two-dimensional feature pixel matrix GMap of the training molecules after respectively performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecules; a second acquisition module 203, configured to acquire a feature calculation value vector m of a prediction molecule, and assign the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain a representation two-dimensional matrix Mol Map of the prediction molecule; a prediction module 204, configured to use the representation two-dimensional matrix Mol Map of the prediction molecule as a molecular representation of the prediction molecule, and use the molecular representation of the prediction molecule to predict drug properties of the prediction molecule.

[0034] Among them, the first acquisition module 201 includes: a calculation unit, configured to respectively calculate N descriptor features and M fingerprint features of the training molecules according to the basic information of the training molecules; an acquisition unit, configured to obtain a high-dimensional descriptor feature distance D1 and a high-dimensional fingerprint feature distance D2 of the training molecules by respectively calculating descriptor feature distances between pairwise descriptor features among the N descriptor features of the training molecules and fingerprint feature distances between pairwise fingerprint features among the M fingerprint features; where N and M are both positive integers.

[0035] Among them, the processing module 202 includes: a dimensionality reduction processing unit, configured to respectively perform dimensionality reduction processing on the high-dimensional descriptor feature distance D1 and the high-dimensional fingerprint feature distance D2 of the training molecules to obtain a two-dimensional descriptor feature matrix MapD1 and a two-dimensional fingerprint feature matrix Map2 of the training molecules; a grid processing unit, configured to respectively perform grid processing on the two-dimensional descriptor feature matrix MapD1 and the two-dimensional fingerprint feature matrix Map2 of the training molecules to obtain a two-dimensional descriptor feature pixel matrix GMap1 and a two-dimensional fingerprint feature pixel matrix GMap2 of the training molecules.

[0036] Among them, the second acquisition module 203 is specifically configured to obtain a descriptor feature calculation value vector m1 according to the N descriptor features of the prediction molecule; and obtain a fingerprint feature calculation value vector m2 according to the M fingerprint features of the prediction molecule.

[0037] The technical solution of the present invention will be specifically described below

[0038] Step 1, data collection. Molecular screening and downloading are performed from public molecular databases to construct a local molecular database. Specifically, the molecular database contains 1 million screened molecules and their basic information. Based on this basic molecular information, molecular features are calculated, and then the representation of molecules is learned according to the molecular features.

[0039] a. Screen a certain number of molecules and download them to the local in multiple processes.

[0040] b. Perform deduplication according to the molecular SMILES strings and similarity. The similarity is calculated using the Tanimoto coefficient. Since the same molecule can have multiple forms of SMILES notations, the Tanimoto coefficient can further determine whether they are duplicate molecules. Specifically, the fingerprint features of each molecule are calculated in parallel and stored as one-hot encoded vectors. The Tanimoto coefficient is calculated using the formula: where a is the number of features of molecule A, b is the number of features of molecule B, and c is the number of common features between A and B. Calculate the Tanimoto coefficient matrix of each molecule with all molecules in parallel. When performing deduplication retrieval, only the molecules with a similarity of 1 need to be removed.

[0041] Step 2, molecular feature calculation. Calculate the descriptor features and fingerprint features of each molecule in a distributed and parallel manner. These features belong to the features of the molecular structure and physical and chemical properties, rather than the features calculated from the atomic level. Therefore, they can represent the local and overall molecular features and there is no problem of atomic distance limitation. Specifically, data splitting is adopted, and multiple processes are started to calculate simultaneously. After all calculations are completed, they are summarized into a complete molecular feature matrix.

[0042] The features of a molecule are composed of 1000 descriptor features and more than 15000 fingerprint features. The finally obtained molecules and their feature calculation results constitute the dataset M for distance matrix calculation, as Figure 3 shown in 201. Each row represents a molecule and its corresponding descriptor features (1000 columns) and fingerprint features (15000 columns), for a total of 16000 columns.

[0043] Step 3, distributed parallel calculation of the molecular feature distance matrix, that is, calculate the distances between pairwise molecular descriptor features and the distances between pairwise molecular fingerprint features to obtain the descriptor feature distance matrix and the fingerprint feature distance matrix.

[0044] Conventional distance matrix calculation only utilizes multi - process operations. The present invention introduces parallelized vector - based distance matrix calculation. Based on the fact that the speed of vector operations far exceeds that of loop - iterative calculations, the loop - iterative calculation logic and method are rewritten, significantly improving the calculation speed. Specifically, each row of the dataset M obtained in step 2 corresponds to 16,000 features of a molecule. Now, what needs to be calculated is the distance matrix between each feature and the remaining features. Let each column of the dataset M be a feature F, and its corresponding dimension is the number of molecules, i.e., the number of rows R of M. The distance between two features, such as the distance between feature F1 in the first column and feature F2 in the second column, refers to the formula:

[0045]

[0046] That is, the distance calculation between two vectors of dimension R. This result corresponds to a number on the distance matrix D, i.e., D21 = cos(F1,F2). The value of the second row and the first column of matrix D is cos(F1,F2). Due to the symmetric property of D, to save the amount of calculation, only the lower - triangular part of D needs to be calculated; for N features, the total amount of calculation of D is Finally, the feature distance matrix D obtained in this step is obtained.

[0047] In addition, the present invention transforms its batch calculation into distributed batch calculation. For example, the sizes of F1 and F2 are both 100w×1000, that is, 1 million molecules and 1000 features. Taking descriptor features as an example, there are actually 15,000 features, which is equivalent to data partitioning in a column - by - column manner. Refer to Figure 3 , 202. It is divided into a total of 15 matrices of 100w×1000. The 1000 features within one matrix can be calculated at one time, and the distance calculation between this matrix and another matrix is also completed at one time, without iterative vector calculation, greatly improving the speed. Finally, the results of each part of the calculation are concatenated in a certain order.

[0048] Step 4, based on the feature distance matrix D, using the manifold algorithm, the high - dimensional molecular features are reduced to a two - dimensional plane to obtain a feature two - dimensional matrix (picture); taking the t - SNE manifold algorithm as an example, t - SNE maps high - dimensional feature points to a distribution. The matrix D obtains the high - dimensional space distance between each feature. Then, feature points with close distances have a higher sampling probability, and feature points with far distances have a lower probability. To ensure that the absolute distance does not affect the distribution, probability normalization is also performed on the feature points. Let two points in the high - dimensional space be: x i ,x j , p j|i represents the center point as x i when x jThe probability of being its neighbor, x j The closer to x i The greater the probability. The formula is as follows:

[0049]

[0050] For different central points x i Its variance σ i Is also different; for the feature points x in the high-dimensional space i , x j , which is mapped to the low-dimensional space, here is the point y in the two-dimensional space i y j , and its probability distribution q j|i Is as follows:

[0051]

[0052] In order to make the points in the high-dimensional space maintain as similar a distribution as possible after being mapped to the low-dimensional space, that is, the points that were originally close are still close, and the points that were far apart are still far apart. Therefore, it is necessary to ensure that the two distributions are as similar as possible. The method used here to measure is to adopt the Ku l l back-Le i b l er D i vergence distance (KL divergence). The calculation formula is as follows:

[0053]

[0054] Measures the difference between distribution P and distribution Q. The smaller the KL, the smaller the difference. t-SNE uses the gradient descent method to solve to minimize the above KL value:

[0055]

[0056] Finally, the high-dimensional feature points can be obtained as the two-dimensional feature points after dimensionality reduction mapping;

[0057] Actually, for 1000 features of the descriptor, a distance matrix D1 is calculated and generated in the same way as in step 3. After step 4, the two-dimensional matrix Map1 after dimensionality reduction of the descriptor features is obtained, and its size is (1000×2), that is, the two-dimensional coordinate points corresponding to 1000 features. For 15000 features of the fingerprint, a distance matrix D2 (step 3) is calculated and generated, and after step 4, the two-dimensional matrix Map2 after dimensionality reduction of the fingerprint features is obtained, and its size is (15000×2), that is, the two-dimensional coordinate points corresponding to 15000 features;

[0058] Step 5, use the linear assignment algorithm to perform grid processing on Map1 and Map2 obtained in step 4. Taking Map1 as an example, first generate a grid, and the size is: After rounding, it is approximately 32×32 grids. The grids are standard grids with a value range between 0 and 1. Each grid has a standard coordinate. The abscissa contains 32 grids, and the length of each grid is 1 / 32. The ordinate contains 32 grids, and the length of each grid is 1 / 32. In this way, each grid corresponds to a coordinate, ensuring that the total number of grids is greater than or equal to the total number of features, which is 1000. Then, calculate the distance matrix DMap between the feature point coordinates in Map1 and all grid coordinates. This serves as the cost matrix for the linear assignment algorithm, that is, the cost (distance) of assigning a feature point to any grid. The purpose of the linear assignment algorithm is to assign all feature points to the grids in appropriate positions under the condition of the minimum cost, thereby forming a standard two-dimensional pixel matrix GMap1. The two-dimensional pixel matrix obtained in this way is suitable for various downstream deep learning tasks, especially the learning of convolutional neural networks. Similarly, for Map2, a standard two-dimensional pixel matrix GMap2 of 123×122 can be obtained.

[0059] Step 6, molecular representation learning. For a new molecular SMILES expression, after Step 2, the calculated value vector m1 of all descriptor features is obtained, and its size is 1×1000 (because there is only one molecule). The calculated value of the fingerprint feature m2, with a size of (1×15000). In Step 5, we have already obtained GMap1 and GMap2 and know which grid point each feature should correspond to. Then, directly assign the calculated values of m1 and m2 to the corresponding grid points of GMap1 and GMap2 to obtain the molecular representation two-dimensional matrices (images) MolMap1 and MolMap2, with sizes of 32×32 and 123×122 respectively. These two matrices are the molecular representations. Next, as the input of the convolutional neural network, through convolutional network, pooling, aggregation, and multi-linear layer network transformation, an output is obtained. For example, during network training, the training data is a batch of hepatotoxic molecules and the standard values measured in experiments. Then, compare the output value of this neural network with the standard hepatotoxicity, calculate the mean square error, and backpropagate the error to learn the correlation between the molecular representation and hepatotoxicity.

[0060] According to the solution provided by the embodiment of the present invention, the calculation efficiency speed has been greatly improved, and the high-dimensional molecular features are converted into two-dimensional pixel matrices with a fixed pattern. Regardless of the molecular structure differences or property differences, direct comparative analysis can be carried out based on the differences in the values (colors) of the same regions on the pictures.

[0061] Although the present invention has been described in detail above, the present invention is not limited thereto. Those skilled in the art of this technology can make various modifications according to the principle of the present invention. Therefore, all modifications made according to the principle of the present invention should be understood to fall within the protection scope of the present invention.

Claims

1. A method for predicting the properties of molecular drugs, characterized in that, Including: Obtain training molecules and basic information of the training molecules from a molecular database, and use the basic information of the training molecules to obtain the high-dimensional feature distance D of the training molecules, which includes: respectively and distributively calculating in parallel the N descriptor features and M fingerprint features of the training molecules according to the basic information of the training molecules; obtaining the high-dimensional descriptor feature distance D1 and high-dimensional fingerprint feature distance D2 of the training molecules by respectively and distributively calculating in parallel the descriptor feature distances between pairwise descriptor features among the N descriptor features of the training molecules and the fingerprint feature distances between pairwise fingerprint features among the M fingerprint features of the training molecules; where N and M are both positive integers. After respectively performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecules, obtain the two-dimensional feature pixel matrix GMap of the training molecules. Obtain the feature calculation value vector m of the prediction molecule, and assign the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain the representation two-dimensional matrix MolMap of the prediction molecule. Use the representation two-dimensional matrix MolMap of the prediction molecule as the molecular representation of the prediction molecule, and use the molecular representation of the prediction molecule to predict the drug properties of the prediction molecule.

2. The method according to claim 1, wherein The obtaining of the two-dimensional feature pixel matrix GMap of the training molecules after respectively performing dimensionality reduction and network processing on the high-dimensional feature distance D of the training molecules includes: Respectively perform dimensionality reduction on the high-dimensional descriptor feature distance D1 and high-dimensional fingerprint feature distance D2 of the training molecules to obtain the two-dimensional descriptor feature matrix MapD1 and two-dimensional fingerprint feature matrix Map2 of the training molecules. Respectively perform grid processing on the two-dimensional descriptor feature matrix MapD1 and two-dimensional fingerprint feature matrix Map2 of the training molecules to obtain the two-dimensional descriptor feature pixel matrix GMap1 and two-dimensional fingerprint feature pixel matrix GMap2 of the training molecules.

3. The method according to claim 2, wherein The obtaining of the feature calculation value vector m of the prediction molecule includes: Obtain the descriptor feature calculation value vector m1 according to the N descriptor features of the prediction molecule; and obtain the fingerprint feature calculation value vector m2 according to the M fingerprint features of the prediction molecule.

4. The method according to claim 3, wherein The obtaining of the feature calculation value vector m of the prediction molecule and assigning the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain the representation two-dimensional matrix MolMap of the prediction molecule includes: Assign the descriptor feature calculation value vector m1 of the prediction molecule to the two-dimensional descriptor feature pixel matrix GMap1 to obtain the representation two-dimensional descriptor matrix MolMap1 of the prediction molecule, and at the same time assign the fingerprint feature calculation value vector m2 of the prediction molecule to the two-dimensional fingerprint feature pixel matrix GMap2 to obtain the representation two-dimensional fingerprint matrix MolMap2 of the prediction molecule.

5. The method according to claim 4, wherein The using of the representation two-dimensional matrix MolMap of the prediction molecule as the molecular representation of the prediction molecule includes: Use the two-dimensional descriptor matrix MolMap1 representing the predicted molecule and the two-dimensional fingerprint matrix MolMap2 as the molecular representation of the predicted molecule.

6. A device for predicting the properties of molecular drugs, characterized in that, It includes: A first acquisition module, configured to acquire training molecules and basic information of the training molecules from a molecular database, and use the basic information of the training molecules to obtain the high-dimensional feature distance D of the training molecules; It includes: a calculation unit, configured to respectively and distributively and parallelly calculate N descriptor features and M fingerprint features of the training molecules according to the basic information of the training molecules; an acquisition unit, configured to obtain the high-dimensional descriptor feature distance D1 and the high-dimensional fingerprint feature distance D2 of the training molecules by respectively and distributively and parallelly calculating the descriptor feature distances between pairwise descriptor features among the N descriptor features of the training molecules and the fingerprint feature distances between pairwise fingerprint features among the M fingerprint features; where N and M are both positive integers; A processing module, configured to obtain the two-dimensional feature pixel matrix GMap of the training molecules after respectively performing dimensionality reduction and grid processing on the high-dimensional feature distance D of the training molecules; A second acquisition module, configured to acquire the feature calculation value vector m of the predicted molecule, and assign the feature calculation value vector m to the two-dimensional feature pixel matrix GMap to obtain the two-dimensional matrix MolMap representing the predicted molecule; A prediction module, configured to use the two-dimensional matrix MolMap representing the predicted molecule as the molecular representation of the predicted molecule, and use the molecular representation of the predicted molecule to predict the drug properties of the predicted molecule.

7. The device according to claim 6, characterized in that, The processing module includes: A dimensionality reduction processing unit, configured to respectively perform dimensionality reduction processing on the high-dimensional descriptor feature distance D1 and the high-dimensional fingerprint feature distance D2 of the training molecules to obtain the two-dimensional descriptor feature matrix MapD1 and the two-dimensional fingerprint feature matrix Map2 of the training molecules; A grid processing unit, configured to respectively perform grid processing on the two-dimensional descriptor feature matrix MapD1 and the two-dimensional fingerprint feature matrix Map2 of the training molecules to obtain the two-dimensional descriptor feature pixel matrix GMap1 and the two-dimensional fingerprint feature pixel matrix GMap2 of the training molecules.

8. The device according to claim 7, wherein The second acquisition module is specifically configured to obtain the descriptor feature calculation value vector m1 according to the N descriptor features of the predicted molecule; and obtain the fingerprint feature calculation value vector m2 according to the M fingerprint features of the predicted molecule.

Citation Information

Patent Citations

  • Implicated crime principle and network topological structural feature based recognition method for drug-target interaction

    CN105117618A

  • 3D display method and system for compounds

    CN107480429A