Molecular property analysis method, device, equipment, storage medium and program product

By clustering the molecular graph data of small molecules and establishing their structure-effect relationship, the problem of insufficient interpretability in traditional algorithms when interpreting the prediction results of molecular properties is solved, and a more intuitive and effective analysis and design of the properties of small molecules is achieved.

CN120220849APending Publication Date: 2025-06-27CONTEMPORARY AMPEREX TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311812935.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional GNN or machine learning algorithms have the problem of insufficient interpretability when predicting the properties of small molecules, making it difficult to explain the specific relationship between molecular structure and properties.

Method used

By using molecular graph data of multiple small molecules and pre-trained molecular property prediction models, high-dimensional vectors and molecular properties of each small molecule are obtained, and cluster analysis is performed to establish the structure-activity relationship of small molecules and explain the prediction results of molecular properties.

Benefits of technology

The establishment of the structure-activity relationship of small molecules is achieved, the prediction results of molecular properties are explained, and researchers are assisted in the screening and design of small molecules, improving the analysis efficiency and design efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220849A_ABST
    Figure CN120220849A_ABST
Patent Text Reader

Abstract

The invention relates to a molecular property analysis method and device, equipment, a storage medium and a program product. The method comprises the following steps: obtaining a high-dimensional vector and a molecular property corresponding to each small molecule according to molecular map data of a plurality of small molecules and a pre-trained molecular property prediction model; performing clustering analysis on the high-dimensional vectors corresponding to the plurality of small molecules to obtain a clustering result; analyzing molecular structures and molecular properties of the plurality of small molecules according to the clustering result to obtain an analysis result; wherein the analysis result is used for representing a corresponding relation between the molecular structure and the molecular property. By adopting the method, the structure-function relationship of small molecules can be established, and the prediction result of molecular properties can be explained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of small molecule technology, and particularly to a method, device, equipment, storage medium, and program product for molecular property analysis. Background Art

[0002] In the field of chemistry, predicting the molecular properties of small molecules is an important research content, which usually requires the use of technologies such as molecular descriptors, machine learning algorithms, or graph neural networks (GNNs). In the field of GNNs, many studies have been conducted to improve the prediction accuracy of molecular properties by adjusting GNNs or machine learning algorithms.

[0003] However, traditional GNNs or machine learning algorithms have certain limitations in the interpretability of prediction results. Summary of the Invention

[0004] Based on the above problems, this application provides a method, device, equipment, storage medium, and program product for molecular property analysis, which can establish the structure-activity relationship of small molecules and explain the prediction results of molecular properties.

[0005] In a first aspect, this application provides a method for molecular property analysis, which includes: obtaining the high-dimensional vectors and molecular properties corresponding to each small molecule according to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model; performing clustering analysis on the high-dimensional vectors corresponding to the multiple small molecules to obtain a clustering result; analyzing the molecular structures and molecular properties of the multiple small molecules according to the clustering result to obtain an analysis result; wherein the analysis result is used to characterize the correspondence between the molecular structure and the molecular property.

[0006] The technical solution of the embodiments of this application establishes the structure-activity relationship of small molecules, which can not only explain the prediction results of molecular properties, but also assist researchers in screening and designing small molecules based on this structure-activity relationship.

[0007] In some embodiments, analyzing the molecular structures and molecular properties of the multiple small molecules according to the clustering result to obtain an analysis result includes: determining the molecular properties corresponding to the categories to which each small molecule belongs according to the clustering result; for a target category among the multiple categories after clustering, obtaining a target functional group according to the functional group similarity among the multiple small molecules; and obtaining an analysis result according to the molecular property corresponding to the target category and the target functional group. The technical solution of the embodiments of this application establishes the correspondence between the functional groups and molecular properties of small molecules, which can not only explain the prediction results of molecular properties, but also assist researchers in screening and designing small molecules with specific functional groups, thereby improving the screening efficiency and design efficiency.

[0008] In some embodiments, determining the molecular properties corresponding to the categories to which each small molecule belongs according to the clustering result includes: performing dimensionality reduction processing on the clustering result to obtain a dimensionality reduction result; and determining the molecular properties corresponding to the categories to which each small molecule belongs according to the dimensionality reduction result. By performing dimensionality reduction processing on the technical solution of the embodiments of the present application, the corresponding relationship between the categories to which small molecules belong and the molecular properties can be more intuitively displayed, and it is more convenient to analyze the structure-activity relationship of small molecules, thereby improving the analysis efficiency of molecular properties.

[0009] In some embodiments, obtaining the high-dimensional vectors and molecular properties corresponding to each small molecule according to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model includes: inputting the molecular graph data of each small molecule into the molecular property prediction model to obtain the high-dimensional vectors corresponding to each small molecule output by the target embedding layer in the molecular property prediction model, and the molecular properties output by the molecular property prediction model. In the technical solution of the embodiments of the present application, a small molecule is represented as a high-dimensional vector, so that there is a certain correlation between the high-dimensional vector and the predicted molecular property, so that the structure-activity relationship of small molecules can be established after clustering analysis, and then assist researchers in screening and designing small molecules.

[0010] In some embodiments, the method further includes: obtaining a training sample set; the training sample set includes a plurality of sample graph data and the molecular property annotations corresponding to each sample graph data; and performing model training based on the training sample set and an initial prediction model to obtain a molecular property prediction model. The technical solution of the embodiments of the present application provides a model training method, so as to obtain the high-dimensional vectors and molecular properties of small molecules by using the trained molecular property prediction model, thereby establishing the structure-activity relationship of small molecules and providing a basis for explaining the prediction results.

[0011] In some embodiments, performing model training based on the training sample set and an initial prediction model to obtain a molecular property prediction model includes: inputting the sample graph data of the training sample set into the initial prediction model to obtain the predicted molecular properties output by the initial prediction model; calculating the loss value between the predicted molecular properties and the molecular property annotations corresponding to the sample graph data according to a preset loss function; in the case that the loss value does not meet the convergence condition, adjusting the model parameters; performing training based on the training sample set and the model with adjusted parameters until the loss value output by the model meets the convergence condition, and then ending the training, and determining the model at the end of the training as the molecular property prediction model. The technical solution of the embodiments of the present application provides a model training method. Through the trained molecular property prediction model, the high-dimensional vectors and molecular properties of small molecules can be obtained, so that the structure-activity relationship of small molecules can be established by performing clustering analysis on the high-dimensional vectors of small molecules, and a basis for explaining the prediction results is provided.

[0012] In some embodiments, obtaining molecular graph data of multiple small molecules includes: obtaining molecular configuration data of each small molecule; respectively performing format conversion on the molecular configuration data of each small molecule to obtain the SMILES code of each small molecule; respectively inputting the SMILES code of each small molecule into a pre-trained graph neural network to obtain the molecular graph data of each small molecule output by the graph neural network. The technical solution of the embodiments of the present application provides a way to obtain molecular graph data, which can facilitate subsequent prediction of molecular properties using a molecular property prediction model and improve the prediction accuracy.

[0013] In a second aspect, the present application also provides a molecular property analysis device, which includes:

[0014] A property prediction module, configured to obtain a high-dimensional vector and molecular properties corresponding to each small molecule according to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model;

[0015] A clustering analysis module, configured to perform clustering analysis on the high-dimensional vectors corresponding to multiple small molecules to obtain a clustering result;

[0016] A structure-activity analysis module, configured to analyze the molecular structures and molecular properties of multiple small molecules according to the clustering result to obtain an analysis result; wherein, the analysis result is used to characterize the corresponding relationship between the molecular structure and the molecular properties.

[0017] In a third aspect, the present application also provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in any item of the first aspect is implemented.

[0018] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any item of the first aspect is implemented.

[0019] In a fifth aspect, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any item of the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] By reading the detailed description of the following optional embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the optional embodiments and are not considered to be a limitation of the present application. And in all the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0021] Figure 1 is a schematic flowchart of a molecular property analysis method according to an embodiment of the present application;

[0022] Figure 2 It is a schematic flowchart of the molecular structure and molecular property analysis steps in an embodiment of the present application;

[0023] Figure 3 It is a schematic diagram of the clustering result and functional groups in an embodiment of the present application;

[0024] Figure 4 It is a schematic flowchart of the step for determining molecular properties in an embodiment of the present application;

[0025] Figure 5a It is one of the schematic diagrams of the dimensionality reduction result in an embodiment of the present application;

[0026] Figure 5b It is another schematic diagram of the dimensionality reduction result in an embodiment of the present application;

[0027] Figure 6 It is a schematic structural diagram of the molecular property prediction model in an embodiment of the present application;

[0028] Figure 7 It is a schematic flowchart of the model training step in an embodiment of the present application;

[0029] Figure 8 It is a schematic flowchart of the model training step in an embodiment of the present application;

[0030] Figure 9 It is a schematic flowchart of the step for obtaining molecular graph data in an embodiment of the present application;

[0031] Figure 10 It is a structural block diagram of the molecular property analysis device in an embodiment of the present application;

[0032] Figure 11 It is a structural block diagram of the molecular property analysis device in an embodiment of the present application;

[0033] Figure 12 It is a structural block diagram of the molecular property analysis device in an embodiment of the present application;

[0034] Figure 13 It is an internal structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners

[0035] Next, embodiments of the technical solution of the present application will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present application, so they are only examples and cannot be used to limit the protection scope of the present application.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion.

[0037] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "a plurality" is more than two, unless otherwise specifically defined.

[0038] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase does not necessarily refer to the same embodiment everywhere in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0039] In the description of the embodiments of this application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0040] Currently, in the field of chemistry, predicting the molecular properties of small molecules is an important research content, which usually requires the use of technologies such as molecular descriptors, machine learning algorithms, or graph neural networks (GNNs). In the field of GNNs, there have been many studies to improve the prediction accuracy of molecular properties by adjusting GNNs or machine learning algorithms. However, traditional GNNs or machine learning algorithms usually only predict the molecular properties, and it is mostly difficult to explain the structure-activity relationship of small molecules, that is, the relationship between the molecular structure and the molecular properties. That is to say, traditional GNNs or machine learning algorithms have certain limitations in the interpretability of prediction results.

[0041] In view of the above problems, the embodiments of the present application provide a molecular property analysis solution. According to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model, the high-dimensional vectors and molecular properties corresponding to each small molecule are obtained; clustering analysis is performed on the high-dimensional vectors corresponding to multiple small molecules to obtain a clustering result; according to the clustering result, the molecular structures and molecular properties of multiple small molecules are analyzed to obtain the corresponding relationship between the molecular structure and the molecular property. The technical solution of the embodiments of the present application establishes the structure-activity relationship of small molecules, which can not only explain the prediction results of molecular properties, but also assist researchers in screening and designing small molecules based on this structure-activity relationship.

[0042] According to some embodiments of the present application, with reference to Figure 1 , a molecular property analysis method is provided. The embodiments of the present application take the application of this method to a terminal as an example for illustration. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The method may include the following steps:

[0043] Step 101, according to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model, obtain the high-dimensional vectors and molecular properties corresponding to each small molecule.

[0044] Among them, the small molecule can be any one of electrolyte small molecules and drug small molecules; the molecular graph data can include at least one of the molecular structure feature graph and the molecular structure feature vector of the small molecule. The molecular property is the property exhibited by the small molecule, which may include at least one of HOMO (Highest Occupied Molecular Orbital), LUMO (Lowest Unoccupied Molecular Orbital), viscosity, ionic conductivity, melting point, and boiling point.

[0045] The terminal can obtain the molecular structures of multiple small molecules, and according to the bond connection relationship in the molecular structure, convert the molecular structure into molecular graph data. For example, an undirected structure graph is drawn according to the bond connection relationship in the molecular structure to obtain the molecular graph data. The terminal can also use a feature extraction model to extract features from the molecular structure of the small molecule to obtain the molecular graph data of the small molecule.

[0046] It should be noted that the acquisition methods of the molecular graph data include but are not limited to the above two methods, and can be determined according to the actual situation.

[0047] The terminal pre-trains a molecular property prediction model. After obtaining the molecular graph data of multiple small molecules, the molecular graph data is input into the molecular property prediction model, and the molecular property prediction model outputs the high-dimensional vector of the small molecule and the molecular property of the small molecule.

[0048] For example, the molecular property prediction model is the LUMO prediction model. The molecular graph data of small molecule 1 is input into the LUMO prediction model, and the LUMO prediction model outputs the high-dimensional vector and LUMO value of small molecule 1. And so on, the molecular graph data of small molecules 2, 3... n are respectively input into the LUMO prediction model, and the LUMO prediction model respectively outputs the high-dimensional vectors and LUMO values of small molecules 2, 3... n.

[0049] Step 102: Perform clustering analysis on the high-dimensional vectors corresponding to multiple small molecules to obtain a clustering result.

[0050] Among them, the clustering result includes the categories to which each small molecule belongs in the high-dimensional space, and one dimension of the high-dimensional space is the molecular property.

[0051] The terminal can determine the positions of the small molecules in the high-dimensional space according to the high-dimensional vectors corresponding to the small molecules, and then perform clustering analysis on the multiple small molecules according to the positions of the small molecules in the high-dimensional space to obtain a clustering result.

[0052] It should be noted that the number of categories in the clustering result can be determined according to the actual situation.

[0053] Step 103: Analyze the molecular structures and molecular properties of multiple small molecules according to the clustering result to obtain an analysis result.

[0054] Among them, the analysis result is used to characterize the correspondence between the molecular structure and the molecular property, that is, to characterize the structure-activity relationship of small molecules.

[0055] According to the categories included in the clustering result, small molecules belonging to each category can be selected from multiple small molecules; then, by analyzing the molecular structures of small molecules in each category, the structural commonalities of small molecules in each category can be summarized. For each category, according to the structural commonalities and the molecular properties of small molecules, the correspondence between the molecular structure and the molecular property can be obtained, that is, the analysis result can be obtained.

[0056] For example, the clustering result includes multiple categories, and one of the categories is that the LUMO value is within a preset range; according to the clustering structure, small molecules with LUMO values within the preset range can be selected from multiple small molecules; by analyzing the molecular structures of such small molecules, it is determined that such small molecules all include structure M. Since the LUMO values of such small molecules are within the preset range, it can be determined that there is a certain correspondence between structure M and the LUMO value within the preset range. It can be seen that the embodiments of the present application establish the structure-activity relationship of small molecules.

[0057] In one embodiment, after obtaining the correspondence between the molecular structure and the molecular properties, small molecules can be screened according to this correspondence. For example, if there is a correspondence between the LUMO value within a preset range and the structure M, small molecules with the structure M can be screened out from multiple small molecules to be screened, and the LUMO values of these small molecules are likely to be within the preset range. It can be understood that based on the above correspondence for small molecule screening, it is not necessary to predict the molecular properties of all small molecules to be screened. Only the molecular properties of small molecules with a specific structure need to be predicted to screen out small molecules whose molecular properties meet the requirements. Therefore, the screening efficiency can be significantly improved.

[0058] In one implementation, after obtaining the correspondence between the molecular structure and the molecular properties, small molecules can also be designed according to this correspondence. For example, if there is a correspondence between the LUMO value within a preset range and the structure M, to obtain small molecules with the LUMO value within the preset range, the structure M can be added to the basic molecular structure during design, and small molecules that meet the requirements can be obtained. It can be understood that based on the above correspondence for small molecule design, the success rate of small molecule design can be improved, thereby improving the design efficiency of small molecules.

[0059] In the above embodiments, according to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model, the high-dimensional vectors and molecular properties corresponding to each small molecule are obtained; cluster analysis is performed on the high-dimensional vectors corresponding to multiple small molecules to obtain a clustering result; according to the clustering result, the molecular structures and molecular properties of multiple small molecules are analyzed to obtain the correspondence between the molecular structure and the molecular properties. The technical solution of the embodiments of the present application establishes the structure-activity relationship of small molecules, which can not only explain the prediction results of molecular properties, but also assist researchers in screening and designing small molecules based on this structure-activity relationship.

[0060] According to some embodiments of the present application, referring to Figure 2 , an implementation manner related to the step "According to the clustering result, analyze the molecular structures and molecular properties of multiple small molecules to obtain an analysis result" in the above embodiments, this manner may include the following steps:

[0061] Step 201, determine the molecular properties corresponding to the categories to which each small molecule belongs according to the clustering result.

[0062] Since the clustering result includes the categories to which each small molecule belongs in the high-dimensional space, and one dimension of the high-dimensional space is the molecular property. Therefore, the correspondence between the categories to which small molecules belong and the molecular properties can be determined.

[0063] Step 202, for the target category among multiple categories after clustering, obtain the target functional group according to the functional group similarity among multiple small molecules.

[0064] Take any one of the multiple categories after clustering as the target category, perform a structural analysis on multiple small molecules in the target category, find the same or similar functional groups, and use the found functional groups as the target functional groups.

[0065] Refer to Figure 3 , take the category with a LUMO value lower than 2.5 eV as the target category, perform a structural analysis on multiple small molecules in the target category, find that the same functional group among these small molecules is O-F, and then use the functional group O-F as the target functional group.

[0066] Step 203, obtain the analysis result according to the molecular properties corresponding to the target category and the target functional group.

[0067] Establish the correspondence between the target category and the target functional group to obtain the analysis result. For example, establish the correspondence between the LUMO value lower than 2.5 eV and the functional group O-F to obtain the analysis result.

[0068] In one of the embodiments, small molecules can be screened based on functional groups. For example, in the sample space of fluorinated ether small molecules, small molecules with functional groups such as ROCF3, ROCH2CF3, ROCF2CH2F, ROCF2CHF2, etc. have a higher LUMO and are more resistant to reduction during electrode charging; while small molecules with functional groups such as ROF and RC(==CH)CH3 have a lower LUMO and are less resistant to reduction. According to the correspondence between the above functional groups and molecular properties, small molecules that better meet the actual requirements can be screened out.

[0069] In the above embodiments, the molecular properties corresponding to the categories to which each small molecule belongs are determined according to the clustering result; for the target category among the multiple categories after clustering, the target functional group is obtained according to the functional group similarity among multiple small molecules; and the analysis result is obtained according to the molecular properties corresponding to the target category and the target functional group. The technical solution of the embodiments of the present application establishes the correspondence between the functional groups and molecular properties of small molecules, which can not only explain the prediction results of molecular properties, but also assist researchers in screening and designing small molecules with specific functional groups, thereby improving the screening efficiency and design efficiency.

[0070] According to some embodiments of the present application, refer to Figure 4 , which relates to an implementation manner of the step "determine the molecular properties corresponding to the categories to which each small molecule belongs according to the clustering result" in the above embodiments, and this manner may include the following steps:

[0071] Step 301, perform dimensionality reduction processing on the clustering result to obtain the dimensionality reduction result.

[0072] Among them, the dimensionality reduction result includes the categories to which each small molecule belongs in the low-dimensional space, and one of the dimensions of the low-dimensional space is the molecular property. In one implementation, the low-dimensional space is a two-dimensional space.

[0073] A preset dimensionality reduction algorithm can be used to perform dimensionality reduction on the clustering result to obtain the dimensionality reduction result. Among them, the preset clustering algorithm can include but is not limited to PCA (Principal Component Analysis) and t-SNE. Figure 5a Shows the dimensionality reduction result obtained by performing linear dimensionality reduction using PCA. Figure 5b Shows the dimensionality reduction result obtained by performing non-linear dimensionality reduction using t-SNE.

[0074] Step 302, determine the molecular property corresponding to the category to which each small molecule belongs according to the dimensionality reduction result.

[0075] Since the dimensionality reduction result includes the categories to which each small molecule belongs in the low-dimensional space, and one of the dimensions of the low-dimensional space is the molecular property. Therefore, the correspondence between the category to which the small molecule belongs and the molecular property can be determined according to the dimensionality reduction result.

[0076] In the above implementation, after obtaining the high-dimensional vector and molecular property of the small molecule using the molecular property prediction model, first perform clustering analysis on the high-dimensional vector, and then perform dimensionality reduction. In another implementation, after obtaining the high-dimensional vector and molecular property of the small molecule using the molecular property prediction model, it is also possible to first perform dimensionality reduction on the high-dimensional vector, and then perform clustering analysis; or directly perform clustering analysis during the dimensionality reduction process.

[0077] In the above embodiment, perform dimensionality reduction on the clustering result to obtain the dimensionality reduction result; determine the molecular property corresponding to the category to which each small molecule belongs according to the dimensionality reduction result. The technical solution of the embodiment of the present application performs dimensionality reduction, which can more intuitively display the correspondence between the category to which the small molecule belongs and the molecular property, and is more convenient for analyzing the structure-activity relationship of the small molecule, thereby improving the analysis efficiency of the molecular property.

[0078] According to some embodiments of the present application, the step in the above embodiment "obtain the high-dimensional vector and molecular property corresponding to each small molecule according to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model" may include: input the molecular graph data of each small molecule into the molecular property prediction model to obtain the high-dimensional vector corresponding to each small molecule output by the target embedding layer in the molecular property prediction model and the molecular property output by the molecular property prediction model.

[0079] Refer to Figure 6, The pre-trained molecular property prediction model may include an input layer, at least one embedding layer, and an output layer. For example, the molecular property prediction model uses a neural network, and the parameters of the embedding layer are 6*50*50, that is, it has 6 embedding layers. It should be noted that the structure and parameters of the molecular property prediction model can be set according to the actual situation.

[0080] Input the molecular graph data of each small molecule into the molecular property prediction model, and the output layer of the molecular property prediction model can output the molecular properties of each small molecule. Moreover, take any embedding layer in the molecular property prediction model as the target embedding layer, and take the vector output by the target embedding layer as the high-dimensional vector corresponding to the small molecule.

[0081] In practical applications, since the last embedding layer among multiple embedding layers is closest to the output layer, the correlation between the vector output by the last embedding layer and the molecular properties output by the output layer is the highest. Therefore, the last embedding layer can be taken as the target embedding layer, and the vector output by the target embedding layer can be taken as the high-dimensional vector corresponding to the small molecule.

[0082] In the above embodiments, input the molecular graph data of each small molecule into the molecular property prediction model to obtain the high-dimensional vectors corresponding to each small molecule output by the target embedding layer in the molecular property prediction model and the molecular properties output by the molecular property prediction model. In the technical solution of the embodiments of the present application, represent the small molecule as a high-dimensional vector, so that there is a certain correlation between the high-dimensional vector and the predicted molecular properties, so that the structure-activity relationship of small molecules can be established after cluster analysis, and then assist researchers in screening and designing small molecules.

[0083] According to some embodiments of the present application, referring to Figure 7 , The embodiments of the present application may further include the step of model training:

[0084] Step 401, obtain a training sample set.

[0085] Among them, the training sample set includes multiple sample graph data and the molecular property annotations corresponding to each sample graph data.

[0086] The terminal obtains the molecular structures of multiple sample small molecules, and can obtain the sample graph data of each sample small molecule according to the bond connection relationship in the molecular structure; or can use a graph neural network to obtain the sample graph data of each sample small molecule. The terminal can obtain the molecular property annotations of each sample graph data input by the user, so as to summarize multiple sample graph data and the molecular property annotations corresponding to each sample graph data into a training sample set.

[0087] Alternatively, the terminal can obtain the pre-stored training sample set from a cloud server or a blockchain node.

[0088] It should be noted that the acquisition methods of the training sample set include, but are not limited to, the above two implementation methods, and can be determined according to the actual situation.

[0089] Step 402: Perform model training based on the training sample set and the initial prediction model to obtain a molecular property prediction model.

[0090] Select a model from multiple machine learning models as the initial prediction model. Then, input the sample graph data in the training sample set into the initial prediction model for model training. During the training process, hyperparameters such as kernel functions and likelihood estimates can be adjusted to make the model converge, thereby obtaining a molecular property prediction model.

[0091] In one implementation, model training can be performed based on the predicted molecular properties output by the initial prediction model and the molecular property annotations to obtain an intermediate prediction model. Refer to the acquisition method of the training sample set to obtain a test sample set, and use the test sample set to test the intermediate prediction model to obtain the test results output by the intermediate prediction model. Evaluate the test results. If the evaluation results indicate that the test results output by the intermediate prediction model meet the requirements, then determine the intermediate prediction model as the molecular property prediction model. If the evaluation results indicate that the test results output by the intermediate prediction model do not meet the requirements, then use the training sample set to train the intermediate prediction model until the test results output by the trained model meet the requirements, and determine the model at the end of the training as the molecular prediction model.

[0092] The above evaluation of the test results can include evaluating the accuracy of the test results and the prediction speed of the intermediate prediction model, etc. It should be noted that the evaluation methods include, but are not limited to, the above description, and appropriate evaluation methods can be selected according to the actual situation.

[0093] In one implementation, after obtaining the intermediate prediction model, pruning, compression, etc. can also be performed on the intermediate prediction model to obtain a molecular property prediction model.

[0094] In the above embodiments, a training sample set is obtained; model training is performed based on the training sample set and the initial prediction model to obtain a molecular property prediction model. The technical solution of the embodiments of the present application provides a model training method to obtain high-dimensional vectors and molecular properties of small molecules by using the trained molecular property prediction model, thereby establishing the structure-activity relationship of small molecules and providing a basis for explaining the prediction results.

[0095] According to some embodiments of the present application, referring to Figure 8 , an implementation of the step "Perform model training based on the training sample set and the initial prediction model to obtain a molecular property prediction model" in the above embodiments, this method may include the following steps:

[0096] Step 501: Input the sample graph data of the training sample set into the initial prediction model to obtain the predicted molecular properties output by the initial prediction model.

[0097] Input any sample graph data of the training sample set into the initial prediction model to obtain the predicted molecular properties corresponding to this sample graph data output by the initial prediction model.

[0098] Step 502: Calculate the loss value between the predicted molecular properties and the molecular property annotations corresponding to the sample graph data according to the preset loss function.

[0099] Input the predicted molecular properties and the molecular property annotations corresponding to the sample graph data into the preset loss function for calculation to obtain the loss value. The preset loss function can adopt mean square error loss function, cross - entropy loss function, etc., and the loss function can be selected according to the actual situation.

[0100] Step 503: When the loss value does not meet the convergence condition, perform model parameter adjustment.

[0101] If the loss value does not meet the convergence condition, the parameters in the model can be adjusted, or the structure of the model can be adjusted.

[0102] Step 504: Based on the training sample set and the model after parameter adjustment, perform training until the loss value output by the model meets the convergence condition, then end the training, and determine the model at the end of the training as the molecular property prediction model.

[0103] Input another sample graph data in the training sample set into the model after parameter adjustment to obtain the new predicted molecular properties output by the model. Calculate the loss between the new predicted molecular properties and the molecular property annotations corresponding to this sample graph data to obtain a new loss value. Determine whether the new loss value meets the convergence condition. If it does not meet the convergence condition, continue to perform model parameter adjustment and model training. If the new loss value meets the convergence condition, end the training, and use the model at the end of the training as the molecular prediction model.

[0104] In the above embodiments, the sample graph data of the training sample set is input into the initial prediction model to obtain the predicted molecular properties output by the initial prediction model; the loss value between the predicted molecular properties and the molecular property annotations corresponding to the sample graph data is calculated according to the preset loss function; in the case where the loss value does not meet the convergence condition, the model parameters are adjusted; if the loss value does not meet the convergence condition, the parameters in the model can be adjusted, or the structure of the model can be adjusted; training is performed based on the training sample set and the model after parameter adjustment until the loss value output by the model meets the convergence condition, and the model at the end of training is determined as the molecular property prediction model. The technical solution of the embodiments of the present application provides a model training method. Through the trained molecular property prediction model, the high-dimensional vectors and molecular properties of small molecules can be obtained, so as to establish the structure-activity relationship of small molecules by performing clustering analysis on the high-dimensional vectors of small molecules, providing a basis for explaining the prediction results.

[0105] According to some embodiments of the present application, referring to Figure 9 , the following steps may further be included:

[0106] Step 601, obtain the molecular configuration data of each small molecule.

[0107] The terminal can obtain the molecular configuration data of each small molecule from the cloud server, or import the molecular configuration data of each small molecule from the preset storage space in the terminal.

[0108] Step 602, perform format conversion on the molecular configuration data of each small molecule respectively to obtain the SMILES codes of each small molecule.

[0109] Among them, SMILES (Simplified molecular input line entry system) is a specification that clearly describes the molecular structure with ASCII strings.

[0110] The molecular configuration data of each small molecule is a file in a preset format, and the file format conversion of the molecular configuration data of each small molecule can be performed through software or toolkits such as openbabel and RDKit to obtain the SMILES codes of each small molecule.

[0111] Step 603, input the SMILES codes of each small molecule into the pre-trained graph neural network respectively to obtain the molecular graph data of each small molecule output by the graph neural network.

[0112] The graph neural network is pre-trained. After determining the SMILES codes of each small molecule, the SMILES codes of each small molecule are input into the graph neural network, and the graph neural network outputs the molecular graph data of each small molecule.

[0113] The training method of the graph neural network can refer to the training method of the above-mentioned molecular property prediction model, which will not be elaborated in the embodiments of the present application.

[0114] In the above embodiments, the molecular configuration data of each small molecule is obtained; the molecular configuration data of each small molecule is respectively subjected to format conversion to obtain the SMILES code of each small molecule; the SMILES code of each small molecule is respectively input into the pre-trained graph neural network to obtain the molecular graph data of each small molecule output by the graph neural network. The technical solution of the embodiments of the present application provides a method for obtaining molecular graph data, which can facilitate subsequent molecular property prediction using the molecular property prediction model and improve the prediction accuracy.

[0115] According to some embodiments of the present application, a molecular property analysis method is provided, which may include the following steps:

[0116] Step 1, obtain a training sample set.

[0117] Among them, the training sample set includes a plurality of sample graph data and the molecular property annotations corresponding to each sample graph data.

[0118] Step 2, perform model training based on the training sample set and the initial prediction model to obtain a molecular property prediction model.

[0119] Step 3, obtain the molecular configuration data of each small molecule.

[0120] Step 4, respectively perform format conversion on the molecular configuration data of each small molecule to obtain the SMILES code of each small molecule.

[0121] Step 5, respectively input the SMILES code of each small molecule into the pre-trained graph neural network to obtain the molecular graph data of each small molecule output by the graph neural network.

[0122] Step 6, input the molecular graph data of each small molecule into the molecular property prediction model to obtain the high-dimensional vectors corresponding to each small molecule output by the target embedding layer in the molecular property prediction model and the molecular properties output by the molecular property prediction model.

[0123] Step 7, perform clustering processing on the high-dimensional vectors corresponding to multiple small molecules to obtain a clustering result.

[0124] Step 8, perform dimensionality reduction processing on the clustering result to obtain a dimensionality reduction result.

[0125] Step 9, determine the molecular properties corresponding to the categories to which each small molecule belongs according to the dimensionality reduction result.

[0126] Step 10, for the target category among multiple categories after clustering, obtain the target functional group according to the functional group similarity between multiple small molecules.

[0127] Step 11: Obtain an analysis result according to the molecular properties and target functional groups corresponding to the target category.

[0128] In the above embodiments, a molecular property prediction model is pre-trained. Subsequently, the molecular property prediction model can be used to obtain high-dimensional vectors representing small molecules and the molecular properties of small molecules. Thus, a structure-activity relationship of small molecules can be established based on the clustering results of the high-dimensional vectors of small molecules, providing a basis for explaining the prediction results and assisting researchers in screening and designing small molecules, thereby improving the screening efficiency and design efficiency.

[0129] It should be understood that although the steps in the above flowchart are sequentially shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0130] Based on the same inventive concept, an embodiment of the present application also provides a molecular property analysis device for implementing the above-mentioned molecular property analysis method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the molecular property analysis device provided below can refer to the limitations on the molecular property analysis method in the above text and will not be repeated here.

[0131] According to some embodiments of the present application, with reference to Figure 10 , a molecular property analysis device is provided, and the device includes:

[0132] A property prediction module 701, configured to obtain high-dimensional vectors and molecular properties corresponding to each small molecule according to the molecular graph data of multiple small molecules and a pre-trained molecular property prediction model;

[0133] A clustering analysis module 702, configured to perform clustering analysis on the high-dimensional vectors corresponding to multiple small molecules to obtain a clustering result;

[0134] A structure-activity analysis module 703, configured to analyze the molecular structures and molecular properties of multiple small molecules according to the clustering result to obtain an analysis result, where the analysis result is used to characterize the corresponding relationship between the molecular structure and the molecular property.

[0135] In some embodiments, the structure-activity analysis module 703 is specifically configured to determine the molecular properties corresponding to the categories to which each of the small molecules belongs according to the clustering result; for the target category, obtain the target functional group based on the functional group similarity among the multiple small molecules; and obtain the analysis result according to the molecular properties corresponding to the target category and the target functional group.

[0136] In some embodiments, the structure-activity analysis module 703 is specifically configured to perform dimensionality reduction processing on the clustering result to obtain a dimensionality reduction result; and determine the molecular properties corresponding to the categories to which each of the small molecules belongs according to the dimensionality reduction result.

[0137] In some embodiments, the property prediction module 701 is specifically configured to input the molecular graph data of each of the small molecules into the molecular property prediction model to obtain the high-dimensional vectors corresponding to each of the small molecules output by the target embedding layer in the molecular property prediction model, and the molecular properties output by the molecular property prediction model.

[0138] In some embodiments, referring to Figure 11 , the device further includes:

[0139] A sample acquisition module 704, configured to acquire a training sample set; the training sample set includes a plurality of sample graph data and the molecular property annotations corresponding to each of the sample graph data;

[0140] A model training module 705, configured to perform model training based on the training sample set and an initial prediction model to obtain the molecular property prediction model.

[0141] In some embodiments, the model training module 705 is specifically configured to input the sample graph data of the training sample set into the initial prediction model to obtain the predicted molecular properties output by the initial prediction model; calculate the loss value between the predicted molecular properties and the molecular property annotations corresponding to the sample graph data according to a preset loss function; adjust the model parameters in the case that the loss value does not meet the convergence condition; perform training based on the training sample set and the model with adjusted parameters until the loss value output by the model meets the convergence condition, and determine the model at the end of training as the molecular property prediction model.

[0142] In some embodiments, referring to Figure 12 , the device further includes:

[0143] A graph data acquisition module 706, configured to acquire the molecular configuration data of each of the small molecules; perform format conversion on the molecular configuration data of each of the small molecules respectively to obtain the SMILES codes of each of the small molecules; and input the SMILES codes of each of the small molecules into a pre-trained graph neural network respectively to obtain the molecular graph data of each of the small molecules output by the graph neural network.

[0144] Each module in the above-mentioned over-molecule property analysis device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0145] According to some embodiments of the present application, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 13 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a method for analyzing molecular properties. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0146] Those skilled in the art can understand that Figure 13 the structure shown in

[0147] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout. According to some embodiments of the present application, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by the processor of the computer device to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0148] According to some embodiments of the present application, a computer program product is further provided. When the computer program is executed by a processor, the above method can be implemented. The computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, part or all of the above method can be implemented in accordance with the process or function described in the embodiments of the present application.

[0149] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0150] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0151] The embodiments described above merely represent several implementation manners of the present application, facilitating the specific and detailed understanding of the technical solution of the present application, but should not be construed as limiting the protection scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. It should be understood that the technical solutions obtained by those skilled in the art through logical analysis, reasoning or limited experiments based on the technical solutions provided by the present application are all within the protection scope of the appended claims of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the content of the appended claims, and the description and drawings can be used to explain the content of the claims.

Claims

1. A method for molecular property analysis, characterized in that, The method includes: Obtaining a high-dimensional vector and a molecular property corresponding to each of the small molecules according to the molecular graph data of a plurality of small molecules and a pre-trained molecular property prediction model; Performing clustering analysis on the high-dimensional vectors corresponding to the plurality of small molecules to obtain a clustering result; Analyzing the molecular structures and molecular properties of the plurality of small molecules according to the clustering result to obtain an analysis result, where the analysis result is used to characterize the corresponding relationship between the molecular structure and the molecular property.

2. The method according to claim 1, wherein The analyzing the molecular structures and molecular properties of the plurality of small molecules according to the clustering result to obtain an analysis result includes: Determining the molecular property corresponding to the category to which each of the small molecules belongs according to the clustering result; For a target category among the multiple categories after clustering, obtaining a target functional group according to the functional group similarity among the plurality of small molecules; Obtaining the analysis result according to the molecular property corresponding to the target category and the target functional group.

3. The method according to claim 2, characterized in that, The determining the molecular property corresponding to the category to which each of the small molecules belongs according to the clustering result includes: Performing dimensionality reduction processing on the clustering result to obtain a dimensionality reduction result; Determining the molecular property corresponding to the category to which each of the small molecules belongs according to the dimensionality reduction result.

4. The method according to claim 1, wherein The obtaining a high-dimensional vector and a molecular property corresponding to each of the small molecules according to the molecular graph data of the plurality of small molecules and a pre-trained molecular property prediction model includes: Inputting the molecular graph data of each of the small molecules into the molecular property prediction model to obtain the high-dimensional vector corresponding to each of the small molecules output by the target embedding layer in the molecular property prediction model, and the molecular property output by the molecular property prediction model.

5. The method according to claim 4, wherein The method further includes: Obtaining a training sample set, where the training sample set includes a plurality of sample graph data and the molecular property annotation corresponding to each of the sample graph data; Performing model training based on the training sample set and an initial prediction model to obtain the molecular property prediction model.

6. The method according to claim 5, wherein The performing model training based on the training sample set and an initial prediction model to obtain the molecular property prediction model includes: Inputting the sample graph data of the training sample set into the initial prediction model to obtain the predicted molecular property output by the initial prediction model; Calculating a loss value between the predicted molecular property and the molecular property annotation corresponding to the sample graph data according to a preset loss function; Performing model parameter adjustment when the loss value does not meet the convergence condition; Performing training based on the training sample set and the model with adjusted parameters until the loss value output by the model meets the convergence condition, and then ending the training and determining the model at the end of the training as the molecular property prediction model.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtaining the molecular configuration data of each of the small molecules; Performing format conversion on the molecular configuration data of each of the small molecules respectively to obtain the SMILES code of each of the small molecules; Inputting the SMILES code of each of the small molecules into a pre-trained graph neural network respectively to obtain the molecular graph data of each of the small molecules output by the graph neural network.

8. A molecular property analysis device, characterized in that, The device includes: A property prediction module, configured to obtain a high-dimensional vector and a molecular property corresponding to each of the small molecules according to the molecular graph data of a plurality of small molecules and a pre-trained molecular property prediction model; A clustering analysis module, configured to perform clustering analysis on the high-dimensional vectors corresponding to the plurality of small molecules to obtain a clustering result; A structure-activity analysis module, configured to analyze the molecular structures and molecular properties of the plurality of small molecules according to the clustering result to obtain an analysis result, wherein the analysis result is used to characterize the corresponding relationship between the molecular structure and the molecular property.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.