Method, apparatus, device, and program for generating structure data
By employing structural and node feature representations and using GCN with wavelet transforms, the method addresses the limitations of GCN-based molecular generation, enhancing the efficiency and diversity of generated molecular structures.
Patent Information
- Application Number
- JP2024530434
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-02-17
- Filing Date
- 2022-12-05
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2042-12-05
AI Technical Summary
Existing methods for generating molecular structures using graph convolutional networks (GCN) suffer from low-pass filtering, leading to insufficient diversity and efficiency in generating new molecules due to the smoothing of graph data, which hinders the reconstruction of the complete original signal during decoding.
A method involving obtaining structural and node feature representations, generating a hidden layer feature representation in multiple frequency bands using GCN and wavelet transforms, and training a decoder to reconstruct the structure, thereby improving generation efficiency and diversity.
The approach allows for quick generation of various reconstructed structures with enhanced diversity and efficiency by iteratively training the decoder, ensuring accurate reconstruction of molecular structures while maintaining chemical rules.
Smart Images

Figure 0007713104000020 
Figure 0007713104000021 
Figure 0007713104000022
Abstract
Description
Related Application
[0001] This application claims priority to a Chinese patent application with an application number of 202210146218.2 and an invention title of "Method, Apparatus, Device, Medium, and Program Product for Generating Structural Data", which was filed on February 17, 2022, and all of its contents are incorporated herein by reference.
Technical Field
[0002] This application relates to the field of artificial intelligence, and particularly to a method, apparatus, device, medium, and program product for generating structural data.
Background Art
[0003] With the development of artificial intelligence (AI), AI has been applied in more and more fields. Here, in the field of intelligent medicine, AI can promote drug discovery and assist experts in researching and developing new drugs.
[0004] In related technologies, the chemical molecular structure is mapped to generate a corresponding molecular graph of the graph structure, and then, through a graph convolutional network (GCN), these molecular graphs are learned based on the message propagation process, and further, a new feature representation is generated by the GCN. In the decision-making process, a new structure corresponding to the new feature representation is added to the conventional graph so as to conform to the organic molecular chemistry rules, thereby obtaining a molecular graph corresponding to a new molecule.
[0005] However, in the above process of generating the structure of a new molecule, due to the low-pass characteristic of the GCN, the graph data representing the molecule is smoothed, so that the complete original signal cannot be reconstructed during decoding, and finally the diversity and effectiveness of the generated molecules are insufficient, and the generation efficiency is low.
Summary of the Invention
Problems to be Solved by the Invention
[0006] Embodiments of the present application provide a method, apparatus, device, medium, and program product for generating structural data that can improve the generation efficiency of a specified structure and the diversity of the generated structure. The present application adopts the following technical solutions.
Means for Solving the Problem
[0007] According to one aspect of the present invention, obtaining a structural feature representation and a node feature representation of sample structural data, wherein the structural feature representation is for indicating the connection status between nodes constituting the sample structural data, and the node feature representation is for indicating the node type corresponding to the nodes constituting the sample structural data; generating a hidden layer feature representation based on the structural feature representation and the node feature representation, wherein the hidden layer feature representation is for indicating the coupling status between nodes in the sample structural data in at least two frequency bands; inputting the hidden layer feature representation into a decoder to be trained to reconstruct a structure, thereby obtaining predicted structural data; training the decoder to be trained based on the predicted structural data to obtain a specified decoder, wherein the specified decoder is for reconstructing a structure for input sampled data to obtain reconstructed structural data, and the sampled data is data obtained by sampling candidate data. A method for generating structural data is provided, including the above steps.
[0008] According to another aspect of the present invention, an acquisition module for acquiring a structural feature representation and a node feature representation of sample structural data, wherein the structural feature representation is for indicating the connection status between nodes constituting the sample structural data, and the node feature representation is for indicating the node type corresponding to the nodes constituting the sample structural data; An encoding module for generating a hidden layer feature representation based on the structural feature representation and the node feature representation, wherein the hidden layer feature representation is for indicating the connection situation between nodes in the sample structure data in at least two frequency bands, and the encoding module; A decoding module for obtaining predicted structure data by inputting the hidden layer feature representation into a decoder waiting to be trained to reconstruct the structure; A training module for obtaining a specified decoder by training the decoder waiting to be trained based on the predicted structure data, wherein the specified decoder is for reconstructing the structure for the input sampling data to obtain reconstructed structure data, and the sampling data is data obtained by sampling candidate data, and the training module; A structure data generation device is provided.
[0009] According to another aspect of the present invention, there is provided a computer device including a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set for implementing the structure data generation method according to any one of the embodiments of the present application when loaded and executed by the processor.
[0010] According to another aspect of the present invention, there is provided a computer-readable storage medium storing at least one program code for implementing the structure data generation method according to any one of the embodiments of the present application when loaded and executed by a processor.
[0011] According to another aspect of the present invention, there is provided a computer program product or computer program including computer instructions stored in a computer-readable storage medium. When the processor of the computer device reads and executes the computer instructions from the computer-readable storage medium, the computer device is caused to execute the structure data generation method according to any of the above embodiments.
[0012] According to the technical solution of the present application, at least the following beneficial effects can be achieved.
Effect of the Invention
[0013] After obtaining the hidden layer feature representation by the structure feature representation and node feature representation corresponding to the sample structure data, and then performing iterative training on the decoder to be trained according to the hidden layer feature representation to obtain the specified decoder, the required structure data can be generated by the sampling data input to the specified decoder. That is, if necessary, various reconstructed structure data can be quickly generated by the specified decoder obtained through training, and the generation efficiency and generation diversity of the structure data can be improved.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0015] First, terms related to the embodiments of the present application will be briefly described.
[0016] Artificial intelligence: It is a theory, method, technology, and application system that uses digital computers or devices controlled by digital computers to simulate, extend, and expand human intelligence, and performs environmental perception, knowledge acquisition, and acquisition of optimal results using knowledge. In other words, artificial intelligence is an integrated technology of computer science, aiming to understand the essence of intelligence and manufacture new intelligent machines that can react in a manner similar to human intelligence. That is, artificial intelligence gives devices the functions of perception, inference, and decision-making by studying the design principles and implementation methods of various intelligent machines.
[0017] Artificial intelligence technology is an integrated discipline related to a wide range of fields, including both hardware technology and software technology. The basic technologies of artificial intelligence generally include, for example, technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0018] Machine Learning (ML) is an interdisciplinary field that intersects multiple domains and is related to many disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in studying how computers simulate or realize human learning behaviors, acquire new knowledge and skills, and continuously improve their performance by reorganizing existing knowledge structures. Machine learning is the core of artificial intelligence and the fundamental route for endowing computers with intelligence, and it is applied across various fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, trust networks, reinforcement learning, transfer learning, inductive learning, and supervised learning.
[0019] Graph Convolutional Network: Convolutional neural networks adopt local perception regions, shared weight values, and downsampling in the spatial domain, have stable characteristics against displacement, scaling, and distortion, and can extract spatial features of images well. The graph structure does not have translational invariance of pictures, and the conventional convolution method cannot be applied to the graph structure. Therefore, an important difficulty in graph convolution is that the number of neighboring nodes of each node in the graph does not match, and it is impossible to extract features with a convolution kernel of the same size. GCN completes the integration of neighboring information through a message passing mechanism in the spatial or spectral domain and performs main feature extraction. The most common GCN performs low-pass filtering on graph signals.
[0020] Wavelet Transform: It is a local analysis of spatial frequency. Through dilation and translation operations, it gradually performs multi-scale subdivision on signals, finally reaching the subdivision of frequency bands, and can automatically adapt to the requirements of time-frequency signal analysis. Thereby, it can focus on any detail of the signal and solve the difficult problems of Fourier transform.
[0021] Variational Auto-Encoder (VAE): It is a deep learning model for data generation. First, it compresses and encodes the input data, calculates and generates hidden variables, and finally restores the original data by a decoder. When generating data, it can generate data close to the distribution of the original data only by sampling from a specific distribution from the hidden variables. Applying the VAE model to molecular generation is to generate effective molecules with properties consistent with the reference molecules and further discover high-quality drugs.
[0022] In the embodiments of the present application, machine learning / deep learning in artificial intelligence technology is applied to the generation of structured data having certain rules or satisfying certain rules.
[0023] Next, the application scenario of the structured data generation method according to the embodiments of the present application will be described by way of example.
[0024] First, it can be applied to the generation scenario of organic chemical molecules in the intelligent pharmaceutical scenario. In the intelligent pharmaceutical scenario, AI assists in the discovery and research and development of new drugs, such as the generation of lead drugs and the optimization of drugs. Here, the above-mentioned lead drug refers to a chemical compound drug having certain activities and chemical structures obtained by certain routes and means, and is used for further structural modification and modification, which is the starting point of new drug research. Drug optimization refers to improving the physicochemical properties of drugs by optimizing the structure according to certain rules for the chemical structure of drugs.
[0025] In the related art, after mapping the chemical molecules of a compound to a molecular graph with a graph structure, feature extraction is performed by a graph convolutional neural network GCN, and then the structure is restored by GCN. That is, the intermediate feature Z = GCN(X,A) is obtained by feature extraction. Here, X is the node feature of the molecular graph, and A is the edge feature of the molecular graph. Then TIFF0007713104000001.tif6170 is generated,, TIFF0007713104000002.tif6170 is the feature after performing the specified conversion on A. However, in the realization process of this method, the interpretability of the decoding method used in the generation process of new molecules is poor, and there is no decoupling principle that is dual to the encoding part. Therefore, second-order smoothing is performed on the graph signal, resulting in low diversity of the generated molecular graphs and low generation efficiency.
[0026] Exemplarily, by learning the chemical molecular structure of an existing drug through the structure data generation method according to the embodiments of the present application, a specified decoder can be obtained. Here, since the decoding process and the encoding process are dual, in the new drug research and development process, high-quality clinical candidate molecules can be efficiently and effectively generated by the specified decoder to support the new drug research and development. Or, it can be applied to the screening process of similar drug molecules having relatively strong potential activity in a known target. Here, the similar drug molecules are those in which the compound corresponding to the molecule has a certain similarity to the known drug and may become a drug. By the structure data generation method according to the embodiments of the present application, for a target that is difficult to become a drug, ideal candidate molecules with a high success rate can be generated.
[0027] Second, it can be applied to the scene of knowledge graph mining and construction. Here, the above knowledge graph is composed of several entities connected to each other and their attributes, and by combining the theories and methods of disciplines such as applied mathematics, graphics, information visualization technology, and information science with methods such as citation analysis and co-occurrence analysis in metrics, the core structure, cutting-edge fields, and overall knowledge architecture of the discipline are visually displayed using the visualized graph to achieve the purpose of multi-disciplinary integration. Specifically, for example, in the intelligent medical scene, the medical knowledge graph, exemplarily, obtains a specified decoder by training an already constructed knowledge graph corresponding to the disease state as training data, and the specified decoder efficiently generates a plurality of knowledge graphs with a certain effectiveness.
[0028] Third, it can be applied to the scenario of intelligent travel automatic planning. Exemplarily, according to the method for generating structural data according to the embodiments of the present application, a specified decoder is obtained by training with a travel course planning graph for training, and the user can generate various travel plan courses by means of the specified decoder, that is, various travel course plans can be provided to the user under specified conditions or random conditions, enriching intelligent travel. The above-mentioned specified conditions may be a specified travel city, a specified type of tourist attraction, etc.
[0029] The above exemplary scenario is only an exemplification of the application scenario of the method for generating structural data according to the embodiments of the present application. The method may be applied to scenarios where information such as a recommendation system based on social relationships between users, text semantic analysis, and road condition prediction can be processed into graph-structured data, and the specific application scenario is not limited here.
[0030] Regarding the implementation environment of the embodiments of the present application, it will be described in combination with the above description of the noun interpretation and application scenario. As shown in FIG. 1, the computer system of the implementation environment includes a terminal device 110, a server 120, and a communication network 130.
[0031] The terminal device 110 includes various types of devices such as mobile phones, tablet computers, desktop computers, mobile notebook computers, smart home appliances, in-vehicle terminals, and airplanes. Exemplarily, the user instructs the server 120 to train the decoder waiting to be trained via the terminal device 110.
[0032] Server 120 is for providing a training function for the decoder waiting for training. That is, server 120 can call the corresponding operation module according to the request of terminal device 110 to train the specified decoder waiting for training. Preferably, the model architecture corresponding to the decoder waiting for training may be pre-stored in server 120 or may be uploaded by terminal device 110 through a model data file. The training dataset used for training the decoder waiting for training may be pre-stored in server 120 or may be uploaded by terminal device 110 through a training data file. In one example, the user uploads a dataset corresponding to sample structure data to server 120 through terminal device 110 and sends a training request for the decoder waiting for training including the model identifier (ID) of the decoder waiting for training. Server 120 reads the model architecture of the decoder waiting for training corresponding to the model ID from the database based on the model ID in the training request and trains the decoder waiting for training with the received dataset.
[0033] Here, in the training process, server 120 obtains a hidden layer feature representation based on the structural feature representation and node feature representation of the sample structure data, and trains the decoder waiting for training based on the hidden layer feature representation to obtain a specified decoder. When server 120 performs training to obtain a specified decoder, server 120 can send the specified decoder to terminal device 110 or set the specified decoder in the application module so that terminal device 110 can call it according to a data generation request.
[0034] In some embodiments, if the computing power of terminal device 110 is sufficient for the training process of the decoder waiting for training described above, the entire training process of the specified decoder may be independently realized by terminal device 110.
[0035] Note that the server 120 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0036] Here, cloud technology refers to a trusted technology that unifies resources such as hardware, software, and networks within a wide area network or a local area network to realize data calculation, storage, processing, and sharing.
[0037] In some embodiments, the server 120 may be implemented as a node in a blockchain system.
[0038] Exemplarily, the terminal device 110 and the server 120 are connected via a communication network 130. Here, the communication network 130 may be a wired network or a wireless network, and is not limited herein.
[0039] FIG. 2 shows a method for generating structural data according to an embodiment of the present application. In the embodiment of the present application, the method is executed by a computer device, and the computer device may be implemented as the terminal device or the server in FIG. 1. In one example, the method is applied to the server shown in FIG. 1 and includes the following steps.
[0040] In step 201, the structural feature representation and the node feature representation of the sample structural data are obtained.
[0041] Here, the above structural feature representation is for indicating the connection status between nodes that make up the sample structure data, and the node feature representation is for indicating the node type corresponding to the nodes that make up the sample structure data.
[0042] Exemplarily, the above sample structure data is training data for training a decoder to be trained, and the sample structure data is data whose data structure is a graph structure, that is, the sample structure data is data consisting of at least two nodes and at least one edge. Preferably, the above graph structure may be any of a non-directed graph, a directed graph, a non-directed complete graph, a directed complete graph, etc., and the specific graph structure can be determined based on the data information corresponding to the sample structure data. For example, when it is necessary to represent a chemical molecule by a graph structure, that is, when the sample structure data corresponds to a chemical molecule, the atoms in the molecule become the nodes in the graph, and the chemical bonds between the atoms become the edges in the graph. Since the edges do not need to indicate a direction, correspondingly, a non-directed graph can be used as the data structure corresponding to the sample structure data.
[0043] Here, the structural feature representation is for indicating the connection status between nodes in the graph corresponding to the sample structure data, and the connection status is related to the structure generation task corresponding to the sample structure data. For example, when the structure generation task is the generation of a chemical molecule, the connection relationship between the above nodes is the chemical bond between each atom in the chemical molecule. When the structure generation task is a recommendation system based on a social network, the connection status between nodes is the mutual relationship between users in the social network (for example, stranger relationship, friend relationship, blacklist relationship, etc.). When the structure generation task is the generation of a travel route, the connection status between nodes is the route status between tourist attractions.
[0044] The node feature representation is for indicating the node type of each node in the graph corresponding to the sample structure data, and the node type is related to the structure generation task corresponding to the sample structure data. For example, when the structure generation task is the generation of chemical molecules, the above node type is the atomic type in the chemical molecule. When the structure generation task is a recommendation system based on a social network, the node type is the user account in the social network. When the structure generation task is the generation of a travel route, the node type is a tourist destination.
[0045] Exemplarily, the structure feature representation and the node feature representation of the sample structure data are obtained by converting the sample structure data according to a preset feature conversion method. Preferably, the above structure feature representation may be a feature representation in matrix form or a feature representation in vector form, and the above node feature representation may be a feature representation in matrix form or a feature representation in vector form, which is not limited herein.
[0046] Preferably, the above preset conversion method may be a network conversion method, that is, the above structure feature representation and node feature representation may also be obtained by performing feature extraction by a feature extraction network. Input the sample structure data into a pre-trained feature extraction network to output the structure feature representation and the node feature representation. The above feature extraction network may be a network capable of completing feature extraction, such as a Convolutional Neural Networks (CNN), a Visual Geometry Group Network (VGGNet), an Alex Network (AlexNet), etc., which is not limited herein.
[0047] Preferably, the above-mentioned preset conversion method may be a matrix conversion method, that is, the above-mentioned structural feature representation and node feature representation may be obtained by performing matrix conversion on the graph structure data corresponding to the sample structure data. Exemplarily, the logical structure of the sample structure data of the above graph structure is divided into two parts: a node set consisting of at least two nodes and an edge set consisting of edges between the nodes. The above structural feature representation is an adjacency matrix, which is two-dimensional data generated by the edge set and used to store edges, and the adjacency matrix is for recording the connection relationship between the above at least two nodes. The above node feature representation is a one-dimensional matrix (array) generated by the node set and is for storing node data in the graph.
[0048] In one example, taking the case where the sample structure data is for indicating a chemical molecule, what is recorded in the adjacent matrix is the chemical bond type between atoms in the chemical molecule, and the node feature representation is a one-dimensional feature matrix generated based on the constituent atoms of the chemical molecule and is for recording the atomic type in the chemical molecule. Exemplarily, a sample chemical molecule consisting of at least two atoms is obtained, and the sample chemical molecule is a known molecule that satisfies the criteria for atomic bonding. The sample chemical molecule is converted into a sample molecular graph whose data structure is a graph structure. The nodes of the sample molecular graph are for representing at least two atoms in the sample chemical molecule, such as carbon atoms, hydrogen atoms, oxygen atoms, etc. The edges in the sample molecular graph are for indicating the chemical bonds between atoms in the sample chemical molecule. The chemical bonds include types such as edgeless, single bond, double bond, triple bond, etc. Among them, in the computer, the atomic type and the chemical bond type can be mapped to different characters or character strings according to a specified mapping relationship. For example, the edgeless can be corresponded to "0", the single bond to "1", the double bond to "2", and the triple bond to "3" respectively. The above mapping relationship may be recorded in a preset mapping table. The adjacent matrix corresponding to the sample molecular graph is determined as the structural feature representation, and the node matrix corresponding to the sample molecular graph is determined as the node feature representation. Since the molecular graph with a graph structure can simply and clearly represent the connection relationship between atoms in the chemical molecule, the acquisition efficiency of the sample molecular graph is improved, and in the process of feature extraction, the features of the atomic type and the chemical bonds between atoms can be retained.
[0049] In the embodiments of the present application, the decoder to be trained is a part of the training model. Exemplarily, by inputting the structural feature representation and the node feature representation corresponding to the sample structure data into the training model, predicted structure data is output, and the entire training model is trained by the deviation between the predicted structure data and the sample structure data. That is, the training of the decoder to be trained is completed in the overall training process for the training model.
[0050] In step 202, a hidden layer feature representation is generated based on the structural feature representation and the node feature representation.
[0051] The above hidden layer feature representation is for indicating the connection status between nodes in sample structure data in at least two frequency bands.
[0052] Exemplarily, the above training model further includes a to-be-trained encoder, and the to-be-trained encoder is for generating a hidden layer feature representation based on the structural feature representation and the node feature representation. Preferably, the encoder structure corresponding to the above to-be-trained encoder may be an autoencoder, a variational auto-encoder (VAE), a low-pass filter, a band-pass filter, etc., and the specifically used filter may be a convolutional neural network, a wavelet filter, a Butterworth filter, a Bessel filter, etc., which is not limited here.
[0053] In the embodiment of the present application, when the sample structure data is graph-structured data, the to-be-trained encoder is a GCN, that is, the structural feature representation and the node feature representation are used as the input of the GCN, and the above hidden layer feature representation is obtained by outputting.
[0054] Exemplarily, since the above hidden layer feature representation is for indicating the connection status between nodes in sample structure data in at least two frequency bands, the GCN in the embodiment of the present application includes a first filtering layer and at least two second filtering layers.
[0055] Preferably, the first filtering layer may be for completing low-pass filtering, high-pass filtering, or band-pass filtering. Specifically, it needs to be determined according to the needs of the actual application scenario. In one example, when applied to the chemical molecule generation scenario, the first filtering layer is a low-pass GCN layer. Here, the first filtering layer only indicates a filtering layer with a consistent function. The first filtering layer may be composed of a single filtering layer or multiple filtering layers. For example, the first filtering layer includes two low-pass GCN layers, that is, it is composed of two layers of neurons, but this is not limited here.
[0056] The second filtering layer is a band-pass filtering layer corresponding to at least two frequency bands. Exemplarily, the data output from the first filtering layer is input into the at least two second filtering layers, and filtering results corresponding to each frequency band are output. Exemplarily, the number of second filtering layers corresponds to the number of divisions of the frequency band. Here, the second filtering layer only indicates a filtering layer with a consistent function. The at least two second filtering layers are at least two second filtering layers connected in parallel. One second filtering layer may be composed of a single filtering layer or multiple filtering layers, but this is not limited here.
[0057] Exemplarily, based on the structural feature representation and the node feature representation, encoding is performed in at least two frequency bands respectively to obtain intermediate feature data. The intermediate feature data is for indicating the connection status between nodes of the sample structure data in the corresponding frequency band. Based on the specified data distribution, the intermediate feature data corresponding to at least two frequency bands is clustered respectively to obtain the hidden layer feature representation. That is, the intermediate feature data in at least two frequency bands is obtained by the above-mentioned first filtering layer and the second filtering layer, and the hidden layer feature representation is obtained by clustering the intermediate feature data.
[0058] Preferably, the clustering method of the above intermediate feature data may be splicing based on the frequency band order between at least two frequency bands. In one example, the above frequency band order may be the order of arranging at least two frequency bands from low frequency to high frequency. Or, the clustering method of the above intermediate feature data is to fit the intermediate feature data corresponding to at least two frequency bands respectively based on the specified data distribution. For example, it may be fitting based on a normal distribution (Gaussian distribution), fitting by a Chebyshev polynomial, fitting based on the least squares method, etc. In some embodiments, when the number of nodes in the sample structure data is small, that is, when the computing power can satisfy the calculation requirements of the structure matrix and the feature matrix, eigenvalue decomposition may be used instead of polynomial fitting, that is, eigenvalue decomposition may be performed on the Laplacian matrix.
[0059] In step 203, by inputting the hidden layer feature representation into the decoder waiting to be trained to reconstruct the structure, predicted structure data is obtained.
[0060] Exemplarily, the hidden layer feature representation obtains predicted structure data by reconstructing the structure by the decoder. In some embodiments, the output of the decoder to be trained is the decoded structure feature representation and the decoded node feature representation, that is, by sampling the hidden layer feature representation by the decoder to be trained, the decoded structure feature representation and the decoded node feature representation are obtained, and based on the decoded structure feature representation and the decoded node feature representation, predicted structure data is generated. Here, the decoded structure feature representation is for indicating the relationship between nodes in the predicted structure data, and the decoded node feature representation is for indicating the nodes in the predicted structure data.
[0061] In one example, FIG. 3 is a diagram showing a model structure according to one exemplary embodiment of the present application. As shown in FIG. 3, the structure feature representation and the node feature representation of the sample structure data are input into the first filtering layer 310, and then the output result of the first filtering layer 310 is input into the second filtering layer 320. The output of the second filtering layer 320 is clustered to obtain the hidden layer feature representation, and the hidden layer feature representation is decoded by the decoder 330 to obtain the decoding result, and the decoding result is the above-mentioned decoded structure feature representation and decoded node feature representation.
[0062] In step 204, by training the decoder to be trained based on the predicted structure data, a specified decoder is obtained.
[0063] Here, the specified decoder is for reconstructing the structure for the input sampled data to obtain the reconstructed structure data, and the sampled data is data obtained by sampling candidate data. In some embodiments, the candidate data is data that satisfies the specified data distribution.
[0064] In the embodiments of the present application, the decoder waiting for training is trained according to the structural difference situation between the predicted structure data output from the decoder waiting for training and the input sample structure data until the decoder waiting for training converges. In some embodiments, since the decoder waiting for training is a part of the training model, its training process depends on the overall training process of the training model, that is, by training the training model according to the structural difference situation between the predicted structure data and the input sample structure data, a converged prediction model is obtained, and the decoder part in the prediction model is decomposed into the specified decoder and used for generating the reconstructed structure data.
[0065] Exemplarily, a training loss value is obtained based on the structural difference situation between the sample structure data and the predicted structure data. In response to the fact that the training loss value reaches the specified loss threshold, that is, it is determined that the model has converged after being trained, it is determined that the training of the decoder waiting for training is completed, and the specified decoder is obtained. Alternatively, in response to the failure of the matching between the training loss value and the specified loss threshold, iterative training is performed on the model parameters of the decoder waiting for training, that is, iterative training is performed by adjusting the model parameters of the training model. Here, the specified loss threshold may be set by the system in advance or customized according to the user's needs. For example, the higher the required model accuracy, the smaller the corresponding specified loss threshold.
[0066] The training loss value is calculated by a specified loss function, and the specified loss function may be a loss function used for regression, reconstruction, and classification. Preferably, the specified loss function may be a loss function such as the mean absolute error loss function, the negative log-likelihood loss function, the exponential loss function, the cross-entropy loss function, and its variants, but is not limited thereto here.
[0067] As described above, the method for generating structural data according to the embodiment of the present application obtains a hidden layer feature representation by means of a structural feature representation and a node feature representation corresponding to sample structural data, and then performs iterative training on a decoder to be trained based on the hidden layer feature representation to obtain a specified decoder. By doing so, the necessary structural data can be generated by the sampling data input to the specified decoder. That is, if necessary, various reconstructed structural data can be quickly generated by the specified decoder obtained through training, improving the generation efficiency and generation diversity of the structural data.
[0068] FIG. 4 shows a method for generating a hidden layer feature representation according to an embodiment of the present application. In the embodiment of the present application, the process of obtaining the hidden layer feature representation by an encoder will be described. The method includes the following steps.
[0069] In step 401, a structural feature representation and a node feature representation of sample structural data are obtained.
[0070] In the embodiment of the present application, the overall training model includes an encoder part used for estimation and a decoder part used for generation. The structural feature representation and the node feature representation of the sample structural data are the inputs of the overall training model.
[0071] Here, the above structural feature representation is for indicating the connection status between nodes constituting the sample structural data, and the node feature representation is for indicating the node type corresponding to the nodes constituting the sample structural data. The above sample structural data is training data for training the training model, and the sample structural data is data with a graph structure as the data structure.
[0072] In step 402, intermediate feature data is obtained by encoding respectively in at least two frequency bands based on the structural feature representation and the node feature representation.
[0073] In the embodiments of the present application, intermediate feature data is obtained by encoding the structural feature representation and the node feature representation by the encoder waiting for training in the training model. Here, the structure of the encoder waiting for training includes a first filtering layer and at least two second filtering layers.
[0074] In some embodiments, the encoding process is completed by wavelet transform, that is, the structure of the encoder waiting for training is a GCN. Here, the first filtering layer is a low-pass GCN layer, and the second filtering layer is a band-pass wavelet layer. Here, each of the at least two band-pass wavelet layers corresponds to a different wavelet basis function, that is, the signal is filtered by the wavelet basis function, that is, referring to the multi-scale principle in wavelet transform, it is shown that it is the basis of convolutional band-pass filtering based on the Taylor expansion of different wavelet basis functions, and the wavelet transform process is a convolution process between the input feature and the wavelet basis function.
[0075] Exemplarily, a scale criterion is obtained based on the structure generation task to be completed by the training model, and at least two corresponding basis functions are calculated based on the scale criterion. The at least two basis functions form a basis function group, and each basis function in the basis function group corresponds to one band-pass wavelet layer, that is, corresponds to one frequency band.
[0076] In step 403, a hidden layer feature representation is obtained by clustering the intermediate feature data corresponding to each of at least two frequency bands based on the specified data distribution.
[0077] In some embodiments, the outputs of at least two band-pass wavelet layers are directly clustered to obtain a hidden layer feature representation. That is, in this case, the intermediate feature data is the encoding result obtained by encoding with the encoder to be trained. Exemplarily, the structural feature representation and the node feature representation are input to the encoder to be trained, and intermediate feature data corresponding to each of at least two frequency bands is output. The intermediate feature data is for clustering to obtain a hidden layer feature representation. That is, in the process of generating a hidden layer feature representation for the structural feature representation and the node feature representation, intermediate feature data is generated by the encoder to be trained, and by jointly training the encoder to be trained and the decoder to be trained, the training result is backpropagated in the training process of the model to assist in the training of the decoder to be trained.
[0078] Here, the clustering method of the intermediate feature data is splicing based on the frequency band order between at least two frequency bands. In one example, the frequency band order may be an order of arranging at least two frequency bands from low frequency to high frequency.
[0079] Exemplarily, FIG. 5 is a diagram showing the acquisition of a hidden layer feature representation according to an exemplary embodiment of the present application. The structural feature representation and the node feature representation 501 are input to GCN0510 to obtain low-pass filtering data 502, and the low-pass filtering data 502 is input to GCN wavelet 520 (GCN wavelet1 、GCN wavelet2 、…、GCN waveletn ), and one intermediate feature data 503 is output from each of GCN wavelet 520, and the intermediate feature data 503 is clustered to obtain a hidden layer feature representation Z504.
[0080] In some other embodiments, in order to present the diversity of the reconstructed structure data generated by the specified decoder obtained by training, the intermediate feature data may be data obtained by performing intermediate calculations on the features obtained by encoding with the encoder to be trained.
[0081] In one example, the data obtained by the above intermediate calculation are the mean value and the variance. Exemplarily, based on the structural feature representation and the node feature representation, the node feature vectors of the nodes of the sample structure data in the feature spaces respectively corresponding to at least two frequency bands are obtained, the mean value data and the variance data between the node feature vectors corresponding to at least two frequency bands are obtained, and the mean value data and the variance data are determined as the intermediate feature data. That is, after filtering by the encoder to be trained, the mean value and the variance are calculated for the node feature vectors corresponding to each frequency band to obtain the mean value data and the variance data, and the mean value data and the variance data are the intermediate feature data for generating the hidden layer feature representation. That is, by realizing the variational estimation process of variational auto-encoding based on the mean value data and the variance data between the node feature vectors corresponding to different frequency bands, the interpretability of the specified decoder obtained by downstream training is improved, and the generation space of the structural data is expanded.
[0082] Specifically, the encoder to be trained is an encoder based on the probability model shown in Equation 1, where Z represents the hidden layer feature representation of the node, X represents the feature matrix corresponding to the node feature representation, the dimension of X is N*1, N is the number of nodes in the sample structure data, N is a positive integer, and A is the adjacency matrix for storing the edges between nodes corresponding to the structural feature representation.
[0083]
Number
[0084] Here, q(z i |X,A) is determined by Equation 2, μ represents the mean value of the node vector representation, σ represents the variance of the node vector representation, diag() is to generate a diagonal matrix, and Equation 2 represents fitting the features corresponding to the node z i to a Gaussian distribution.
[0085]
Number
[0086] As can be seen from Equation 1 and Equation 2, when the encoder waiting for training takes the structural feature representation and the node feature representation as known conditions, the i-th prediction node in the prediction node set determines the connection probability of establishing a connection relationship with each prediction node in the prediction node set, and determines the connection probability distribution corresponding to the i-th prediction node based on the connection probabilities between the i-th prediction node and each prediction node. Based on the fusion result of the connection probability distributions of all prediction nodes in the prediction node set, intermediate feature data (distributions corresponding to the hidden layer feature representations) corresponding to at least two frequency bands are determined respectively. Here, the above prediction nodes are for constructing the finally output prediction structure data. Exemplarily, each prediction node in the prediction node set corresponds to one connection probability distribution, and the distribution obtained by continuously multiplying the connection probability distributions corresponding to each prediction node can be determined as the distribution corresponding to the hidden layer feature representation Z. In the embodiments of the present application, the distributions of the frequency band hidden layer feature representations respectively correspond to at least two frequency bands, and by fusing the distributions of the frequency band hidden layer feature representations in the different frequency bands, the distribution corresponding to the hidden layer feature representation can be obtained.
[0087] Taking the case where the structure data is a chemical molecule as an example, a set of different types of chemical atoms is used as the prediction node set. For whether the i-th type of chemical atom can establish a connection relationship with all types of chemical atoms including itself in the prediction node set, it is associated with a probability value of establishing one connection relationship. When comprehensively observing the connection relationships between the i-th type of chemical atom and all types of chemical atoms in the set, the above-mentioned connection probability distribution, which is the distribution of the above-mentioned probability values corresponding to the i-th type of chemical atom, is formed. Each chemical atom corresponds to one connection probability distribution, and the connection probability distribution is fitted to a Gaussian distribution in the training process of the encoder waiting to be trained. When the connection probability distribution of each node is fitted to a Gaussian distribution, the distribution of the hidden layer feature representation Z obtained by continuously multiplying the connection probability distributions is also a Gaussian distribution.
[0088] If GCN is adopted in the above probability model, the mean value data and variance data can be obtained through encoding. Taking a single-layer GCN as an example, the computational representation of the single-layer GCN is as shown in Equation 3, where A is the structural feature representation input to the GCN, X is the node feature representation input to the GCN, W0 is the model parameter of the GCN model, and here D is the degree matrix corresponding to the sample structure data of the graph structure. Preferably, the mean value data and variance data may be obtained by a single-layer GCN or by a multi-layer GCN, and this is not limited here.
[0089]
Number
[0090] After obtaining the mean value data and variance data, the mean value data and variance data can be fitted based on the specified data distribution to obtain the hidden layer feature representation. In one example, the specified data distribution is a Gaussian distribution.
[0091] Exemplarily, FIG. 6 is a diagram showing the acquisition of hidden layer feature representation according to another exemplary embodiment of the present application. The structural feature representation and the node feature representation 601 are input into GCN0610 to obtain low-pass filtering data 602, and the low-pass filtering data 602 is input into GCN wavelet 620 (GCN wavelet1 , GCN wavelet2 , …, GCN waveletn included), and corresponding node feature vectors 603 are output from each of GCN wavelet 620, and average value data and variance data 604 corresponding to the node feature vectors 603 are obtained by intermediate calculation. Taking the average value data and the variance data 604 as intermediate feature data, fitting is performed based on the intermediate feature data to obtain a hidden layer feature representation Z605 of a Gaussian distribution.
[0092] In the embodiment of the present application, for GCN0 shown in FIG. 6, the corresponding calculation is as shown in Equation 3, that is, TIFF0007713104000006.tif6170. The node feature representation X and the structural feature representation A are input into GCN0, and are converted by an activation function to obtain X0. The input of GCN wavelet is the above X0 and A, and taking the case where GCN wavelet includes GCN wavelet1 , GCN wavelet2 as an example, GCN wavelet1 (X0,A)=A wavelet1 X0W1, GCN wavelet2 (X0,A)=A wavelet2 X0W2, where W1 is the model parameter of GCN wavelet1 , and W2 is the model parameter of GCN wavelet2 , and A wavelet1 is the wavelet transform result of A with respect to the wavelet basis function corresponding to GCN wavelet1 , and A wavelet2 is the wavelet transform result of A with respect to the wavelet basis function corresponding to GCN wavelet2It is the wavelet transform result of A with respect to the corresponding wavelet basis function. Here, the activation function may be a sigmoid activation function (S-shaped growth curve), a tanh non-linear function, a ReLU activation function, or a modification thereof, and is not limited herein.
[0093] In one example, after determining the mean data and variance data, a reparameterization trick is used for calculation to obtain a hidden layer feature representation. That is, as shown in Equation 4, where μ is the mean data, σ is the variance data, and ε is a normal Gaussian distribution, that is, p(ε)=N(0,I).
[0094]
Number
[0095] As described above, the method for generating a hidden layer feature representation according to the embodiment of the present application performs filtering on the structural feature representation and node feature representation corresponding to the sample structure data by means of GCN to obtain intermediate feature data corresponding to a plurality of frequency bands, and obtains a hidden layer feature representation by clustering the intermediate feature data. That is, the process of graph compression encoding is realized by wavelet transform, and multi-scale subdivision is performed on the features to achieve multi-frequency banding of the features, and the features are focused on the feature details corresponding to a plurality of frequency bands respectively, thereby ensuring the diversity of the reconstructed structure data generated. At the same time, by realizing the encoding prediction process using the data of the graph structure, the accuracy requirement of the data for the reconstruction of the structure is guaranteed.
[0096] FIG. 7 shows a training method for a decoder according to an embodiment of the present application. In the embodiment of the present application, the decoder part in the training network is described. Here, steps 701 to 703 (including steps 7031 and 7032) are realized after step 403, and the method includes the following steps.
[0097] In step 701, the decoder waiting for training reconstructs the structure for the hidden layer feature representation to obtain the decoded structure feature representation and the decoded node feature representation.
[0098] In the embodiment of the present application, the hidden layer feature representation Z is the input of the decoder waiting for training. The decoder waiting for training reconstructs the decoded structure feature representation and the decoded node feature representation according to the probability that there is an edge between nodes in the decoding process. Here, the probability model corresponding to the decoder is as shown in Equation 5, where N is the number of nodes, z i and z j are the nodes in the hidden layer feature representation respectively.
[0099]
Equation
[0100] Here, p(A i,j |z i ,z j ) in the above Equation 5 is obtained by Equation 6. Here, sigmoid() represents the activation function (S-shaped growth curve). TIFF0007713104000009.tif6170 represents performing a transpose operation on z i . The sigmoid activation function used above is only exemplary. In actual applications, other activation functions may be used and are not limited here.
[0101]
Equation
[0102] As can be seen from the above Equation 5 and Equation 6, the obtained TIFF0007713104000011.tif6170 is as shown in Equation 7. Here, Z is the hidden layer feature representation, and σ() is the activation function similar to the above sigmoid().
[0103]
Number
[0104] In the embodiments of the present application, the decoder waiting for training reconstructs the prediction structure data according to the inverse transformation theory for the wavelet basis to obtain the decoded structure feature representation and the decoded node feature representation. Here, the process of completing the reconstruction of the structure by the above wavelet inverse transformation can also be popularized to some other high-frequency basis functions, for example, any wavelet basis without a high-pass filtering function, a mother function, etc.
[0105] In the decoding process, it is necessary to discretize the scale in the wavelet transformation process. In one example, the inverse transformation representation of the kernel function g is defined as g-1, and the inverse function at a = 1 TIFF0007713104000013.tif6170, the inverse function at a = 2 TIFF0007713104000014.tif6170, the inverse function at a = 3 TIFF0007713104000015.tif6170 are respectively obtained correspondingly. Here, the above a is the scale in the wavelet basis function. Then, by performing a third-order Taylor expansion on the three inverse functions, the inverse representation coefficients are obtained. Here, the above division of the scale (that is, a = 1, 2, 3) is only exemplary, and in actual applications, the scale can be divided in different ways, and this is not limited here. By convolving the hidden layer feature representation with the above inverse representation coefficients, the decoded structure feature representation and the decoded node feature representation are obtained.
[0106] In step 702, based on the decoded structure feature representation and the decoded node feature representation, prediction structure data is generated.
[0107] Exemplarily, when the decoded structural feature representation is obtained, the connection relationship between nodes in the predicted structural data can be known. When the decoded node feature representation is obtained, the nodes in the predicted structural data can be known. Based on the nodes and the connection relationship between the nodes, the predicted structural data of the corresponding graph structure can be obtained. The predicted structural data is for participating in the training process of the training model as the output result of training.
[0108] In one example, the decoded structural feature representation is realized as an adjacency matrix obtained by decoding, and the decoded node feature representation is realized as a node matrix obtained by decoding. Here, the adjacency matrix is a matrix for representing the adjacency relationship between nodes, and the node matrix is for indicating the node type corresponding to each node in the predicted structural data. The matrix element at the i-th row and j-th column in the adjacency matrix represents the connection status between the i-th node and the j-th node. For example, when the matrix element at the i-th row and j-th column is 0, it means that there is no edge between the i-th node and the j-th node. Based on the adjacency matrix, the connection status between each node in the node matrix can be determined, and further, the predicted structural data can be generated.
[0109] In step 7031, in response to the failure of the matching between the training loss value between the sample structural data and the predicted structural data and the specified loss threshold, the model parameters of the training model are repeatedly trained.
[0110] Exemplarily, a training loss value is obtained based on the structural difference situation between the sample structure data and the predicted structure data. In the embodiments of the present application, during the training process, the training loss value is determined based on both the distance metric between the generated predicted structure data and the originally input sample structure data and the divergence between the node distribution and the Gaussian distribution, that is, the distance metric data of the sample structure data and the predicted structure data in the feature space is obtained, the divergence data between the node distribution corresponding to the predicted structure data and the specified data distribution is obtained, the node distribution is for indicating the distribution situation of the node feature vectors of the predicted structure data in the feature space, the divergence data is for indicating the difference degree between the node distribution and the specified data distribution, and the training loss value is obtained based on the distance metric data and the divergence data. That is, the feature similarity between the two is indicated based on the distance metric data of the sample structure data and the predicted structure data in the feature space, and the training loss value used for model training is jointly determined based on the divergence between the node distribution of the predicted structure data and the specified data distribution, so as to perform model training from two angles of feature difference and node distribution, improve the prediction accuracy of the structure data by the model during the training process, make the node distribution fit the specified data distribution as much as possible, and improve the diversity of the generated structure data by sampling the specified data distribution downstream to generate the corresponding reconstructed structure data.
[0111] Here, the training loss value is calculated by a specified loss function, and preferably, the specified loss function may be a loss function such as the mean absolute error loss function, negative log-likelihood loss, exponential loss, cross-entropy loss function and its variants.
[0112] Preferably, the distance metric data may be the Euclidean distance, Hamming distance, cosine similarity, Manhattan distance, Chebyshev distance, etc. of the sample structure data and the predicted structure data in the feature space, and this is not limited here.
[0113] In one example, the specified loss function for determining the training loss value is as shown in Equation 8, where TIFF0007713104000016.tif6170 is the cross-entropy loss function between the structural features and the node features, p(Z) is as shown in Equation 9, and KL() is the Kullback-Leibler divergence function.
[0114]
Number
Number
[0115] Exemplarily, when the matching between the calculated training loss value and the preset specified loss threshold fails, the model parameters in the training model are adjusted to train the entire model, so that the training decoder converges as the model converges. Here, the parameters of the training model include the model parameters of the encoder and the model parameters of the decoder.
[0116] In step 7032, in response to the training loss value between the sample structure data and the predicted structure data reaching the specified loss threshold, a prediction model is obtained.
[0117] When the training loss value obtained by the specified loss function reaches the specified loss threshold, it is determined that the training of the entire training model is completed, that is, a prediction model is obtained. Here, the decoder part in the prediction model corresponds to the specified decoder. In one example, since the predicted structure data to be output in the training process needs to be as close as possible to the input sample structure data, in response to the training loss value between the sample structure data and the predicted structure data being smaller than the specified loss threshold, it is determined that the training of the training model is completed, and a prediction model is obtained.
[0118] Preferably, the obtained prediction model through training may be stored in the server and called by the terminal according to a generation request. That is, the terminal sends the generation request to the server, and the server calls the prediction model to generate reconstructed structure data and returns the generated reconstructed structure data to the terminal. Alternatively, the obtained prediction model through training may be sent by the server to the terminal, and the terminal uses the prediction model to generate reconstructed structure data.
[0119] Preferably, in the application process, the generation of reconstructed structure data can be performed by a complete prediction model. Exemplarily, candidate structure data is input into the prediction model, and through the encoding of the prediction model, a hidden layer feature representation is obtained. Further, by performing the reconstruction of the structure through decoding, reconstructed structure data having a structural similarity relationship with the candidate structure data is obtained. This method can be applied to a scenario where it is necessary to generate structure data having a strong similarity with a specified structure. Alternatively, in the application process, the generation of reconstructed structure data may be performed only by a specified decoder in the prediction model. That is, sampling data is obtained by sampling candidate data of a specified data distribution, the sampling data is used as the input of the specified decoder, and the specified decoder performs the reconstruction of the structure based on the sampling data to obtain corresponding reconstructed structure data. This method is applicable to an application scenario where it is necessary to generate reconstructed structure data with unknown properties, and can improve the diversity of the generated structure while ensuring the rationality of the generated structure.
[0120] As described above, the decoder training method according to the embodiment of the present application realizes the training process of the decoder by reconstructing the structure of the hidden layer feature representation obtained by the decoder to obtain corresponding predicted structure data, and training the entire training model based on the difference between the predicted structure data and the sample structure data. Here, the decoder restores the high-frequency characteristics compressed and reduced in the hidden layer according to the inverse transformation process of wavelet transform, thereby realizing the functions of reconstructing high-frequency signals and removing noise. When the data is secondarily smoothed due to performing reconstruction directly using the GCN after filtering using the GCN, there is an accuracy accumulation effect (the final accuracy is proportional to the Nth power of the model prediction accuracy), that is, the prediction accuracy decreases in the encoding process and further decreases in the decoding process, solving the problem that the diversity of the generated structure data is low and the generation rate is low, and improving the accuracy of the reconstruction result of the decoder in the application process.
[0121] Exemplarily, when applying the above decoder training method to the scene of generating chemical molecules and obtaining new chemical molecules by reconstructing atomic types, compared with the method of restoring the high-frequency characteristics compressed and reduced in the hidden layer and directly completing the reconstruction process by the GCN, the root mean squared error (RMSE) corresponding to the method according to the present application can be reduced by about 10%, ensuring the reconstruction accuracy by both predicting atomic types and graph structures, that is, significantly improving the stability of the property of generating new chemical molecules and ensuring the effectiveness of generating new chemical molecules.
[0122] FIG. 8 shows a method for generating structure data according to an embodiment of the present application, schematically explaining the application of the trained prediction model. In the embodiment of the present application, the generation of reconstructed structure data is completed by the trained prediction model. The method includes the following steps.
[0123] In step 801, candidate structure feature representations and candidate node feature representations of candidate structure data are obtained.
[0124] Exemplarily, the candidate structure data is data that waits for the generation of similar structure data. The candidate structure data is data whose data structure is a graph structure. That is, the candidate structure data is data composed of at least two nodes and at least one edge. Preferably, the graph structure may be any of a graph structure such as an undirected graph, a directed graph, an undirected complete graph, a directed complete graph, etc. The specific graph structure can be determined based on the data information corresponding to the candidate structure data.
[0125] In one example, taking the case where a prediction model is used for the generation of chemical molecules as an example, the candidate structure data corresponds to a candidate chemical molecule. Here, the atoms in the molecule are the nodes in the graph, and the chemical bonds between the atoms are the edges in the graph. Exemplarily, a corresponding candidate molecule graph is generated based on the chemical structure of the candidate chemical molecule, and the candidate molecule graph is used as the candidate structure data, and a candidate structure feature representation and a candidate node feature representation can be obtained based on the candidate molecule graph. Here, the candidate structure feature representation is an adjacency matrix that records the connection status between atoms in the candidate chemical molecule, and the candidate node feature representation is a matrix that records the atomic types that make up each atom in the candidate chemical molecule.
[0126] In step 802, a candidate hidden layer feature representation is generated based on the candidate structure feature representation and the candidate node feature representation.
[0127] Exemplarily, the candidate structure feature representation and the candidate node feature representation are encoded by a specified encoder in the prediction model to obtain intermediate encoded data, and the intermediate encoded data is clustered to obtain a candidate hidden layer feature representation. In the embodiments of the present application, the intermediate encoded data is data obtained by performing intermediate calculations based on the features encoded by the specified encoder. That is, the data obtained by the intermediate calculations are the mean value and the variance.
[0128] Based on the candidate structure feature representation and the candidate node feature representation, obtain the node feature vectors of the nodes of the candidate structure data in the feature space corresponding to each of at least two frequency bands, obtain the mean value data and variance data between the node feature vectors corresponding to at least two frequency bands, and determine the mean value data and variance data as intermediate encoded data. After determining the mean value data and variance data, a hidden layer feature representation is obtained by calculation using the reparameterization trick. The specific determination process is the same as that in steps 402-403. Since this is the data processing process in the application stage in this embodiment, the description thereof is omitted here.
[0129] In step 803, by inputting the candidate hidden layer feature representation into the specified decoder for prediction, reconstructed structure data is obtained.
[0130] There is a structural property similarity relationship between the above reconstructed structure data and the candidate structure data.
[0131] Exemplarily, by inputting the candidate hidden layer feature representation into the specified decoder and having the specified decoder perform structure reconstruction based on the candidate hidden layer feature representation, a reconstructed structural feature representation and a reconstructed node feature representation can be obtained. Based on the reconstructed structural feature representation and the reconstructed node feature representation, reconstructed structure data is obtained. Specifically, taking the case where the candidate structure data is a candidate molecular graph corresponding to a candidate chemical molecule as an example, the predicted reconstructed structure data is a candidate molecular structure, and the candidate molecular structure is similar in chemical properties to the input candidate chemical molecule. When applied correspondingly to the intelligent medicine scenario, there may be a case where the drug properties are similar between the candidate molecular structure and the input candidate chemical molecule, thereby assisting in the process of studying alternative drugs and optimizing drugs.
[0132] As described above, the method for generating structural data according to the embodiment of the present application generates new structural data having a structural property similarity relationship by encoding and reconstructing the input candidate structural data using the complete prediction model obtained through training. Here, since the prediction model is a model that employs wavelet encoding and decoding, it plays the role of reconstructing high-frequency signals and removing noise, improves the accuracy of the reconstructed structural data obtained through reconstruction, guarantees the structural property similarity between the input and output structural data, and can be applied to the generation of similar structural data.
[0133] FIG. 9 shows a method for generating structural data according to an embodiment of the present application, and schematically illustrates the application of a specified decoder obtained through training. In the embodiment of the present application, the generation of the reconstructed structural data is completed by the specified decoder obtained through training, and the method includes the following steps.
[0134] In step 901, candidate data of a specified data distribution is acquired.
[0135] In the embodiment of the present application, after the training of the prediction model is completed, the specified decoder in the prediction model is separated and applied as a model for generating reconstructed structural data. The above-mentioned specified data distribution may be a Gaussian distribution, the candidate data may be custom candidate data input from a terminal, and may also be curve data of structural properties corresponding to a structure generation task after being trained until the prediction model converges. The above-mentioned structure generation task is a task corresponding to the prediction model, that is, the above-mentioned curve data is data obtained by learning the sample structural data in the training set. In some embodiments, the above-mentioned candidate data may also be candidate data generated after encoding the input candidate structural data to obtain mean value data and variance data.
[0136] In step 902, sampling is performed from the candidate data to obtain a preset number of sampling data.
[0137] Exemplarily, candidate data is randomly sampled to obtain a preset number of sampling data, where the preset number is the quantity of reconstructed structure data that needs to be generated by a specified decoder as instructed by the terminal. The sampling data obtained by the above sampling is for instructing the hidden layer representation between nodes and edges in the reconstructed structure data to be generated. Here, the quantity of corresponding nodes in each sampling data may be randomly generated or specified, and the quantity of nodes among the preset number of sampling data may be the same or different. In one example, sampling data Z input to the specified decoder is obtained by sampling according to Equation 10, where N(0, I) represents candidate data following a normal distribution.
[0138]
Number
[0139] In step 903, a preset number of sampling data is input to the specified decoder to obtain a preset number of reconstructed structure data.
[0140] When the sampled sampling data Z is input to the specified decoder, the specified decoder makes node predictions based on the sampling data and reconstructs the structure feature representation and node feature representation corresponding to the reconstructed structure data according to the probability that there is an edge between nodes, and obtains the reconstructed structure data from the structure feature representation and node feature representation.
[0141] In one example, the designated decoder is used for the generation of chemical molecules. That is, the candidate molecular structure is generated by the designated decoder, where the candidate molecular structure is composed of at least two atomic nodes. Exemplarily, a preset number of sampling data are input into the designated decoder, and a preset number of candidate molecular structures are obtained based on the connection relationship between the atomic nodes in the molecular structure learned by the designated decoder during training. That is, the designated decoder can generate a candidate molecular structure that satisfies chemical rules based on the sampling data. The preset number of generated molecular structures has a certain effectiveness, that is, the molecular structure is effective under chemical rules and further supports the generation of lead compounds.
[0142] As described above, the method for generating structural data according to the embodiment of the present application samples from candidate data of a designated data distribution to obtain a certain number of sampling data, inputs the sampling data into a designated decoder to perform structure reconstruction, thereby obtaining reconstructed structural data with a certain effectiveness and improving the diversity of the generated reconstructed structural data.
[0143] Exemplarily, the method for generating structural data according to the embodiment of the present application is applied to the publicly available dataset ZINC for testing. A designated decoder is obtained through training, and a new chemical molecule is generated by the designated decoder. Here, in the process of generating the new chemical molecule, 104 random samplings are performed on the data N(0, I) of the normal distribution, and the sampling results are input into the designated decoder obtained by training, and 104 newly generated chemical molecules are obtained. The obtained new chemical molecules are verified by the open-source platform rdkit. As a result, the new chemical molecules have high uniqueness and novelty on the premise that the effectiveness is guaranteed. Among them, the effectiveness is 98.0%, the uniqueness is 99.1%, and the novelty is 96.5%. The higher the uniqueness and novelty, the higher the diversity of molecular generation. As can be seen as a whole, the diversity of the generated molecules is improved, so the generation space can be enlarged.
[0144] It should be noted that when the above embodiments of the present application are applied to specific products or technologies, and when it involves user data (for example, the method is applied to a recommendation system), it is necessary to obtain the permission or consent of the user for the acquisition of such data. At the same time, regarding the research of compounds and the use and processing of data, it is necessary to comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0145] FIG. 10 is a block diagram showing the structure of a structure data generation device according to an exemplary embodiment of the present application. The device includes an acquisition module 1010 for acquiring the structural feature representation and node feature representation of the sample structure data, wherein the structural feature representation is for indicating the connection status between nodes constituting the sample structure data, and the node feature representation is for indicating the node type corresponding to the nodes constituting the sample structure data, an encoding module 1020 for generating a hidden layer feature representation based on the structural feature representation and the node feature representation, wherein the hidden layer feature representation is for indicating the coupling status between nodes in the sample structure data in at least two frequency bands, a decoding module 1030 for obtaining predicted structure data by inputting the hidden layer feature representation into a decoder to be trained to perform structure reconstruction, and a training module 1040 for obtaining a specified decoder by training the decoder to be trained based on the predicted structure data, wherein the specified decoder is for performing structure reconstruction on the input sampling data to obtain reconstructed structure data, and the sampling data is data obtained by sampling candidate data.
[0146] In some alternative embodiments, as shown in FIG. 11, the encoding module 1020 It is for obtaining intermediate feature data by encoding in each of the at least two frequency bands based on the structural feature representation and the node feature representation, and the intermediate feature data is for indicating the connection situation between nodes of the sample structure data in the corresponding frequency band. An encoding unit 1021 Based on the specified data distribution, it is for clustering the intermediate feature data corresponding to the at least two frequency bands respectively to obtain the hidden layer feature representation, and the candidate data is data that satisfies the specified data distribution. A clustering unit 1022 is further included.
[0147] In some selectable embodiments, the encoding unit 1021 is further used to input the structural feature representation and the node feature representation into a to-be-trained encoder. When the structural feature representation and the node feature representation are regarded as known conditions, the to-be-trained encoder determines the connection probability that the i-th prediction node in the prediction node set establishes a connection relationship with each prediction node in the prediction node set, and determines a connection probability distribution corresponding to the i-th prediction node based on the connection probability between the i-th prediction node and each prediction node, determines the intermediate feature data corresponding to each of the at least two frequency bands based on the fusion result of the connection probability distributions of all prediction nodes in the prediction node set. The prediction node is for constructing the prediction structure data, and i is a positive integer.
[0148] In some selectable embodiments, the encoding unit 1021 further obtains the node feature vectors of the nodes of the sample structure data in the feature spaces corresponding to the at least two frequency bands respectively based on the structural feature representation and the node feature representation, obtains the average value data and the variance data between the node feature vectors corresponding to the at least two frequency bands, and is also used to determine the average value data and the variance data as the intermediate feature data.
[0149] In some selectable embodiments, the decoding module 1030 a reconstruction unit 1031 for obtaining a decoded structure feature representation and a decoded node feature representation by reconstructing the structure of the hidden layer feature representation by the decoder waiting for training; and a generation unit 1032 for generating the predicted structure data based on the decoded structure feature representation and the decoded node feature representation.
[0150] In some selectable embodiments, the apparatus a training loss value is obtained based on the structural difference situation between the sample structure data and the predicted structure data, and in response to the training loss value reaching a specified loss threshold, it is determined that the training of the decoder waiting for training is completed, and the specified decoder is obtained, or, in response to the failure of the matching between the training loss value and the specified loss threshold, a training module 1040 for repeatedly training the model parameters of the decoder waiting for training is further included.
[0151] In some selectable embodiments, the training module 1040 an acquisition unit 1041 for obtaining distance metric data of the sample structure data and the predicted structure data in the feature space; and a determination unit 1042 for obtaining the training loss value based on the distance metric data and the divergence data.
[0152] The acquisition unit 1041 is also used to further obtain divergence data between the node distribution corresponding to the predicted structure data and the specified data distribution, and the node distribution is for indicating the distribution situation of the node feature vectors of the predicted structure data in the feature space.
[0153] In some selectable embodiments, the acquisition module 1010 is also used to further obtain a candidate structure feature representation and a candidate node feature representation of candidate structure data. The encoding module 1020 is further used to generate a candidate hidden layer feature representation based on the candidate structure feature representation and the candidate node feature representation. The decoding module 1030 is further used to obtain the reconstructed structure data by inputting the candidate hidden layer feature representation into the specified decoder for prediction. There is a structural property similarity relationship between the reconstructed structure data and the candidate structure data.
[0154] In some alternative embodiments, the apparatus further includes a sampling module 1050 for obtaining candidate data of the specified data distribution and sampling from the candidate data to obtain a preset number of sampling data.
[0155] The decoding module 1030 is further used to input the preset number of sampling data into the specified decoder to obtain the preset number of reconstructed structure data.
[0156] In some alternative embodiments, the specified decoder is used to generate a candidate molecular structure composed of at least two atomic nodes. The decoding module 1030 is further used to input the preset number of sampling data into the specified decoder and obtain the preset number of candidate molecular structures based on the connection relationship between atomic nodes in the molecular structure learned during training by the specified decoder.
[0157] In some selectable embodiments, when the trained specified decoder is used for generating a molecular structure, the acquisition model 1010 further acquires a sample chemical molecule which is a known molecule satisfying an atomic bond criterion and composed of at least two atoms, converts the sample chemical molecule into a sample molecular graph whose data structure is a graph structure, the nodes of the sample molecular graph are for representing the at least two atoms in the sample chemical molecule, the edges in the sample molecular graph are for representing chemical bonds between atoms in the sample chemical molecule, and is also used to determine an adjacency matrix corresponding to the sample molecular graph as the structural feature representation and determine a node matrix corresponding to the sample molecular graph as the node feature representation.
[0158] As described above, after obtaining a hidden layer feature representation by a structural feature representation and a node feature representation corresponding to sample structure data, the structure data generation device according to the embodiment of the present application performs iterative training on a decoder to be trained based on the hidden layer feature representation to obtain a specified decoder. Thus, the specified decoder can generate structure data based on the input sampling data. That is, if necessary, various structure data can be rapidly generated by the trained specified decoder, and the generation efficiency and diversity of the structure data can be improved.
[0159] It should be noted that the structure data generation device according to the above embodiment is described by taking the division of each of the above function modules as an example. However, in actual applications, if necessary, the above functions can be assigned to different function modules to be completed. That is, the internal structure of the device can be divided into different function modules to complete all or part of the above functions. In addition, the structure data generation device according to the above embodiment belongs to the same concept as the embodiment of the structure data generation method, and its specific implementation process can refer to the embodiment of the method, so the description is omitted here.
[0160] FIG. 12 is a diagram showing the structure of a server according to one exemplary embodiment of the present application. Specifically, it includes the following structure.
[0161] Server 1200 includes a CPU (Central Processing Unit) 1201, a system memory 1204 including a RAM (Random Access Memory) 1202 and a ROM (Read Only Memory) 1203, and a system bus 1205 connecting the system memory 1204 and the CPU 1201. Server 1200 further includes a mass storage device 1206 for storing an operating system 1213, an application program 1214, and other program modules 1215.
[0162] The mass storage device 1206 is connected to the CPU 1201 via a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1206 and the associated computer-readable medium provide non-volatile memory for the server 1200.
[0163] Without loss of generality, the computer-readable medium may include computer storage media and communication media. The above system memory 1204 and mass storage device 1206 may be collectively referred to as memory.
[0164] According to various embodiments of the present application, the server 1200 may be connected to a network 1212 by a network interface unit 1211 connected to the system bus 1205, or may be connected to another type of network or a remote computer system (not shown) by using the network interface unit 1211.
[0165] The above memory further includes one or more programs, and the one or more programs are stored in the memory and arranged to be executed by the CPU.
[0166] Embodiments of the present application include a processor and a memory, and the memory stores at least one instruction, at least one program, a code set, or an instruction set that can be loaded and executed by the processor to implement the method for generating structure data according to each of the above method embodiments. A computer device is further provided. Preferably, the computer device may be a terminal or a server.
[0167] Embodiments of the present application further provide a computer-readable storage medium that stores at least one instruction, at least one program, a code set, or an instruction set that can be loaded and executed by a processor to implement the method for generating structure data according to each of the above method embodiments.
[0168] Embodiments of the present application further provide a computer program product or a computer program that includes computer instructions stored in a computer-readable storage medium. By reading and executing the computer instructions from the computer-readable storage medium by a processor of a computer device, the computer device is caused to execute the method for generating structure data according to any of the above embodiments.
[0169] Preferably, the computer-readable recording medium may include a ROM (Read Only Memory), a RAM (Random Access Memory), an SSD (Solid State Drives), an optical disk, etc. Here, the RAM may include a ReRAM (Resistance Random Access Memory) and a DRAM (Dynamic Random Access Memory). The serial numbers of the embodiments of the present application above are only for the purpose of description and do not represent the superiority or inferiority of the embodiments.
Claims
1. A method for generating structural data executed by a computer device, comprising: obtaining a structural feature representation and a node feature representation of sample structural data, wherein the structural feature representation is for indicating the connection status between nodes constituting the sample structural data, and the node feature representation is for indicating the node type corresponding to the nodes constituting the sample structural data; generating a hidden layer feature representation based on the structural feature representation and the node feature representation, wherein the hidden layer feature representation is for indicating the coupling status between nodes in the sample structural data in at least two frequency bands; obtaining predicted structural data by inputting the hidden layer feature representation into a decoder to be trained for reconstructing the structure; obtaining a specified decoder by training the decoder to be trained based on the predicted structural data, wherein the specified decoder is for reconstructing the structure of the input sampling data to obtain reconstructed structural data, and the sampling data is data obtained by sampling candidate data; A method for generating structural data.
2. The step of generating a hidden layer feature representation based on the structural feature representation and the node feature representation comprises: obtaining intermediate feature data by encoding the structural feature representation and the node feature representation respectively in the at least two frequency bands, wherein the intermediate feature data is for indicating the coupling status between nodes of the sample structural data in the corresponding frequency band; obtaining the hidden layer feature representation by clustering the intermediate feature data respectively corresponding to the at least two frequency bands based on a specified data distribution, wherein the candidate data is data satisfying the specified data distribution; The method according to claim 1.
3. The step of obtaining intermediate feature data by encoding the structural feature representation and the node feature representation respectively in the at least two frequency bands comprises: Input the structural feature representation and the node feature representation into the encoder to be trained. When the encoder to be trained takes the structural feature representation and the node feature representation as known conditions, it determines the connection probability for the i-th predicted node in the predicted node set to establish a connection relationship with each predicted node in the predicted node set. At the same time, based on the connection probability between the i-th predicted node and each predicted node, it determines the connection probability distribution corresponding to the i-th predicted node, and determines the intermediate feature data corresponding to the at least two frequency bands respectively based on the fusion result of the connection probability distributions of all the predicted nodes in the predicted node set. The predicted node is for constructing the predicted structure data, and i is a positive integer, including the step of The method according to claim 2.
4. The step of obtaining intermediate feature data by encoding respectively in the at least two frequency bands based on the structural feature representation and the node feature representation is Based on the structural feature representation and the node feature representation, the step of obtaining the node feature vectors in the corresponding feature spaces of the nodes of the sample structure data in the at least two frequency bands respectively The step of obtaining the mean value data and the variance data between the node feature vectors corresponding to the at least two frequency bands The step of determining the mean value data and the variance data as the intermediate feature data, including The method according to claim 2.
5. The step of obtaining predicted structure data by inputting the hidden layer feature representation into the decoder to be trained for structure reconstruction is The step of obtaining the decoded structure feature representation and the decoded node feature representation by reconstructing the structure for the hidden layer feature representation by the decoder to be trained The step of generating the predicted structure data based on the decoded structure feature representation and the decoded node feature representation, including The method according to any one of claims 1 to 4.
6. The step of obtaining a specified decoder by training the decoder to be trained based on the predicted structure data is The step of obtaining a training loss value based on the structural difference situation between the sample structure data and the predicted structure data In response to the training loss value reaching a specified loss threshold, it is determined that the training of the decoder waiting for training is completed, and the specified decoder is obtained, or in response to the failure of the matching between the training loss value and the specified loss threshold, iteratively training the model parameters of the decoder waiting for training, and The method according to any one of claims 1 to 4.
7. The step of obtaining a training loss value based on the structural difference situation between the sample structure data and the predicted structure data includes obtaining distance metric data of the sample structure data and the predicted structure data in the feature space, and obtaining divergence data between the node distribution corresponding to the predicted structure data and the specified data distribution, where the node distribution is for indicating the distribution situation of the node feature vectors of the predicted structure data in the feature space, and the divergence data is for indicating the difference degree between the node distribution and the specified data distribution, and obtaining the training loss value based on the distance metric data and the divergence data. The method according to claim 6.
8. obtaining a candidate structure feature representation and a candidate node feature representation of candidate structure data, and generating a candidate hidden layer feature representation based on the candidate structure feature representation and the candidate node feature representation, and obtaining the reconstructed structure data by inputting the candidate hidden layer feature representation into the specified decoder for prediction, where there is a structural property similarity relationship between the reconstructed structure data and the candidate structure data. The method according to any one of claims 1 to 4.
9. obtaining candidate data of the specified data distribution, and sampling from the candidate data to obtain a preset number of sampling data, and inputting the preset number of sampling data into the specified decoder to obtain the preset number of the reconstructed structure data. The method according to any one of claims 1 to 4.
10. The specified decoder is used to generate a candidate molecular structure composed of at least two atomic nodes The step of inputting the preset number of sampling data into the designated decoder to obtain the preset number of reconstructed structure data is: including the step of inputting the preset number of sampling data into the designated decoder, and obtaining the preset number of candidate molecular structures based on the connection relationship between atomic nodes in the molecular structure learned by the designated decoder during training. The method according to claim 9.
11. Before the step of obtaining the structural feature representation and node feature representation of the sample structure data when the designated decoder obtained by training is used for generating a molecular structure, a step of obtaining a sample chemical molecule, where the sample chemical molecule is a known molecule that satisfies the atomic bonding criterion and is composed of at least two atoms, a step of converting the sample chemical molecule into a sample molecular graph whose data structure is a graph structure, where the nodes of the sample molecular graph are for representing the at least two atoms in the sample chemical molecule, and the edges in the sample molecular graph are for representing the chemical bonds between atoms in the sample chemical molecule, a step of determining the adjacency matrix corresponding to the sample molecular graph as the structural feature representation, and a step of determining the node matrix corresponding to the sample molecular graph as the node feature representation. The method according to any one of claims 1 to 4.
12. An acquisition module for acquiring the structural feature representation and node feature representation of sample structure data, where the structural feature representation is for indicating the connection status between nodes constituting the sample structure data, and the node feature representation is for indicating the node type corresponding to the nodes constituting the sample structure data. An encoding module for generating a hidden layer feature representation based on the structural feature representation and the node feature representation, where the hidden layer feature representation is for indicating the bonding status between nodes in the sample structure data in at least two frequency bands. A decoding module for obtaining predicted structure data by inputting the hidden layer feature representation into a decoder waiting for training to perform structure reconstruction. A training module for obtaining a designated decoder by training the decoder awaiting training based on the predicted structure data, wherein the designated decoder is for reconstructing a structure for input sampling data to obtain reconstructed structure data, and the sampling data is data obtained by sampling candidate data, and the training module. A structure data generation device. **Claim 13** Including a processor and a memory. The memory stores at least one instruction, at least one program, code set or instruction set that realizes the structure data generation method according to any one of claims 1 to 4 when loaded and executed by the processor. A computer device. **Claim 14** A computer program that realizes the structure data generation method according to any one of claims 1 to 4 when executed by a processor.
Citation Information
Patent Citations
Substance structure analysis device, method and program
JP2020139914A
Chemical compound structure automatic creation device for automatically creating chemical compound structure, chemical compound structure automatic creation system and chemical compound structure automatic creation method
JP2021081769A
Target-to-catalyst translation networks
WO2021081390A1