Molecular classification method based on heterogeneous graph embedding
By combining the heterogeneous graph embedding method with simplex structure and hash algorithm, the problem of difficult to capture and computational overhead in heterogeneous graph embedding is solved, and efficient molecular classification is achieved, especially efficient modeling under large-scale molecular graph data.
Patent Information
- Application Number
- CN202510430570.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing heterogeneous graph embedding methods are difficult to effectively capture high-order interactive information. Heterogeneous information modeling has limitations, high computing overhead, and it is difficult to adapt to the efficient modeling needs of large-scale molecular graphs.
Using a method of combining simplex structure modeling and hashing algorithms, a simplex structure is extracted and the Hochlaplace matrix is calculated, local and global information enhancement operators are constructed, and iterative updates are performed in combination with local sensitive hashing algorithms to generate graph-level embeddings.
Effectively capture high-order interactive information and heterogeneous features, improve computing efficiency, adapt to the efficient modeling needs of large-scale molecular maps, significantly reduce computing resource consumption, and improve classification accuracy.
Smart Images

Figure CN120355989A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of molecular classification, and relates to a molecular classification method based on heterogeneous graph embedding. Background Art
[0002] Molecular classification is an important research direction in the field of bioinformatics and is widely applied in multiple fields such as drug design, environmental science, and materials science. By accurately classifying the activity or function of compounds, it not only helps to accelerate the discovery and screening process of new drugs but also has important value in practical applications such as disease diagnosis and toxicity prediction. In practical applications, molecules usually exist in the form of a graph structure, where nodes represent different types of atoms and edges represent chemical bonds between atoms. To achieve molecular classification, it is necessary to encode the molecular graph into a fixed-length feature vector while retaining its key structural information and attribute information. Since the molecular graph contains multiple types of nodes and edges, it can be modeled as a heterogeneous graph and transformed into a low-dimensional representation by means of a heterogeneous graph embedding method to support downstream classification tasks.
[0003] However, the existing heterogeneous graph embedding methods still face the following challenges:
[0004] (1) Difficulty in capturing high-order interaction information: Most methods only consider the direct connection relationship between nodes and edges and are difficult to model the high-order topological structures (such as triangles, tetrahedrons, etc.) in the graph, resulting in insufficient expression ability for complex structural information.
[0005] (2) Limitations in heterogeneous information modeling: The types of nodes and edges in a heterogeneous graph are rich, and these heterogeneous features are crucial for understanding the molecular structure. However, existing methods often rely on manually defined meta-paths or meta-graph structures, which limit the flexibility and generalization ability of the model.
[0006] (3) High computational cost: Although the recently emerging heterogeneous graph embedding methods based on graph neural networks (such as HGNNs) have certain expression ability, they rely on a large number of parameters for training and need to optimize the weights through backpropagation, consuming high computational and storage resources and being difficult to meet the efficient modeling requirements of large-scale molecular graphs. Summary of the Invention
[0007] To solve the above problems, the present invention aims to provide a molecular classification method based on heterogeneous graph embedding that combines the modeling ability of the simplex structure and the efficiency of the hash algorithm to effectively capture the high-order interaction information and heterogeneous features in the molecular structure, while taking into account the embedding efficiency and model scalability, providing an efficient and practical solution for molecular classification.
[0008] The present invention provides a molecular classification method based on heterogeneous graph embedding, including the following steps:
[0009] Step 1: Extract simplices from the heterogeneous graph g;
[0010] Process the structural information of the heterogeneous graph g based on the extracted simplices to obtain the local and global information enhancement operator N k ;
[0011] Process the feature information of the heterogeneous graph g based on the extracted simplices to obtain the feature matrix
[0012] Step 2: Perform hash iteration on the heterogeneous graph g based on the local and global information enhancement operator N k and the feature matrix to obtain the graph-level embedding x g ;
[0013] Step 3: Classify the graph-level embedding x g to obtain the molecular classification result.
[0014] Furthermore, the heterogeneous graph g is from the molecular heterogeneous graph dataset in the known TUDataset, and the molecular heterogeneous graph dataset contains various types of heterogeneous graphs g.
[0015] Furthermore, the specific process of extracting simplices from the heterogeneous graph g is as follows:
[0016] Take all the nodes in the node set V of the heterogeneous graph g as the 0-dimensional simplices and add them to the 0-dimensional simplex set S0, where S0 contains all 0-simplices;
[0017] Take all the edges in the edge set E of the heterogeneous graph g as the 1-dimensional simplices and add them to the 1-dimensional simplex set S1, where S1 contains all 1-simplices;
[0018] Initialize the Simplex tree with the simplex sets S0 and S1, and use the expansion method in Simplextree to recursively expand from the 2-dimensional simplices, layer by layer to obtain higher-dimensional simplices until reaching the maximum dimension K, so as to extract the 2-dimensional simplices to the K-dimensional simplices respectively, and obtain S2, …, S K , S2, …, S K are the sets from the 2-dimensional to the maximum dimension K respectively.
[0019] Furthermore, the specific process of obtaining the local and global information enhancement operator N k is as follows:
[0020] Calculate the Hodge Laplacian matrix L k corresponding to the k-dimensional simplex set S k , k ∈ {0, 1, …, K};
[0021] Based on the Hodge Laplacian matrix L k , construct the local information enhancement operator M k ;
[0022] Based on the local information enhancement operator M k , further construct the local and global information enhancement operator N k .
[0023] Furthermore, the specific process of obtaining the feature matrix is as follows:
[0024] Create a zero matrix of size |S k | × |T v |
[0025] Traverse each k-simplex in S k and construct the heterogeneous information vectors of all the nodes it contains;
[0026] Traverse each k-simplex in S k and perform an element-wise logical OR operation on the heterogeneous information vectors of all the nodes it contains to obtain a binary vector of length |T v |, which serves as the feature vector;
[0027] Fill the feature vector of each k-simplex into the zero matrix to obtain the feature matrix from dimension 0 to the maximum dimension K
[0028] Furthermore, the specific process of obtaining the graph-level embedding x g is as follows:
[0029] (ⅰ) Initialization;
[0030] Let r = 1 and k = 0. Where: r represents the r-th iteration, and k represents the k-th dimensional simplex;
[0031] (ⅱ) Iterative loop condition judgment;
[0032] Condition: If r ≤ R, then execute step (ⅲ); otherwise, jump to step (ⅷ);
[0033] (ⅲ) Simplex dimension loop condition judgment;
[0034] Condition: If k ≤ K, then execute step (ⅳ); otherwise, jump to step (ⅶ);
[0035] (ⅳ) Temporary feature matrix calculation;
[0036] Calculate the temporary feature matrix of the k-simplex in the r-th iteration
[0037]
[0038] where: is the feature matrix of the k-simplex after the (r - 1)-th iteration; W (r-1,k) is the random hash function matrix used to update .
[0039] (v), Sign operation;
[0040] Apply the sign operation sgn to the temporary feature matrix to convert each element of into a binary hash code, obtaining the updated feature matrix
[0041] The updated feature matrix is expressed as follows:
[0042]
[0043] (vi), Update the simplex dimension to k′;
[0044] Increment k by 1, i.e., k′ = k + 1.
[0045] Jump back to step (iii);
[0046] (vii), Update the iteration number to r′;
[0047] Increment r by 1, i.e., r′ = r + 1;
[0048] Reset k = 0;
[0049] Jump back to step (ii);
[0050] (viii), Generate the graph-level embedding x g ;
[0051] Flatten the final simplex feature matrix into vectors respectively, and then concatenate them to obtain the graph-level embedding x g .
[0052] Furthermore, the specific process of molecular heterogeneous graph classification for the graph-level embedding x g is as follows:
[0053] S3.1, Perform a bitwise negation operation on each binary bit of x g , that is, convert 0 to 1 and 1 to 0, obtaining the negated vector flip(x g );
[0054] S3.2. Concatenate the original vector x g and the vector flip(x g ) after bitwise inversion into a 2L-dimensional extended vector x e ;
[0055] S3.3. Input x e into a logistic regression classifier for classification;
[0056] Specifically: Let x g = [1, 0, 1], then flip(x g ) = [0, 1, 0]. After concatenation, x e = [1, 0, 1, 0, 1, 0], and finally input it into a logistic regression classifier for classification.
[0057] Where: flip(x g ) is the vector after bitwise inversion of x g ; x e is the extended 2L-dimensional vector, obtained through .
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] (1) The present invention proposes a molecular classification method based on heterogeneous graph embedding. This method first preprocesses the heterogeneous graph data. By extracting simplices and calculating the corresponding Hodge Laplacian matrix, and then performing element-wise product and second-order operations on the Hodge Laplacian matrix, a local and global graph structure information enhancement operator is constructed. At the same time, the simplex features are initialized based on simplices and heterogeneous information. Next, through the local and global information enhancement operator, using the random hashing algorithm, the simplex features are iteratively updated. Finally, the simplex features of different dimensions obtained from the last iteration are respectively vectorized, and through the feature concatenation operation, the final heterogeneous graph-level embedding is generated, and the embedding can be used for downstream tasks such as graph classification. This method effectively integrates the high-order topological information of the heterogeneous graph structure and the efficient characteristics of the hashing algorithm, providing an effective framework for heterogeneous graph analysis.
[0060] (2) By combining simplices with the hashing algorithm, the present invention can not only effectively capture high-order interaction information and heterogeneous features, but also significantly improve the computational efficiency, providing an efficient and scalable solution for heterogeneous graph analysis.
[0061] (3) When performing graph embedding on the molecular heterogeneous graph, this method is based on the message passing mechanism, combines local and global information enhancement operators, and introduces a random hashing algorithm for optimization. Traditional Heterogeneous Graph Neural Networks (HGNNs) rely on a large number of trainable parameters, update weights through backpropagation, have complex calculations and are prone to overfitting, especially with low efficiency when dealing with large-scale molecular graph data. The present invention uses a locality-sensitive hashing function to replace the non-linear activation function in the message passing process of the traditional HGNNs method, and uses a randomly generated hash mapping to replace the parameter learning process, effectively reducing the consumption of computing resources and improving the embedding efficiency of the model in the large-scale molecular graph scenario.
[0062] (4) Verified by experiments, this method shows high accuracy and computational efficiency in the molecular graph classification task based on heterogeneous graph embedding.
[0063] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the figures for a further detailed description of the present invention. Description of the Drawings
[0064] The drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0065] Figure 1 It is a schematic flowchart of a molecular classification method based on heterogeneous graph embedding in an embodiment of the present invention. Detailed Embodiments
[0066] To make the above purposes, features, and advantages of the present invention more clearly understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the drawings. It should be noted that the drawings of the present invention are all in simplified forms and use non-precise scales, only for conveniently and clearly assisting in explaining the implementation of the present invention; the several mentioned in the present invention are not limited to the specific quantities in the drawing examples; the orientation or positional relationships indicated by 'front', 'middle', 'back', 'left', 'right', 'up', 'down', 'top', 'bottom', 'middle', etc. in the present invention are all based on the orientation or positional relationships shown in the drawings of the present invention, and do not indicate or imply that the devices or components referred to must have a specific orientation, nor can they be understood as a limitation to the present invention.
[0067] Embodiment:
[0068] See Figure 1 As shown, a molecular classification method based on heterogeneous graph embedding proposed by the present invention is used to efficiently complete the molecular graph embedding and classification tasks; it includes the following steps:
[0069] Step 1: Data preprocessing.
[0070] Existing heterogeneous graph embedding methods mainly model based on the direct binary relationships between nodes and edges, or rely on predefined substructures such as meta-paths and meta-graphs, making it difficult to effectively capture the high-order interaction features in the graph, resulting in limited model performance. Therefore, the present invention introduces simplices to model the high-order interaction relationships in heterogeneous graphs by capturing high-order topological structures (such as triangles and tetrahedrons). Heterogeneous graphs contain rich heterogeneous semantic information, which is crucial for graph embedding. Therefore, in the data preprocessing stage, the present invention extracts simplices to capture high-order interactions and incorporates heterogeneous information into the simplices, thereby better preserving the structural and semantic features of the graph.
[0071] The data used in this application comes from the molecular heterogeneous graph dataset in TUDataset. TUDataset is a collection of publicly available graph datasets, containing various types of heterogeneous graph datasets covering multiple fields such as chemical molecules, social networks, and biological networks. In the molecular heterogeneous graph dataset, each heterogeneous graph corresponds to a specific molecular compound.
[0072] Specifically, the heterogeneous graph g can be formally represented as:
[0073] g = (V, E);
[0074] where V is the set of nodes of the heterogeneous graph g, corresponding to the atoms in the molecular compound, |V| is the total number of nodes in the node set V; E is the set of edges of the heterogeneous graph g, corresponding to the chemical bond connection relationships between atoms, |E| is the total number of edges in the edge set E; T v is the set of node types, |T v | is the number of types in the set of node types, v is any node in the heterogeneous graph g, v ∈ V; t v is the type label of any node v, t v ∈ T v and t v exists in integer form.
[0075] Specifically, the specific process of preprocessing the heterogeneous graph g is as follows:
[0076] S1.1: Simplex extraction;
[0077] S1.2: Processing the structural information and feature information based on the extracted simplices.
[0078] Furthermore, a simplex is a fundamental concept in topological data analysis, used to describe geometric structures in high-dimensional spaces. In the present invention, the Simplex tree data structure (Simplex tree is an existing data structure that represents simplices and their inclusion relationships through a tree structure, supporting fast insertion, query, and traversal of simplices) is used to extract a set of simplices {S0, S1, …, S k , … S K} from the input heterogeneous graph g = (V, E) in the following way:
[0079] Input: The node set V, edge set E of the heterogeneous graph g, and the maximum dimension K of the simplices.
[0080] Output: The set of simplices {S0, S1, …, S k , … S K};
[0081] Wherein: S0 is the set of simplices of the 0th dimension, containing all 0-simplices; S1 is the set of simplices of the 1st dimension, containing all 1-simplices; S k is the set of simplices of the kth dimension, containing all k-simplices, k ∈ {0, 1, …, K}; S K is the set of simplices of the maximum degree K, containing all K-simplices.
[0082] Furthermore, the specific process of simplex extraction is as follows:
[0083] All nodes in the node set V are used as 0th dimension simplices and added to the S0 set;
[0084] All edges in the edge set E are used as 1st dimension simplices and added to the S1 set;
[0085] The S0 set and the S1 set are used to initialize the Simplex tree, and the expansion method in the Simplex tree is used to recursively expand from the 2nd dimension simplices, layer by layer to obtain higher dimension simplices until the maximum dimension K is reached, so as to extract the 2nd dimension simplices to the Kth dimension simplices respectively (i.e., obtain S2, …, S K ).
[0086] In the present invention, through the above specific process, the extraction of the set of simplices {S0, S1, …, S k , … S K} from the 0th dimension to the maximum dimension K is realized, so as to extract simplices of different dimensions from the heterogeneous graph g.
[0087] Furthermore, for the set of simplices S kTaking as an example, the specific process of structural information processing (i.e., constructing local and global information enhancement operators) on the heterogeneous graph g is as follows:
[0088] ①. Calculate the Hodge Laplace matrix L corresponding to the simplex set of the kth dimension k ;
[0089] Hodge Laplace matrix L k The expression is as follows:
[0090]
[0091] in: is the boundary operator of the simplex of dimension k, for The transpose of is the boundary operator of the simplex of dimension k+1, for The transpose of .
[0092] In actual calculations, Usually through the matrix B k Indicates. Definition | S k | is a simplex set S k The number of k-simplices in |S k-1 | is a simplex set S k-1 The number of (k-1)-simplices in the definition is a simplex set S k The jth k-simplex in is a simplex set S k-1 The i-th (k-1)-simplex in k The dimension is |S k-1 |×|S k |, matrix B k The expression is as follows:
[0093]
[0094] Among them: B k (y,z) is the matrix B k The value of the yth row and zth column in .
[0095] Special Note: B k (y,z) Here the characters y and z correspond strictly to and The subscript index of .
[0096] ② In order to mine more potential information, the present invention performs the Hodge Laplace matrix L kFurther transformation processing was carried out, and a method for constructing an operator that enhances local and global information was proposed, which is as follows:
[0097] Based on the Hodge Laplacian matrix L k , construct the local information enhancement operator M k ;
[0098] The expression of the local information enhancement operator M k is as follows:
[0099] M k = L k ⊙ L k ;
[0100] Where: ⊙ is the element-wise product.
[0101] In the k-simplex and its Hodge Laplacian matrix L k , the local information is defined by the self-information of each k-simplex and its adjacency relationship. The element-wise product ⊙ non-linearly amplifies each element of L k , enhancing both the self-information of a single k-simplex (represented by the diagonal elements of L k ) and the adjacency relationship between k-simplices (represented by the non-diagonal elements of L k ). This operation enhances the distinguishability of local information by expanding existing differences to varying degrees. This improvement provides richer local structure information for subsequent graph embedding learning.
[0102] Based on the local information enhancement operator M k , further construct the local and global information enhancement operator N k ;
[0103] The expression of the local and global information enhancement operator N k is as follows:
[0104] N k = (M k ) 2 .
[0105] The global information of the heterogeneous graph involves the high-order connection information between k-simplices. The Hodge Laplacian matrix L k cannot directly retain this information because its corresponding non-diagonal elements are zero. Even after applying the element-wise product to obtain the local information enhancement operator M k , this limitation still exists. By performing a second-order operation on the local information enhancement operator M k , M kThe non - diagonal elements that are zero become non - zero, indicating that some high - order connection relationships are extracted. The second - order operation further increases the relative differences between elements by more significantly amplifying strong connections and important information, thereby enhancing the distinguishability of global information.
[0106] Furthermore, in the molecular heterogeneous graph, the heterogeneous features of nodes and edges contain rich chemical properties and structural semantic information, which is of great significance for accurately understanding the molecular structure and achieving effective graph embedding. Therefore, after the simplex extraction, the present invention initializes the simplex features of each dimension based on the heterogeneous type information to fully retain the semantic structure features of the molecule and enhance the embedding expression ability.
[0107] Specifically, taking the simplex set S of the k - th dimension k as an example, the method for processing the feature information (i.e., simplex feature initialization) of the heterogeneous graph g is as follows:
[0108] Input: Simplex sets {S0, S1, …, S k , … S K}, and the node set V of the heterogeneous graph g;
[0109] Output: Feature matrices of simplices from the 0 - th dimension to the K - th dimension are obtained respectively
[0110] Even further, define as the feature matrix of the k - simplex after the r - th iteration. Each row in the matrix represents the feature vector of a k - simplex, and the dimension of the matrix is |S k | × |T v |; Taking the initialization of the feature matrix of the k - simplex as an example, the specific process of processing the feature information (i.e., simplex feature initialization) of the heterogeneous graph g is as follows:
[0111] (Ⅰ) Create a zero matrix of size |S k | × |T v |
[0112] (Ⅱ) Traverse each k - simplex in S k and construct the heterogeneous information vector of all the nodes it contains. Specifically: Define as the p - th k - simplex in S k . For each node v in , convert the type label t v of the node v into a one - hot encoded vector, that is, a binary vector of length |T v |, where: only tv The corresponding index position is 1, and the rest are 0.
[0113] (Ⅲ), Traverse S k For each k-simplex in it, construct the feature vector of each k-simplex. Specifically: for the k p-th k-simplex in S as an example, perform an element-wise logical OR operation on the heterogeneous information vectors of all the nodes it contains to obtain a binary vector of length |T v |, which serves as the feature vector.
[0114] (Ⅳ), Fill the zero matrix Specifically: fill the feature vector of each k-simplex into the zero matrix .
[0115] In this application, after initializing the features of the simplices in the 0th dimension to the Kth dimension respectively by adopting the above steps, the feature matrices
[0116] Step 2, Hash iteration.
[0117] Existing heterogeneous graph embedding methods based on heterogeneous graph neural networks achieve feature learning by optimizing neural network parameters. However, these methods require frequent execution of forward propagation and backward propagation calculations during training, resulting in a sharp increase in computational complexity and an increase in training time, restricting the practical application efficiency of the methods. To improve this problem, this method constructs local and global information enhancement operator N k , and after initializing the simplex features to obtain the feature matrix , iteratively updates the simplex features through locality-sensitive hashing (LSH).
[0118] The process of performing hash iteration on the heterogeneous graph g is as follows:
[0119] Input: Feature matrix Local and global information enhancement operator N k and a set of randomly generated random hash functions W for updating (r,k) ; among them, the expression of the randomly generated random hash function W (r,k) is as follows:
[0120] {W (r,k) |r ∈ {0, 1, …, R - 1}, k ∈ {0, 1, …, K}};
[0121] where: R is the number of iterations.
[0122] Output: Graph-level embedding x g .
[0123] Specifically, for the r-th iteration of the k-simplex, a random hash function matrix W (r,k) is generated, with a shape of |S k |×25. Each element of the matrix W( r,k) is randomly sampled from the standard normal distribution, that is, randomly generated. Thus, a set of random hash functions is obtained.
[0124] Preferably, the specific steps for hashing iteration on the heterogeneous graph g are as follows:
[0125] (ⅰ) Initialization;
[0126] Let r = 1 and k = 0. Where: r represents the r-th iteration, and k represents the k-th dimensional simplex.
[0127] (ⅱ) Iterative loop condition judgment;
[0128] Condition: If r ≤ R, then execute step (ⅲ); otherwise, jump to step (ⅷ).
[0129] (ⅲ) Simplex dimension loop condition judgment;
[0130] Condition: If k ≤ K, then execute step (ⅳ); otherwise, jump to step (ⅶ).
[0131] (ⅳ) Temporary feature matrix calculation;
[0132] Calculate the temporary feature matrix of the k-simplex in the r-th iteration
[0133]
[0134] Where: is the feature matrix of the k-simplex after the (r - 1)-th iteration; W (r-1,k) is the random hash function matrix used to update .
[0135] (ⅴ) Sign operation;
[0136] Apply the sign operation sgn (specifically, for the input elements, if the element value is greater than 0, then return binary 1, if the element value is less than or equal to 0, then return binary 0) to the temporary feature matrix , and convert each element of into a binary hash code to obtain the updated feature matrix
[0137] Updated feature matrix The expression is as follows:
[0138]
[0139] (ⅵ) Update the simplex dimension to k′;
[0140] Increment k by 1, i.e., k′ = k + 1.
[0141] Jump back to step (ⅲ).
[0142] (ⅶ) Update the iteration count to r′;
[0143] Increment r by 1, i.e., r′ = r + 1.
[0144] Reset k = 0.
[0145] Jump back to step (ⅱ).
[0146] (ⅷ) Generate the graph-level embedding x g ;
[0147] Flatten the final simplex feature matrix into vectors respectively, and then concatenate them to obtain the graph-level embedding x g .
[0148] Step 3: Heterogeneous graph molecule classification to obtain the molecule classification result.
[0149] To further verify the effectiveness and efficiency of the proposed heterogeneous graph embedding method in the molecule classification task, the molecule heterogeneous graph embedding generated by this method is applied to the molecule graph classification task, and the classification accuracy and computing time consumption are evaluated.
[0150] The x obtained in step 2 g is an L-dimensional binary hash code. The L-dimensional binary hash code x is extended through the bitwise inversion and concatenation operation g . The specific steps are as follows:
[0151] S3.1 Perform a bitwise inversion operation on each binary bit of x g , that is, convert 0 to 1 and 1 to 0, to obtain the inverted vector flip(x g );
[0152] S3.2 Concatenate the original vector x g and the inverted vector flip(x g ) into a 2L-dimensional extended vector x e ;
[0153] S3.3 Input x e into a logistic regression classifier for classification. For example: Assume x g= [1, 0, 1], then flip(x g ) = [0, 1, 0], and after concatenation, we get x e = [1, 0, 1, 0, 1, 0], and finally input it into the logistic regression classifier for classification.
[0154] Related character definitions:
[0155] x g : An L - dimensional binary hash code, representing the graph - level embedding of the heterogeneous graph g.
[0156] flip(x g ) : The vector obtained by bit - wise inversion of x g .
[0157] x e : The extended 2L - dimensional vector, obtained by ...
[0158] Experimental example:
[0159] In this application, two molecular heterogeneous graph classification datasets are selected in the examples: the sr - ARE dataset and the nr - BIO dataset. These two datasets are used for molecular classification tasks in the biomedical field and have important practical significance and application value.
[0160] sr - ARE: It contains 7167 molecular graphs, where different types of nodes represent specific kinds of atoms and the edges represent chemical bonds between atoms. The sr - ARE dataset contains a large number of molecular graph samples, which pose high requirements for the performance evaluation of the algorithm and its performance in large - scale data processing, and help to comprehensively examine the generalization ability and computational efficiency of the embedding method.
[0161] nr - BIO: The dataset contains 48542 molecular graphs and is a large - scale molecular heterogeneous graph classification dataset. The structural characteristics of its nodes and edges are similar to those of the sr - ARE dataset, and both construct the molecular graph structure based on atom types and chemical bond information. This dataset is large - scale and suitable for evaluating the scalability and accuracy of the model in large - scale molecular graph classification tasks, and poses high challenges to the computational efficiency and expressive ability of the embedding method.
[0162] The process of molecular classification of specific molecules using the method of this application:
[0163] In the method of the present invention, molecular graph data is first input. These molecular graphs contain nodes (representing atoms) and edges (representing chemical bonds between atoms), as well as related isomeric information. The data is preprocessed. First, simplices are extracted and the corresponding Hodge Laplacian matrix is calculated. Next, the Hodge Laplacian matrix is processed through element-wise product and second-order operations to construct local and global graph structure information enhancement operators. Then, the simplex features are iteratively updated through these enhancement operators and a random hashing algorithm, and finally, a simplex feature vector with rich information is generated. Finally, through a feature concatenation operation, simplex features in multiple dimensions are concatenated into the final heterogeneous graph graph-level embedding. The heterogeneous graph embedding is input into a logistic regression classifier after being expanded by inverse concatenation, and the classifier outputs the classification result (0 or 1) of the molecule.
[0164] Results of molecular classification of specific molecules using the method of this application:
[0165] For the sr-ARE dataset, when using the method of this application for molecular classification, the results output by the logistic regression classifier are 0 or 1. The classification result of 0 indicates that the molecule is inactive in the antioxidant response element (ARE) stress response test, and 1 indicates that the molecule is active in the test. This classification result can help in the initial screening stage of drug development to quickly identify compounds with potential antioxidant activity, thus significantly improving the screening efficiency, shortening the drug screening time, and enhancing the screening accuracy.
[0166] For the nr-BIO dataset, when using the method of this application for molecular classification, the logistic regression classifier outputs 0 or 1. The classification result of 0 indicates that the molecule has low or no activity in the nuclear receptor signal transduction experiment, and 1 indicates that the molecule has strong activity. This result provides an important reference for drug discovery, especially in the drug development targeting nuclear receptors, and can effectively guide researchers to preferentially select compounds with high activity for subsequent research and development, further improving the success rate and efficiency of drug development.
[0167] A schematic description of the dataset is shown in Table 1:
[0168] Table 1 Schematic table of dataset description
[0169]
[0170]
[0171] The existing muxGNN heterogeneous graph classification method, HGCNs heterogeneous graph classification method, and the heterogeneous graph classification method of this application are respectively used to evaluate the classification effect on the above datasets. muxGNN introduces a multi-graph neural network for heterogeneous graphs, models heterogeneity from relation-specific graphs and coupled graphs, and then captures multi-faceted semantic contexts in the coupled attention mechanism. HGCNs integrates relational GCN layers into a heterogeneous graph classification framework, performs message passing on different types of nodes and edges, and then conducts feature aggregation to obtain graph-level representations.
[0172] Experimental configuration: Five-fold cross-validation is adopted, and logistic regression is used as the classifier. The sizes of the simplex sets are fixed as |S0| = 20, |S1| = 30, and |S2| = 20 to ensure fixed-length embeddings.
[0173] For the sr-ARE dataset, the number of iterations R = 2, and for the nr-BIO dataset, the number of iterations R = 3; the maximum simplex dimension K extracted from both datasets is 2. The time limit for the experiment to run is 24 hours. When the method fails to complete the calculation within the specified time, the result is represented by the symbol "\", indicating a timeout. The classification results and efficiency of the molecular heterogeneous graph are shown in Table 2:
[0174] Table 2 Schematic table of performance comparison of heterogeneous graph classification methods
[0175]
[0176] In the prior art, the molecular classification methods adopted also use 0 or 1 as the classification results on the sr-ARE and nr-BIO datasets, respectively indicating whether the compound meets specific classification criteria. It can be seen from Table 2 that these methods rely on complex graph neural network structures and a large number of parameter learnings, resulting in a high computational complexity. Specifically, in the sr-ARE dataset, although the existing methods can provide relatively accurate classification results, on large-scale datasets, especially the nr-BIO dataset, they often face problems such as timeouts or storage bottlenecks. This is because the nr-BIO dataset contains a large number of molecular graph samples, which require a large amount of computing resources to process. Therefore, when dealing with large-scale datasets, the existing methods often cannot effectively cope with these computing and storage challenges, limiting their scalability in practical applications.
[0177] In contrast, the method of the present application performs heterogeneous graph embedding by randomly generating multiple groups of hash functions, avoiding the complex mathematical operations and a large number of parameter learning processes in traditional methods, thereby significantly reducing the time and storage overhead. Especially on the nr-BIO dataset, the method of the present application can effectively avoid the timeout problem and shows significant advantages in running time. In addition, the classification accuracy of the method of the present application on the sr-ARE dataset is better than that of HGCNs. Although it is slightly lower than that of muxGNN, the gap is small and it still maintains high accuracy. Therefore, the method of the present application comprehensively considers accuracy, running time, and storage overhead, and shows higher efficiency and better performance when dealing with large-scale datasets.
[0178] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A molecular classification method based on heterogeneous graph embedding, characterized in that, It includes the following steps: Step 1: Extract simplices from the heterogeneous graph g; Process the structural information of the heterogeneous graph g based on the extracted simplex to obtain the local and global information enhancement operator N k ; Process the feature information of the heterogeneous graph g based on the extracted simplex to obtain a feature matrix Step 2: Based on the local and global information enhancement operator N k and the feature matrix perform hash iteration on the heterogeneous graph g to obtain the graph-level embedding x g ; Step 3. Perform molecular classification on the graph-level embedding x g to obtain the molecular classification result.
2. The molecular classification method based on heterogeneous graph embedding according to claim 1, wherein The heterogeneous graph g is from the molecular heterogeneous graph dataset in the known TUDataset, and the molecular heterogeneous graph dataset contains various types of heterogeneous graphs g.
3. The molecular classification method based on heterogeneous graph embedding according to claim 2, wherein The specific process of extracting simplices from the heterogeneous graph g is as follows: All nodes in the node set V of the heterogeneous graph g are used as simplices of the 0th dimension and added to the simplex set S0 of the 0th dimension. S0 contains all 0-simplices; All edges in the edge set E of the heterogeneous graph g are used as simplices of the 1st dimension and added to the simplex set S1 of the 1st dimension. S1 contains all 1-simplices; Initialize the Simplex tree with the simplex sets S0 and S1, and use the expansion method in the Simplex tree to recursively expand starting from the 2nd - dimensional simplex, expanding layer by layer to obtain higher - dimensional simplices until the maximum dimension K is reached, so as to extract the 2nd - dimensional simplex to the K - dimensional simplex respectively, obtaining S2, …, S K , S2, …, S K which are the sets from the 2nd dimension to the maximum dimension K respectively.
4. The molecular classification method based on heterogeneous graph embedding according to claim 3, wherein Obtain the local and global information enhancement operator N k The specific process is as follows: Calculate the simplex set S of the k-th dimension k corresponding Hodge Laplacian matrix L k , where k ∈ {0, 1, …, K}; Based on the Hodge Laplacian matrix L k , construct the local information enhancement operator M k ; Based on the local information enhancement operator M k we further construct the local and global information enhancement operator N k .
5. The molecular classification method based on heterogeneous graph embedding according to claim 4, characterized in that Obtain the feature matrix The specific process is as follows: Create a zero matrix of size |S k |×|T v | Traverse S k For each k-simplex in it, construct the heterogeneous information vectors of all the nodes it contains; Traverse each k-simplex in S k and perform an element-wise logical OR operation on the heterogeneous information vectors of all the nodes it contains to obtain a binary vector of length |T v |, which is used as the feature vector of Fill the eigenvectors of each k-simplex into a zero matrix to obtain an eigenmatrix from dimension 0 to the maximum dimension K 6. The molecular classification method based on heterogeneous graph embedding according to any one of claims 1-5, characterized in that, Obtain the graph-level embedding x g The specific process is as follows: (i) Initialization; Let r = 1 and k = 0. Where: r represents the rth iteration, and k represents the simplex of the kth dimension; (ii) Judgment of the iteration loop condition; Condition: If r ≤ R, then execute step (iii); otherwise, jump to step (viii); (iii) Judgment of the simplex dimension loop condition; Condition: If k ≤ K, then execute step (iv); otherwise, jump to step (vii); (iv) Calculate the temporary feature matrix; Calculate the temporary feature matrix of the k-simplex in the r-th iteration : Wherein: is the feature matrix of the k-simplex after the (r-1)-th iteration; W (r-1,k) is used to update random hash function matrix; (v) Symbol operation; Apply the sign operation sgn to the temporary feature matrix and convert each element of to a binary hash code to obtain the updated feature matrix Updated feature matrix The expression is as follows: (vi) Update the simplex dimension to k'; Increase k by 1, that is, k' = k + 1. Jump back to step (iii); (vii) Update the iteration number to r'; Increase r by 1, that is, r' = r + 1; Reset k = 0; Jump back to step (ii); (ⅷ), Generate graph-level embedding x g ; Flatten the final simplex feature matrix into vectors respectively and then concatenate them to obtain the graph-level embedding x g .
7. The molecular classification method based on heterogeneous graph embedding according to claim 6, wherein For the graph-level embedding x g The specific process of molecular heterogeneous graph classification is as follows: S3.
1. Invert each binary digit of x g , that is, convert 0 to 1 and 1 to 0, to obtain the inverted vector flip(x g ); S3.
2. Concatenate the original vector x g and the inverted vector flip(x g ) into a 2L-dimensional extended vector x e ; S3.
3. Input x e into the logistic regression classifier for classification; Specifically: Let x g = [1, 0, 1], then flip(x g ) = [0, 1, 0]. After concatenation, we get x e = [1, 0, 1, 0, 1, 0], and finally input it into the logistic regression classifier for classification. where: flip(x g ) is the vector obtained by bitwise inversion of x g ; x e is the extended 2L-dimensional vector, obtained by .
Citation Information
Patent Citations
Heterogeneous graph data processing method and device
CN111626311A
Self-supervised heterogeneous graph node classification method based on potential pattern embedding
CN118245929A