A Link Prediction Method Based on Node Features and Topological Structure
The integration of node features and topology through non-negative matrix factorization enhances link prediction accuracy by addressing the limitations of traditional topology-based methods, particularly in complex networks.
Patent Information
- Application Number
- CN202311512603.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-11-14
AI Technical Summary
The existing topological-based link prediction methods have limitations in terms of accuracy and reliability, and cannot effectively utilize the deep-level characteristics and potential complex relationships of nodes, and the inaccuracy of observation data affects the prediction effect.
The link prediction method based on node characteristics and topological structure is adopted, and the adjacency matrix of the network is decomposed into low-dimensional factor matrix U and V through non-negative matrix decomposition, and combined with the node feature matrix C, the objective function is optimized using the standard gradient descent method to improve prediction accuracy.
Improves the accuracy of link prediction, is suitable for both attributes and non-attribute networks, and can more accurately predict the connection possibilities between nodes.
Smart Images

Figure CN117556380B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of complex network analysis, and particularly relates to a link prediction method based on node features and topological structures. Background Art
[0002] At present, most traditional link prediction methods mainly perform link prediction based on topological structures. However, the topological structure between nodes is not the only reason for the existence of an edge between two nodes, so that ideal prediction performance cannot always be obtained. With the rapid development of data representation and storage technologies, more and more node semantic attributes of networks have been mined. For example, for users in a social network, their node semantic attributes usually include the age, gender, hobbies, occupation, etc. of the users, and these attributes affect the possibility of generating edges between nodes to a certain extent. Generally speaking, the greater the possibility that nodes with similar semantic attributes are connected. That is, the more similar the attributes of two nodes are, the higher the possibility that there is a connection between them. Node similarity metrics based on local and semi-local information have certain applicability in link prediction in medium-sized and large-scale complex networks due to their simplicity and efficiency. However, these metrics usually rely on the common neighbors of two nodes, that is, they calculate the connection probability between two nodes according to the number or topological structure of the common neighbors. Although these metrics can quickly provide similarity estimates between nodes, they may ignore some deeper network features and potential complex associations between nodes, so there are limitations in the prediction accuracy and reliability of some complex networks. In addition, there may be unobserved or misobserved data in the observed or extracted data, which makes the data unable to truly and comprehensively reflect the information contained in the real network, thus affecting the processing of data analysis. Therefore, reconstructing the original network based on the non-negative matrix factorization method is a more effective method to solve such problems. Summary of the Invention
[0003] In order to overcome the above technical problems, the purpose of the present invention is to provide a link prediction method based on node features and topological structures, which not only considers the topological structure in the network, but also flexibly combines the contributions of the topological structure and node features of the network, improving the prediction accuracy.
[0004] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0005] A link prediction method based on node features and topological structures, comprising the following steps;
[0006] Step 1: Initialize AUC, and set AUC = 0; AUC is the accuracy of this link prediction method on a network with n nodes;
[0007] Step 2: Divide the set of known edges E in a network into a training set E T and a test set E P ;
[0008] Step 3: For the training set E T and the test set E P , construct the adjacency matrices A T and A P of their corresponding networks respectively; and generate the corresponding node feature matrix C for the corresponding network;
[0009] Step 4: Perform non - negative matrix factorization on the adjacency matrix of the network, that is:
[0010]
[0011] Step 5: Perform non - negative matrix factorization on the node feature matrix of the network, that is also:
[0012]
[0013] Then, use the link prediction method based on node features and topological structure to calculate the similarity S between any two nodes of the network, that is:
[0014] S=(UV)
[0015] Step 6: Calculate the precision AUC using the training set E T , the test set E P and the similarity matrix S;
[0016] Step 7: Return the precision AUC avg and the similarity matrix S on the network.
[0017] In the said Step 2, use the ten - fold cross - validation method to divide the set of known edges E of the network into a training set E T and a test set E P , E = E T ∪E P , E T ∩E P = Ф. Specifically, first randomly divide the set E of known edges in the network into ten parts. Then, each time take one of them alone as the test set E P , and the remaining nine parts as the training set E T , obtain an AUC result, and repeat this process ten times. Finally, use the average AUC avg result on the test set as its performance.
[0018] In the said Step 3, the process steps of generating the node feature matrix C include:
[0019] Step 3.1: For the attribute network, the present invention directly uses the semantic attributes of each node. In order to make each attribute have the same influence, after normalizing the node semantic attribute matrix, PCA is used to reduce the dimension of the feature matrix, improving the calculation efficiency.
[0020] Step 3.2: For the non-attribute network, the present invention calculates the covariance between the node edge connection information, that is, the adjacency matrix, and uses it as the node feature matrix.
[0021] In the said Step 5, the process steps of the link prediction method of the present invention include:
[0022] Step 5.1: Let f = 0, AUC * = 0, S * = 0, where f ∈ [0, 2] is a super parameter for measuring the importance of node feature information between any two nodes of the network, AUC * is the optimal accuracy of the link prediction model under different super parameters f, and S * is the similarity matrix for generating AUC * ;
[0023] Step 5.2: Calculate the similarity between any two nodes of the network according to the f value, and save the result in the similarity matrix S;
[0024] S = (UV)
[0025] This method flexibly combines the contribution of the network topology structure and node semantic attributes in link prediction to improve the prediction accuracy. Its objective function is as follows:
[0026]
[0027] Update the values of U, V, and Y based on the standard gradient descent method, so that the objective function gradually approaches the minimum value. The U, V, and Y obtained at the end of the iteration are the optimal U, V, and Y. The update rules of U, V, and Y and
[0028] Step 5.3: Use the training set E T , test set E P and the similarity matrix S to calculate the accuracy index AUC of the link prediction model under this f value;
[0029] Step 5.4: If AUC is greater than AUC * , then S * = S;
[0030] Step 5.5: Optimize f so that it can quickly search for the best super parameter;
[0031] Step 5.6: Finally, return the similarity matrix S.
[0032] In step 7, for the link prediction method based on the topological structure and node features of the network, the similarity matrix S between nodes is predicted through the mapping matrices U and V. This matrix contains the possibility of forming a link between nodes that are not yet connected in the network. Specifically, the larger the value at a certain position in the similarity matrix S, the greater the possibility of forming a link between the two nodes corresponding to that position. Using the similarity matrix S, the possibility of connection between two nodes without a link in the network can be predicted.
[0033] Advantages of the present invention.
[0034] The present invention provides a link prediction method based on node features and topological structure. First, it decomposes the adjacency matrix A of the network based on non - negative matrix factorization. That is, A is first decomposed into two low - dimensional factor matrices U and V. Secondly, it deeply excavates the node features, and the node feature matrix C of the network is decomposed into U and Y, improving the accuracy of link prediction. This method is applicable not only to attribute - less networks but also to attributed networks. Brief Description of the Drawings
[0035] Figure 1 is the calculation process of the link prediction method of the present invention.
[0036] Figure 2 is the schematic diagram of the decomposition process of the link prediction method of the present invention. Detailed Embodiments
[0037] In the embodiments of the present invention, the specific implementation process of the technical solution is clearly described in a diagrammatic way. The diagram includes the extraction of the network topological structure, the generation of the node feature structure, the non - negative matrix factorization process, etc. The diagram can help readers better understand the technical implementation details of the present invention and obtain clear operation guidance. It should be clear that the technical solutions described in the embodiments of the present invention only represent some embodiments, not all embodiments. All other embodiments that can be obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0038] Therefore, through detailed description and illustration, the present invention shows a technical person an innovative link prediction method based on node features and topological structure. This method combines the topological structure of the network, node features, and non - negative matrix factorization technology, effectively predicting the connection probability between nodes.
[0039] The present invention will be further described in detail below with reference to the accompanying drawings.
[0040] A link prediction method based on node features and topological structure includes the following steps;
[0041] Step 1: Initialize AUC and set AUC = 0; AUC is the accuracy of this link prediction method on a network with n nodes.
[0042] Step 2: Divide the set E of known edges in the citation network into a training set E T and a test set E P . In this step, use the ten-fold cross-validation method to divide the set E of known edges in the citation network into a training set E T and a test set E P , E = E T ∪E P , E T ∩E P = Ф. Specifically, first randomly divide the set E of known edges in this network into ten parts. Then, each time take one of them alone as the test set E P , and the remaining nine parts as the training set E T , obtain an AUC result, and repeat this process ten times. Finally, use the average AUC avg result on the test set as its performance.
[0043] Step 3: For the training set E T and the test set E P , respectively construct the adjacency matrices A T and A P of their corresponding networks; directly use the semantic attributes of each paper in the citation network, and after normalizing the node semantic attribute matrix, then use PCA to generate its corresponding node feature matrix C;
[0044] Step 4: Perform non-negative matrix factorization on the adjacency matrix of this network, that is:
[0045]
[0046] Step 5: Perform non-negative matrix factorization on the node feature matrix of this network, that is also:
[0047]
[0048] Then, use the link prediction method based on node features and topological structure to calculate the similarity S between any two nodes in this network. The schematic diagram of the process of S is shown in the appendix Figure 2 ;
[0049] Step 5.1: Let f = 0, AUC * = 0, S * = 0, where f ∈ [0, 2] is a hyperparameter measuring the importance of node feature information between any two nodes in this network, AUC *is the optimal accuracy of the link prediction model under different hyperparameters f, S * is to generate the AUC * similarity matrix;
[0050] Step 5.2: Calculate the similarity between any two nodes in the network according to the f value, and save the result in the similarity matrix S;
[0051] S=(UV)
[0052] The present invention flexibly combines the contributions of the network's topological structure and node semantic attributes in link prediction to improve the prediction accuracy. Its objective function is as follows, and an example of its calculation process is shown in the appendix Figure 1 :
[0053]
[0054] Update the values of U, V, and Y based on the standard gradient descent method to gradually approximate the minimum value of the objective function. The U, V, and Y obtained at the end of the iteration are the optimal U, V, and Y. The update rules of U, V, and Y and
[0055] Step 5.3: Use the training set E T , the test set E P and the similarity matrix S to calculate the accuracy metric AUC of the link prediction model under this f value;
[0056] Step 5.4: If AUC is greater than AUC * , then S * =S;
[0057] Step 5.5: Optimize f so that it can quickly search for the best hyperparameters;
[0058] Step 5.6: Finally, return the similarity matrix S.
[0059] Step 6: Use the training set E T , the test set E P and the similarity matrix S to calculate the accuracy AUC;
[0060] Step 7: Return the accuracy AUC avg and the similarity matrix S. The link prediction method based on the network's topological structure and node features predicts the similarity matrix S between nodes through the mapping matrices U and V. This matrix contains the possibility of generating edges between nodes that are not yet connected in the network. Specifically, the larger the value at a certain position in the similarity matrix S, the greater the possibility of generating an edge between the two nodes corresponding to that position. Using the similarity matrix S, the possibility of generating a connection between two nodes that do not have an edge in the network can be predicted.
Claims
1. A link prediction method based on node features and topological structure, characterized in that It includes the following steps: Step 1: Initialize AUC and set AUC = 0; AUC is the accuracy of this link prediction method on a citation network with n nodes; Step 2: Divide the set of known edges E in the citation network into a training set E T and a test set E P , and use the ten-fold cross-validation method to divide the set of known edges E in the citation network into a training set E T and a test set E P , E = E T ∪E P , E T ∩E P = Ф; specifically, first randomly divide the set of known edges E in the citation network into ten parts; then, each time take one of them alone as the test set E P , and the remaining nine parts as the training set E T , obtain an AUC result, and repeat this process ten times; finally, use the average AUC avg result on the test set as its performance; Step 3: For the training set E T and the test set E P , respectively construct their adjacency matrices A T and A P ; and generate its corresponding node feature matrix C for this citation network; The process steps for generating the node feature matrix C include: Step 3.1: For the citation network, directly use the semantic attributes of each paper node in the citation network; in order to make each attribute have the same influence, after normalizing the node semantic attribute matrix, then use PCA to reduce the dimension of the feature matrix to improve the calculation efficiency; Step 3.2: For the network where the paper nodes in the citation network do not have semantic attributes, based on the node connection information, that is, the adjacency matrix, calculate the covariance between the node connection information and use it as the node feature matrix; Step 4: Perform non-negative matrix factorization on the adjacency matrix of this citation network, that is: Step 5: Perform non-negative matrix factorization on the node feature matrix of this citation network, that is: Then, use the link prediction method based on node features and topological structure to calculate the similarity S between any two nodes in this citation network, that is: S=(UV) Step 6: Use the training set E T , the test set E P and the similarity matrix S to calculate the accuracy AUC; Step 7: Return the precision AUC on the citation network avg and the similarity matrix S.
2. The link prediction method based on node features and topological structure according to claim 1, wherein In the said Step 5, the process steps of the link prediction method include: Step 5.1: Let f = 0, AUC * = 0, S * = 0, where f ∈ [0, 2] is a hyperparameter for measuring the importance of node feature information between any two nodes in the citation network. Among them, f > 1 indicates that the contribution degree of node features is greater than the topological structure; f < 1 indicates that the contribution of topological structure information is large; f = 1 means that the contribution degrees of the two are the same. AUC * is the optimal accuracy of the link prediction method under the hyperparameter f, and S * is the similarity matrix for generating AUC * . Step 5.2: Calculate the similarity between any two nodes in this citation network according to the f value, and save the result to the similarity matrix S; S=(UV) This link prediction method flexibly combines the contributions of the topological structure of the citation network and the node semantic attributes in link prediction to improve the prediction accuracy. Its objective function is as follows: Update the values of U, V, and Y based on the standard gradient descent method so that the objective function gradually approaches the minimum value. The U, V, and Y obtained at the end of the iteration, that is, the optimal U, V, and Y, and the update rules for U, V, and Y and Step 5.3: Use the training set E T , the test set E P and the similarity matrix S to calculate the precision metric AUC of the link prediction method at this f value; Step 5.4: If the AUC is greater than the AUC * , then S * = S; Step 5.5: Optimize f to quickly search for the best hyperparameters; Step 5.6: Finally, return the similarity matrix S.
3. A link prediction method based on node features and topological structure according to claim 1, characterized in that, In the said Step 7, the link prediction method based on the topological structure and node features of the citation network predicts the similarity matrix S between nodes through the mapping matrices U and V, and this matrix contains the possibility of generating connections between nodes that are not yet connected in the citation network.
Citation Information
Patent Citations
Link prediction method and system based on deep non-negative matrix factorization
CN110858311A
Link prediction method based on convolutional neural network
CN112860977A