Method and system for solving performance degradation problem of link prediction model through representation weighting
By introducing weighted terms of the distance information between structural features in the link prediction model, the problem of performance degradation of the link prediction model on the sparse graph is solved, and more robust and fair prediction performance is achieved.
Patent Information
- Application Number
- CN202510218239.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
Existing link prediction models based on graph neural networks show performance degradation problems on sparse graphs, especially when the test ratio increases, the model's prediction performance decreases.
By adding a vertex characterization weighting term controlled by the distance information between structural features, the model can make link predictions based on the importance of structural features. The specific method is to use the degrees of nodes to estimate the distance between structural features, thereby obtaining the weight of the weighted term and applying it to the existing link prediction model with enhanced structural features.
This method can effectively balance the relationship between structural features and vertex characterization, improve the performance of the link prediction model, especially on sparse graphs at higher test proportions, which significantly improves the prediction accuracy and fairness of the model.
Smart Images

Figure CN120218142A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to a method and system for solving the problem of performance degradation of a link prediction model through feature weighting. Background Art
[0002] The correlation relationships between data are ubiquitous, such as friendship relationships, purchase relationships, browsing relationships, etc. The graph data format can well represent the correlation relationships between data. Therefore, graph data has become one of the most important data types. A basic graph data consists of vertices and edges. For example, in a friendship relationship, each person constitutes a vertex, and the friendship relationship is the edge. The link prediction task aims to discover the existence rules of edges in graph data. Thanks to the development of graph neural networks, in recent years, link prediction models based on graph neural networks process graph data through graph neural networks to obtain the representations of vertices, and then process the representations of two target vertices, such as through the way of vector dot product, to reconstruct the link probability of the vertex pair, and train the model with training data. The link prediction model based on structural feature enhancement is based on the link prediction model based on graph neural networks. It manually constructs a structural quantity feature vector between target vertex pairs, such as the number of common neighbors, etc., and combines it with the representations of two vertices obtained by the graph neural network to obtain the representation of the vertex pair. Finally, the representation of the vertex pair is input into a classifier to obtain the link probability of the two vertices.
[0003] The traditional link prediction method based on graph neural networks obtains the representation of vertices based on the paradigm of message passing neural networks. The message passing neural network aggregates information according to the target vertex and the information of its surrounding vertices to obtain the representation of the target vertex. Thus, it can be seen that the message passing neural network is essentially for a single vertex. For the link prediction task, due to the nature of the task itself being to predict the relationship between vertex pairs, the structure information specific to vertex pairs is very crucial for prediction. For example, the number of common neighbors between vertex pairs means a higher link probability between vertex pairs in many scenarios, etc. Therefore, the graph neural network designed for a single vertex cannot specifically capture this information between target vertex pairs, thus limiting the performance of the model.
[0004] Since the link prediction model enhanced based on structural features uses the number information of structural features between manually selected point pairs to form the representation of point pairs, the above performance limitations are avoided. However, these methods ignore the impact of graph sparsity on structural features. Since graph sparsity will greatly affect the distribution of the number of structures, and the distribution of structures is closely related to the classification result, graph sparsity will greatly affect the prediction result of the model. Existing studies only test the model under the condition of 10% missing edges (also called the test ratio), which means that the structural features in the graph are relatively rich, so this impact has not been tested and studied. Since many graphs in reality are sparse, the existing work has serious performance hidden dangers in practical applications. By performing performance tests on the powerful structural feature-enhanced link prediction models WalkPool (abbreviated as WP) and ELPH on the commonly used Cora dataset at test ratios of 10%, 25%, and 45%, and comparing them with the traditional graph neural network-based link prediction model GAE, as Figure 1 shown, we find that although WP and ELPH utilize more feature information than GAE, as the test ratio increases, they both obtain a performance worse than GAE. The present invention refers to this as the performance degradation problem. From this problem, it can be seen that the performance of the current most powerful structural feature-enhanced link prediction models on relatively sparse graphs has extremely serious problems. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method and system for solving the performance degradation problem of link prediction models through representation weighting.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A method for solving the performance degradation problem of link prediction models through representation weighting, comprising the following steps:
[0008] Determine the importance of structural features according to the distance between structural features, and assign similar importance to similar structural features;
[0009] According to the importance of structural features, use a link prediction model enhanced based on structural features to perform link prediction.
[0010] Further, the performance degradation problem of the link prediction model refers to that as the test ratio increases, the distance between structural features becomes smaller, thereby causing the prediction performance of the link prediction model to decline.
[0011] Further, by adding a vertex representation weighting term controlled by the distance information between structural features, the link prediction model can assign similar importance to similar structural features.
[0012] Further, the link prediction model enhanced based on structural features after adding the weighted term is represented in the following form:
[0013]
[0014] where y i,j is the predicted link probability between nodes i and j; f(x i,j ||g(z i ·z j )) represents the classifier of the existing link prediction model enhanced based on structural features; x i,j is the structural feature between nodes i and j; z i and z j are the embeddings of nodes i and j; f(·) is a neural network classifier that maps the feature vector of a point pair to the predicted link probability; g(·) is a dot product or element-wise product operation; d(·) represents the distance between a structural feature and its nearby structural features; represents a trainable operation that maps distance information to the weight of the dot product term, and is continuous and used to correct the importance.
[0015] Further, the degree of a node is used to estimate the distance between structural features, so as to obtain the weight value of the weighted term, and the link prediction model enhanced based on structural features is represented in the following form:
[0016]
[0017] where deg(i) represents the degree of node i, and the sum of the degrees of the node pair is used to reflect the distance between the structural features of the node pair and its nearby structural features.
[0018] A commodity recommendation method uses the above method to evaluate the importance of structural features in commodity data, enabling the link prediction model to assign weights to the importance of the attribute information of users and commodities, thereby improving the accuracy of commodity recommendations.
[0019] A friend recommendation method based on a social network uses the above method to evaluate the importance of structural features in the social network, enabling the link prediction model to assign weights to the importance of user feature information, thereby improving the performance of friend recommendations.
[0020] A system for solving the problem of performance degradation of a link prediction model through representation weighting includes:
[0021] A structural feature importance evaluation module, which is used to determine the importance of structural features according to the distance between structural features and assign similar importance to similar structural features;
[0022] A link prediction module, which is used to perform link prediction by adopting a link prediction model based on enhanced structural features according to the importance of structural features.
[0023] The beneficial effects of the present invention are as follows:
[0024] The present invention studies the common problem existing in the current most powerful link prediction model, that is, the performance degradation problem of the link prediction model based on enhanced structural features, and proposes a method DIP that only brings negligible additional model complexity and graph data processing overhead to solve this problem. Experiments show that the model enhanced by DIP can obtain performance that is consistently as good as that of themselves and GAE, which indicates that DIP can balance the relationship between structural features and vertex representations and solve the performance degradation problem. In addition, the experiments also verify that DIP using vertex degree information does not cause unfairness in performance on vertices with different degrees, but instead improves the fairness of the model, thus further demonstrating the advantages of the DIP method. Description of the Drawings
[0025] Figure 1 It is the AUC performance index of different models under different test ratios on the Cora dataset.
[0026] Figure 2 It is the visualization of structural features of the Cora dataset under test ratios of 10% (a), 25% (b), and 45% (c).
[0027] Figure 3 It is the comparison of the average standard deviation of the link probabilities predicted by the structure-enhanced link prediction model using complete features and only structural features.
[0028] Figure 4 It is the visualization diagram of the structural features of low-degree (a) and high-degree (b) samples of the Cora dataset under a 10% test ratio.
[0029] Figure 5 It is the comparison of the average standard deviation of the link probabilities predicted by the existing model and the model enhanced by DIP.
[0030] Figure 6 It is the performance comparison of WP and DIP+WP on vertices with different degrees under test ratios of 10% (a) and 45% (b) of the Cora dataset. Detailed Implementation Manner
[0031] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and the drawings.
[0032] The present invention aims to solve the problem of performance degradation existing in the existing link prediction models. Starting from the impact of the test ratio on the structural features, the present invention conducts theoretical analysis and finds that a high test ratio will lead to a decrease in the distance between structural features, which further causes the model to be unable to effectively classify through structural features, thereby reducing the performance of the model and resulting in the problem of performance degradation. In order to further verify the correctness of the theoretical analysis, the present invention also empirically verifies the above conclusion, making the analysis more comprehensive and reliable. According to the above analysis, the present invention believes that the model should determine the importance of structural features in model prediction based on the distance relationship between structural features. However, none of the existing models can achieve this: on the one hand, since similar structural features have a similar distance relationship with other structural features, they should have similar importance. However, the existing models either cannot determine the importance of structural features based on the distance between structural features or will assign different importance to similar structural features; on the other hand, there is currently a lack of an efficient method to evaluate the distance between structural features in link prediction tasks. To overcome these difficulties, the present invention proves that by adding a simple vertex representation weighting term controlled by the distance information between structural features, the model can assign similar importance to similar structural features. In addition, the present invention also verifies that the distance between structural features can be reflected by the degree of vertices, enabling the model to efficiently estimate the distance between structural features. Combining the above two points, the present invention proposes the DIP (distribution perspective, distribution perspective weighting) method, thus solving the problem of performance degradation.
[0033] The present invention first theoretically and empirically analyzes the reasons for the performance degradation problem. Then, on this basis, it conducts theoretical analysis on the challenges faced by existing models in solving the degradation problem and proposes the DIP method.
[0034] (1) Theoretical analysis of the degradation problem:
[0035] Consider a graph with n nodes, G=(V, E, X), where V is the set of nodes, E is the set of edges, and X is the set of node features. The present invention focuses on undirected graphs, which means that when the edge (i, j) ∈ E, there will also be an edge (j, i) ∈ E. The adjacency matrix is A ∈ {0, 1} n×n , if the edge (i, j) ∈ E, then A(i, j)=1; if the edge then A(i, j)=0. Classical graph neural networks use the node feature set X and the adjacent matrix A to obtain the embedding of each node, denoted as Z = GNN(A, X). After obtaining the node embedding, models such as GAE use the dot product to decode the link probability of node pairs: y i,j = z i ·z j , where y i,jis the predicted link probability, z i or z j is the embedding of node i or j. If there is a link between the node pair, the link probability is constrained to be close to 1, and the node pair is called a positive sample. Otherwise, the link probability is constrained to be close to 0, and the node pair is called a negative sample.
[0036] The structural feature between nodes i and j is defined as the vector x i,j ={s i,j,1 , s i,j,2 , … s i,j,l}, where s i,j,k is the value of the k-th item of the structural feature, which is the number of a certain structure (such as a triangle or a path of a certain length), and l represents the length of the feature. A structure often consists of a fixed number of links. If this number is h, it is called an h-hop structure. Let the hop of the structure related to the k-th item of the structural feature be h k , and let When any link in a structure is missing, the theme itself will also be missing. The predictor of the link prediction model based on enhanced structural features can be expressed in the following form:
[0037] y i,j = f(x i,j || g(z i , z j )) (1)
[0038] where f(·) is a neural network classifier that maps the feature vector of the point pair to the predicted link probability, that is, the classification score, such as a multi-layer perceptron (MLP) using the sigmoid activation function. || represents the concatenation operation, and g(·) is an operation such as the dot product or element-wise product. y i,j is the predicted link probability between nodes i and j.
[0039] To analyze the degradation problem, the present invention needs to find the relationship between the test ratio of the model and the performance. Intuitively, the higher the test ratio means the fewer structures between node pairs. Therefore, the test ratio will affect the distribution of structural features. In addition, the distribution of features has a great impact on the performance of the model. Therefore, starting from analyzing the influence of the test ratio on the distribution of structural features, the present invention can discover the influence of the test ratio on the link prediction task.
[0040] The present invention denotes the random variable composed of all values of the k-th item of the structural feature as S k , E(S k ) is the expected value of S k , and Var(S k ) is the variance of S kThe variance. For the test ratio r, to simplify the analysis process, the present invention assumes that the probability of missing for each side is the same. Assuming that the links in different patterns are independently distributed, for a structural feature x i,j ={s i,j,1 , s i,j,2 , …, s i,j,l} The random variable obtained by this structural feature at the test ratio r is where B(·) is the binomial distribution, represents the random variable obtained by the l-th dimension of the structural feature at the test ratio r, and h l represents the jump of the structure related to the l-th item of the structural feature. For convenience, denote as p k . The present invention can analyze the expected value and variance of each item of the structural feature in the above situation:
[0041]
[0042] In the formula, represents the random variable composed of the values of the k-th item of the structural feature at the test ratio r. Therefore, the expected value of the structural feature is exponentially proportional to 1 - r, and the variance of the structural feature is restricted by the test ratio. Define as the random variable representing the difference between the values of the k-th item of the structural feature, and there is:
[0043]
[0044] According to the Chebyshev inequality, there is:
[0045]
[0046] In other words:
[0047]
[0048] where δ > 0. Therefore, for a small δ like 0.01, it can be approximately said that is bounded within . This means that as the test ratio increases, the distance between structural features becomes smaller.
[0049] Then, the present invention analyzes the influence of the test ratio on model prediction. Based on the universal approximation theorem, a two-layer MLP with a non-linear function is a very powerful link prediction model predictor. Therefore, the present invention considers a predictor that takes only structural features as input to analyze the influence of the distance between structural features on model prediction. Assume that the maximum absolute value of the hidden layer parameter W h is The maximum absolute value of the output layer parameter W o is The output dimension of the hidden layer is d h , the hidden layer uses Relu as the activation function, and the output layer uses the Sigmoid activation function (abbreviated as S(·)). Since the absolute value of the derivative of the Sigmoid activation function is not greater than 0.25, for any positive sample (i + , j + ) and negative sample (i - , j - ), there is:
[0050]
[0051] where B h represents the bias term parameter of the hidden layer, B0 represents the bias term parameter of the output layer, represents the structural features of a positive sample.
[0052] The present invention finds that the difference between the prediction scores of positive and negative samples is limited by the test ratio. As the test ratio increases, this difference approaches 0. Therefore, when making predictions by combining the representations of target nodes, as the test ratio increases, the model is easily interfered by less useful structural features, resulting in a decline in performance, and the performance degradation problem is exactly caused by this.
[0053] (2) Empirical analysis of the degradation problem:
[0054] In this section, the present invention empirically verifies the correctness of the theoretical analysis in the previous section. First, the present invention visualizes the structural features of different test ratios to verify the correctness of Formula 7. Inspired by WalkPool, the present invention selects the values of node-level random walk attributes and link-level random walk attributes with different hop numbers as the structural features for visualization:
[0055]
[0056] where A h i,i represents the value of the i-th row and i-th column after taking the h-th power of the adjacency matrix, and its structural meaning is the number of routes from vertex i back to vertex i after h hops; A h i,j represents the value of the i-th row and j-th column after taking the h-th power of the adjacency matrix, and its structural meaning is the number of routes from vertex i to vertex j after h hops; A h j,j represents the value of the j-th row and j-th column after taking the h-th power of the adjacency matrix, and its structural meaning is the number of routes from vertex j back to vertex j after h hops.
[0057] Considering that methods such as SEAL or WalkPool operate on an ego - centered graph where nodes are no more than two hops away from the target node pair, when h = 4 is selected, the present invention can obtain a practical and reasonable structural feature. Figure 2 For the t - SNE visualization of the structural feature under h = 4, the test ratios are 10%, 25%, and 45% respectively. It is easy to find that the higher the test ratio, the closer the structural features of positive and negative samples are, and the more difficult it is for the model to classify positive and negative samples in the structural feature space. Therefore, the distance between structural features changes according to the test ratio, which is consistent with the result of Equation 7.
[0058] To verify the correctness of Equation 8, the present invention calculates the average standard deviation of the link probabilities of test samples given by four models: well - trained WalkPool, WalkPool with only structural features (denoted as WalkPool -), ELPH, and ELPH with only structural features (denoted as ELPH -) on the Cora dataset with different test ratios. As Figure 3 shown, the evaluation metrics obtained by models using only structural features are always lower than those using complete information. In addition, the higher the test ratio, the closer the evaluation metric is to 0, and the smaller the evaluation metric, the smaller the standard deviation, and the closer the evaluation metric is to 0, the closer the standard deviation is to 0. Therefore, this experiment shows that the difference in predicted values is restricted by the test ratio, which verifies the conclusion of Equation 8.
[0059] (III) Challenges faced and proposed methods:
[0060] To solve the problem of performance degradation, for node pairs with useless structural features, the model should pay more attention to their representation information. However, directly evaluating the usefulness of structural features for the link prediction task during the training process is very challenging. Based on the analysis in the previous two sections, the test ratio affects the distance between structural features, which in turn leads to changes in the usefulness of structural features for the link prediction task. The overall change in the distance of structural features reflects the overall usefulness, while the model actually needs to evaluate the usefulness of each sample at a finer granularity during prediction. Therefore, the present invention measures the distance between a structural feature and its nearby structural features to reflect their usefulness. In other words, the importance of a node representation for prediction depends on the distance between the structural feature and its nearby structural features. Considering that the predictor uses the dot product of target node pair representations, this goal can be formalized as:
[0061]
[0062] d(·) represents the distance between a structural feature and its neighboring structural features, and t(·) represents the projection of distance information onto the importance of node representations. Since similar structural features have similar distances to other structural features, and similar distance relationships imply similar classification difficulties, for different node pairs with similar structural features, the importance of their node representations for prediction should also be similar, which means that t(d(·)) should be continuous. Consider a structure-enhanced link prediction predictor. When the predictor takes the dot product of both structural features and node representations as inputs, we have:
[0063]
[0064] where denotes the Hadamard product, is the derivative of Relu, [: , l + 1] represents the column vector formed by the (l + 1)-th column of the hidden layer weight term, and ":" here means taking all values of the row index. l is the length of the structural feature. If Boole[1, k] = 1, otherwise Boole[1, k] = 0, where B h [k] represents the value of the k-th term of the hidden layer bias term. Therefore, when the model parameters are fixed, the importance of node representations depends on y i,j (1 - y i,j ) and Boole. For a trained model, y i,j (1 - y i,j ) essentially depends on the label of the node pair, and each term of Boole is discontinuous with respect to the structural feature, unless the input of Relu is always positive or non-positive. However, when the input of Relu is always positive or non-positive, the value of that term is constant. Therefore, the importance of node representations is either constant or discontinuous with respect to the structural feature. So, the existing link prediction models based on structure feature enhancement cannot meet the requirements of Equation 10, and the performance of the models is limited. To meet the requirements, the present invention designs a simple weighting term to enhance the existing model:
[0065]
[0066] where f(x i,j ||g(z i ·z j )) represents the classifier of the existing structure feature enhanced link prediction method, represents a trainable operation that maps distance information to the weight of the dot product term, and is continuous. It is used to correct the importance to achieve the importance target t(·). If for any k, the k-th term of Boole is discontinuous, then W o [1, k] = 0 or Wh [k, l + 1] = 0, then is continuous. In this way, the goal of Equation 10 can be achieved.
[0067] After that, the present invention hopes to find an effective method to measure the distance between structural features and nearby structural features. In link prediction, the most useful structures contain links directly connected to the target node. Therefore, the number of links a node has, i.e., the degree of the node, may reflect the distribution characteristics of structural features. As can be seen from the foregoing, the upper bound of the distance between structural features is determined by the variance and mean of their distribution. Therefore, the present invention assumes that the node degree can reflect the distance relationship between structural features. To verify this assumption, in Figure 4 , the present invention sums the degrees of target nodes and respectively emphasizes node pairs with lower degrees and higher degrees, and visualizes the structural features at a 10% test ratio of the Cora dataset. It can be seen that node pairs with different degree distributions have different structural feature distributions, indicating that the degree can reflect the distribution characteristics of structural features. In addition, the present invention also finds that the positive and negative samples with lower degrees are closer than those with higher degrees, which shows that the degree can indeed reflect distance information. By using the sum of the degrees of node pairs to reflect the distance between the structural features of node pairs and nearby structural features, the link prediction model based on structural feature enhancement designed by the present invention has the following paradigm:
[0068]
[0069] where deg(i) represents the degree of node i. Since the function has only inputs and outputs of length 1, and the degrees of target node pairs are easily obtained, the additional cost added by DIP to the original method is very small.
[0070] Key points of the present invention:
[0071] 1. The present invention notices the performance degradation problem existing in the link prediction model based on structural feature enhancement. The present invention proves that a higher test ratio will lead to a decrease in the distance between structural features, thus resulting in a degradation problem.
[0072] 2. To solve the degradation problem by balancing the relationship between node representations and structural features, the present invention proposes the DIP method. DIP meets the relationship requirements between node representations and structural features by adding a simple weighted term of node representations. In addition, DIP uses the degree of nodes to estimate the distance information between structural features, thereby obtaining the weight value of the weighted term.
[0073] 3. The present invention conducts experiments on six datasets using three different test ratios. The experimental results show that the performance of the DIP-enhanced model is always better than that of itself or the GAE, which verifies that DIP can balance the relationship between structural features and node embeddings and solves the problem of performance degradation.
[0074] Effects of the present invention:
[0075] 1) Datasets: The present invention conducts experiments on six datasets: Cora, Citeseer, Pubmed, Texas, Wisconsin, and Cornell. Their statistics are shown in Table 1. The present invention randomly selects 10%, 25%, or 45% of the edges as the test set, 5% of the links as the validation set, and considers the remaining 85%, 70%, or 50% of the edges as the training set. The same number of negative samples are selected for training, validation, and testing.
[0076] 2) Comparative methods: To show the general effectiveness of DIP, the present invention combines it with the structure-enhanced link prediction models WalkPool (abbreviated as WP) and ELPH, and compares it with the basic WP (the currently optimal model on the evaluation dataset), ELPH, and GAE (using the same strategy as WP and ELPH to obtain node embeddings). The DIP-enhanced WP is denoted as DIP+WP, and the DIP-enhanced ELPH is denoted as DIP+ELPH. To better demonstrate the effect on solving the problem of performance degradation, the present invention forces the basic GNN structures of GAE, WP, and ELPH to be the same. The present invention uses a two-layer MLP as
[0077] 3) Evaluation metrics: The present invention conducts repeated experiments using 5 different random seeds and reports the mean and standard deviation of the AUC / AP metrics.
[0078] 4) Experimental results: Tables 2-5 give the performance comparisons of different models. If the performance of the DIP-enhanced link prediction model is better than the basic model and GAE, it is marked in bold. Based on these results, the present invention can draw the following conclusions:
[0079] First, at any test ratio, DIP can improve the performance of WP and ELPH, which shows that DIP can improve the model's ability to utilize structural features. In addition, at a higher test ratio, the accuracy of DIP tends to decline less, which shows that the performance of DIP is more robust.
[0080] Secondly, even in the case of severe performance degradation problems, such as Cora or Citeseer at a 45% test ratio, the DIP-enhanced link prediction model can achieve better performance than GAE, which demonstrates the effectiveness of DIP in solving performance degradation problems.
[0081] Finally, in most cases, in the experimental settings of the present invention, the accuracy of ELPH is lower than that of WP. However, when the test ratio is high, the decrease in the accuracy rate of ELPH is less than that of WP. Even on Citeseer, when the test ratio is 45%, the accuracy rate of ELPH is even higher than that of WP. Therefore, the present invention believes that although WP makes better use of structural features, its ability to balance the relationship between structural features and node embeddings is poor. It is noted that the performance of DIP+WP is always better than that of DIP+ELPH, which indicates that DIP can significantly improve the model's ability to balance the relationship between structural features and node embeddings.
[0082] 5) Impact on prediction scores: To more comprehensively evaluate the performance of DIP, the present invention calculates the average standard deviation of the predicted link probabilities of test samples given by WP, DIP+WP, ELPH, and DIP+ELPH trained on the Cora dataset with different test ratios, as Figure 5 shown. The present invention can find that the DIP-enhanced model obtains better evaluation metrics than the ordinary model. Therefore, DIP can give more distinguishable predictions, and the node representations and structural features cooperate better with each other. The present invention also finds that when the test ratio increases, ELPH obtains higher evaluation metrics than WP. This verifies the last conclusion obtained from the analysis in the experimental results, that is, ELPH can balance the relationship between structural features and node embeddings better than WP.
[0083] 6) Fairness analysis: DIP uses the degree of nodes to learn the weighted scores of node embedding terms. Therefore, it may lead to unfairness in performance on nodes with different degrees. The AUC metric on nodes with different degrees is analyzed using WP and DIP+WP on the Cora dataset, as Figure 6 shown. DIP has a certain improvement in the performance of nodes with different degrees, and the lower the performance of the vertices of this degree on WP, the stronger the degree of improvement. In addition, the standard deviation of AUC on nodes with different degrees is also calculated. When the test ratio is 10%, the standard deviation of the WP model is 0.0208, and the standard deviation of DIP+WP is 0.0175; at a 45% test ratio, the standard deviation of WP is 0.0446, and the standard deviation of DIP+WP is 0.0333. It can be seen that DIP+WP has more stable performance among nodes with different degrees. Therefore, enhancing the model through DIP can bring better fairness to the model.
[0084] Table 1 Statistical data of the dataset
[0085] Dataset Cora Citeseer Pubmed Texas Wisconsin Cornell Number of vertices 2708 3327 19717 183 251 183 Number of edges 5278 4552 44324 279 450 277 Average degree 3.90 2.74 4.50 3.05 3.59 3.03 Standard deviation of degree 5.23 3.38 7.43 7.83 7.94 7.04
[0086] Table 2 Experimental results of Cora, Citeseer, and Pubmed datasets under the AUC evaluation metric
[0087]
[0088] Table 3 Experimental results of Texas, Wisconsin, and Cornell datasets under the AUC evaluation metric
[0089]
[0090] Table 4 Experimental results of Cora, Citeseer, and Pubmed datasets under the AP evaluation metric
[0091]
[0092] Table 5 Experimental results of Texas, Wisconsin, and Cornell datasets under the AP evaluation metric
[0093]
[0094] The present invention can be used in fields such as e-commerce and social networks.
[0095] For example, for e-commerce, the method of the present invention can be used to evaluate the importance of structural features such as the same purchase relationship and the order relationship of purchased items in commodity data, enabling the model to assign weights to the importance of user and commodity attribute information, thereby improving the accuracy of commodity recommendations.
[0096] For example, for social networks, the method of the present invention can be used to evaluate the importance of structural features such as common friends or tightly connected groups in social networks, enabling the model to assign weights to the importance of user feature information, thereby improving the performance of friend recommendations.
[0097] Another embodiment of the present invention provides a system for solving the problem of performance degradation of the link prediction model through feature weighting, which includes:
[0098] A structural feature importance evaluation module, configured to determine the importance of structural features according to the distance between structural features, and assign similar importance to similar structural features;
[0099] A link prediction module, configured to perform link prediction using a link prediction model enhanced based on structural features according to the importance of structural features.
[0100] The division of the above modules is only for illustrative purposes. In actual applications, the above functions can be assigned to different functional modules as needed to complete all or part of the functions described in the foregoing method. For the specific working processes of the above modules, reference can be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.
[0101] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the method of the present invention.
[0102] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc). When the computer program stored in the computer-readable storage medium is executed by a computer, the steps of the method of the present invention are implemented.
[0103] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention shall be subject to the scope defined by the claims.
Claims
1. A method for solving the performance degradation problem of a link prediction model by characterizing weights, characterized in that: The following steps are involved: The importance of structural features is determined according to the distance between them, and similar structural features are given similar importance. According to the importance of structural features, a link prediction model based on structural feature enhancement is used to perform link prediction.
2. The method according to claim 1, characterized in that The link prediction model performance degradation problem refers to the fact that as the test ratio increases, the distance between structural features becomes smaller, thereby causing the prediction performance of the link prediction model to decrease.
3. The method according to claim 1, characterized in that: By adding a vertex representation weighting term controlled by the distance information between structural features, the link prediction model can assign similar importance to similar structural features.
4. The method according to claim 3, characterized in that The link prediction model based on structural feature enhancement after adding the weighted term is expressed in the following form: Among them, y i,j is the predicted link probability between nodes i and j; f(x i,j ||g(z i ·z j )) represents the classifier of the existing structural feature enhanced link prediction model; x i,j is the structural feature between nodes i and j; z i 、z j is the embedding of nodes i and j; f(·) is a neural network classifier that maps the feature vector of a point pair to a predicted link probability; g(·) is a dot product or element-wise product operation; d(·) represents the distance between a structural feature and its nearby structural features; represents a trainable operation that maps distance information to dot product weights, and is continuous and is used to correct for importance.
5. The method according to claim 4, characterized in that The distance between structural features is estimated by using the degree of the node, so as to obtain the weight of the weighted item, and the link prediction model based on structural feature enhancement is expressed as follows: Among them, deg(i) represents the degree of node i, and the sum of the degrees of node pairs is used to reflect the distance between the structural features of the node pairs and the nearby structural features.
6. A product recommendation method, characterized in that: The importance of structural features in product data is evaluated by using the method described in any one of claims 1 to 5, so that the link prediction model can weight the importance of attribute information of users and products, thereby improving the accuracy of product recommendations.
7. A friend recommendation method based on social network, characterized in that: The importance of structural features in social networks is evaluated by using the method described in any one of claims 1 to 5, so that the link prediction model can weight the importance of user feature information, thereby improving friend recommendation performance.
8. A system for solving the performance degradation problem of link prediction model by characterization weighting, characterized in that: include: The structural feature importance evaluation module is used to determine the importance of structural features according to the distance between them and to assign similar importance to similar structural features; The link prediction module is used to perform link prediction based on the importance of the structural features by using a link prediction model enhanced based on the structural features.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.