Graph comparison recommendation method based on dual adaptive data enhancement
Through a graph comparison learning method based on node popularity mask and adaptive perturbation data enhancement, the problems of user-project interaction data sparsity and noise interference in the recommendation system are solved, and the robustness and accuracy of the model are improved, especially the recommendation effect for unpopular users or projects.
Patent Information
- Application Number
- CN202510440436.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The existing recommendation system faces the problems of user-project interaction data sparsity and noise interference, which leads to insufficient attention from the model to unpopular users or projects. The traditional random enhancement strategy fails to fully consider the heterogeneity of the user-project interaction graph, which may lead to misleading self-supervised signals.
A dual adaptive data enhancement strategy based on node popularity mask and adaptive perturbation data enhancement is introduced. Combined with the comparison learning mechanism, the message delivery mechanism of the graph neural network is adaptively adjusted to the edge strength and information propagation of higher-order nodes to generate multiple enhancement views for comparison learning.
It significantly improves the robustness and accuracy of the recommendation system, pays more attention to sparse areas, dynamically adjusts the node disturbance intensity, reduces noise interference, improves the attention of unpopular users or projects, and improves the generalization ability and recommendation accuracy of the model.
Smart Images

Figure CN120372084A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of recommendation systems in the field of machine learning, and specifically relates to a personalized recommendation method based on graph neural networks and contrastive learning. Background Art
[0002] With the development of Internet technology, recommendation systems have become a core component of many online service platforms (such as e-commerce, music streaming, social networks, etc.). The goal of a recommendation system is to learn user preferences and provide personalized recommendation services by mining historical interaction data between users and items. In recent years, significant progress has been made in the application of graph neural networks (GNNs) in the field of collaborative filtering. For example, models such as NGCF and LightGCN capture high-order connectivity in the user-item bipartite graph structure through an iterative message passing mechanism and encode complex interaction patterns into low-dimensional embedding representations.
[0003] However, existing recommendation systems still face the following challenges. User-item interaction data usually exhibits characteristics of high sparsity and noise interference, which limits the number of supervision signals and poses a severe challenge to the generalization ability of the model. A small number of high-popularity nodes (such as popular users or items) occupy most of the interaction connections, while most low-popularity nodes (such as niche users or items) have only a small number of interaction records. This uneven distribution results in insufficient attention of the model to niche users or items. Traditional graph contrastive learning methods rely on random augmentation strategies (such as random node / edge deletion) to generate contrastive views, but these strategies do not fully consider the heterogeneity of the user-item interaction graph and may lead to misleading self-supervised signals. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides a graph contrastive learning recommendation method based on dual adaptive data augmentation. By introducing two adaptive data augmentation strategies (data augmentation based on node popularity masking and adaptive perturbation data augmentation) and combining with a contrastive learning mechanism, the robustness and accuracy of the recommendation system are significantly improved. In addition, based on the user-item interaction graph, this method makes full use of the message passing mechanism of graph neural networks to achieve attention to information propagation in sparse regions and adaptive adjustment of the edge strength of high-order nodes, thereby effectively alleviating the challenges brought by the data sparsity problem and the long-tail distribution characteristic. Experimental results show that the recommendation effect of the present invention on key indicators such as Recall and NDCG on public datasets is better than that of existing methods.
[0005] A graph contrastive recommendation method based on dual adaptive data augmentation disclosed by the present invention includes the following steps:
[0006] Step 1: Data preprocessing. Clean the obtained user-item interaction data and delete useless redundant data;
[0007] Discard the interaction data with low scores and construct a user-item interaction graph;
[0008] Step 2: Construction of the supervision signal module. Optimize the preference ranking of users for positive and negative samples through the supervised learning paradigm;
[0009] Use the graph convolution method based on GCN to aggregate the node neighborhood feature information and capture the high-order collaborative filtering signals between users and items;
[0010] Step 3: Data augmentation based on node popularity masking;
[0011] According to the global node degree distribution characteristics, design an adaptive masking strategy, perform a weighting operation on the original adjacency matrix to generate a new masked adjacency matrix, and perform graph convolution feature transformation on the new adjacency matrix to generate the first augmented view for contrast learning tasks with the node embeddings after graph convolution. The contrast learning loss is as follows: where is the contrast learning loss of data augmentation based on node popularity masking, is the batch size, s(·) is the cosine similarity function, τ is the temperature hyperparameter, h represents the embedding of the augmented view, e represents the embedding of the original view, the subscript i represents the selected node among all nodes, and the subscript j represents other nodes except i as the contrast node for node i;
[0012] Step 4: Adaptive perturbation data augmentation;
[0013] Generate a perturbation vector for the embedding of each node using a multi-layer perceptron, and apply the perturbation vector to the node embeddings obtained from the basic graph propagation training to generate the second augmented view. The formula for the contrast learning loss of adaptive perturbation data augmentation is as follows: where, represents the training batch, τ is the temperature parameter for contrast learning, sim(·) represents the calculation function, and the dot product is used here; t′ i is the normalized representation of the embedding e i of node i after being perturbed once to get e′ i , t″ i is the normalized representation of the embedding e i of node i after being perturbed twice to get e″ i , t″ j is the normalized representation of the embedding e j of node j after being perturbed twice to get e″ j , where and e″i is the vector after the first perturbation of node i, e″ i represents the vector after the second perturbation of node i;
[0014] Step 5: Contrastive learning and multi-task optimization;
[0015] Perform graph convolutional feature transformation on the first augmented view and the second augmented view respectively, and optimize the model parameters by combining the contrastive learning mechanism; use the InfoNCE loss function as the objective function of contrastive learning;
[0016] The final multi-task loss function combines the BPR loss, the regularization loss, the data augmentation contrastive learning loss based on the popularity mask, and the adaptive perturbation contrastive learning loss, and the formula is as follows:
[0017]
[0018] where is the task loss of the supervised signal module, is the regularization loss, is the data augmentation task contrastive learning loss based on the popularity mask, is the adaptive perturbation data augmentation contrastive learning task loss, and α and β are hyperparameters;
[0019] Step 6: Model training and recommendation generation;
[0020] Train the model parameters based on the gradient descent method to generate the final user and item embedding representations; calculate the prediction scores according to the user and item embedding representations, and generate a recommendation list.
[0021] Furthermore, the specific method of step 1 is: convert the interaction records of users and items into a sparse adjacency matrix A, where A ij = 1 indicates that user i interacts with item j, otherwise A ij = 0.
[0022] Furthermore, the specific method of step 2 is:
[0023] Embed the initialized embedding vector E u of user u and the embedding vector E v of user v, and embed users and items into a d-dimensional latent space; use the graph convolutional method based on GCN to aggregate the node neighborhood feature information, where represents the user-item interaction matrix, I is the number of users, and J is the number of items; is the normalized adjacency matrix, and is the degree matrix of angles, whose diagonal elements are the degrees of users and items respectively; GCN represents the Graph Convolutional Neural Network, D (u) represents the degree matrix of users, D (v) represents the degree matrix of items, is the adjacency matrix obtained by normalizing the user-item interaction matrix according to the degree matrix;
[0024] In each layer of the GCN neural network, the embeddings of users and items are updated through neighborhood information; for users and items, their new embedding representations are respectively: and The superscript T represents the matrix transpose, that is is the transpose matrix of the (u) matrix; where, z (v) and z
[0025] respectively represent the information aggregated from neighboring items and users to the central node; in order to suppress the over-smoothing effect, a residual connection is applied in the aggregation stage; and The subscript l represents the item number, and the final user and item embeddings are obtained by summing the embeddings of all layers and then average pooling: E final are the final user and item embeddings;
[0026] The BPR loss function is used as the core optimization objective of the supervised signal module, and the formula of the BPR loss is as follows: where, is the predicted score of user u for the positive sample item v pos ; is the predicted score of user u for the negative sample item v neg ; σ(x) is the Sigmoid function; the BPR loss function optimizes the preference ranking of users for items by maximizing the score difference between positive and negative samples.
[0027] Furthermore, the specific method of step 3 is:
[0028] Data augmentation contrast learning based on node popularity masking, by analyzing the degree distribution of nodes, generates an adaptive edge masking matrix to adjust the connection strength of high-popularity nodes in the original adjacency matrix; first, calculate the order of each node; specifically, for users and items, the formula for their degrees is: where, A ui represents the interaction between user u and item i in the user interaction matrix, d u and d iDenote the degrees of user \(u\) and item \(i\) respectively, \(N\) represents the number of users, and \(M\) represents the number of items; perform a logarithmic transformation on the degrees to smooth their distribution: \(d\) u ' = log₂(d u + 1) \(d\) i ' = log₂(d i + 1), where \(d'\) u and \(d'\) i are the degrees after transformation respectively. Subsequently, calculate the maximum and minimum values of the logarithm of the degrees in the global nodes of users and items respectively, which are used for normalization calculation in the subsequent adjacency matrix generation process:
[0029] max ui = max(d u ') + max(d i '), min ui = min(d u ') + min(d i ')
[0030] where, max ui and min ui represent the maximum and minimum values of the logarithm of the degrees in the global nodes of users and items respectively;
[0031] Generate an adaptive graph mask graph mask matrix for dynamically adjusting the user-item interaction graph; this mask is used in combination with the relationship matrix through element-wise multiplication: where is the new adjacency matrix generated after masking, represents the degree of masking of the interaction between users and items. The closer the value is to 0, the lower the importance of this interaction; the closer the value is to 1, the higher the importance of this interaction; ⊙ represents element-wise multiplication;
[0032] Each element in the mask matrix is calculated as follows:
[0033] Normalize to obtain the normalized adjacency matrix represents the element of the matrix in the \(i\)-th row and \(j\)-th column, and is specifically defined as:
[0034]
[0035] where, represents the new adjacency matrix after masking For the element in the $i$-th row and $j$-th column, based on the normalized adjacency matrix, the message passing process is defined as follows:
[0036]
[0037] Wherein, and respectively represent the updated embeddings of users and items, represents the normalized adjacency matrix
[0038] Furthermore, the specific method of step 4 is as follows;
[0039] Let Wherein, $E$ l represents the embedding matrix of the $l$-th layer, $L$ is the number of layers of the graph convolutional network, and $E$ is the final embedding representation; Use a multi-layer perceptron MLP to learn the mean adaptation of user and item embeddings to obtain a perturbation vector, and the calculation formula of the perturbation vector is as follows: $p=(e$ i $\cdot W + b)\odot sign(e$ i $)$, where $p$ represents the perturbation vector adaptively generated by MLP according to the characteristics of each node itself, $W$ is the learned weight, $b$ is the bias term, $\odot$ represents element-wise multiplication, and $sign(e$ i $)$ is the sign function, which is used to retain the direction information of the original embedding, and $e$ i represents the node embedding;
[0040] Given a node $i$ and its representation $e$ i , by applying adaptive perturbation to it, two new representations are generated: $t'$ i $=e$ i $+p'$ i $\cdot \epsilon$, $t''$ i $=e$ i $+p''$ i $\cdot \epsilon$, where $t'$ i , $t''$ i are the node embedding representations in the data augmentation views generated by two independent perturbations respectively, is a small constant used to control the intensity of the perturbation;
[0041] The embeddings of the same node in different views in the two independently generated perturbation views are used for contrastive learning training; Maximize the similarity of embeddings from the same class, while minimizing the similarity between embeddings from different classes.
[0042] Furthermore, in step 5 Wherein, $W$ i represents the parameter of the $i$-th layer in the model, represents the L2 norm of the parameter, and $\lambda$ is the hyperparameter of regularization, which is used to adjust the intensity of regularization.
[0043] Furthermore, after the training in step 6 is completed, the model will use the final embedding vectors for prediction to generate the predicted scores of users for items. During prediction, assume that the embedding vector of user u is u and the embedding vector of item v is v. For user u and item v, their predicted scores are calculated through the inner product: where, is the predicted score of user u for item v. After calculating the predicted scores of each user for all items, they are sorted from high to low to generate the recommendation list for each user.
[0044] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0045] 1. Through the data augmentation strategy based on node popularity masking, the dominant role of high-popularity nodes is reduced, and more attention is paid to the information dissemination in sparse regions, thereby enhancing the attention to cold-start users or items.
[0046] 2. Through the adaptive perturbation data augmentation strategy, the perturbation intensity of each node is dynamically adjusted, avoiding the noise interference that may be introduced by the random augmentation strategy, and significantly enhancing the robustness of the model.
[0047] 3. Through the multi-task optimization strategy, combined with multiple loss functions, the recommendation accuracy, generalization ability, and overall performance of the model are further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the flowchart of the method of the present invention.
[0049] Figure 2 is the schematic diagram of the node popularity masking constructed by the present invention.
[0050] Figure 3 is the overall model architecture diagram of the present invention. SPECIFIC IMPLEMENTATION MANNER
[0051] In order to make the purpose and innovative points of the present invention clearer, the present invention is further introduced in detail below with reference to the accompanying drawings.
[0052] The specific implementation manner of the present invention is as follows:
[0053] 1. First, preprocess the original data, and convert the interaction records of users and items into a sparse adjacency matrix A, where A ij = 1 indicates that there is an interaction between user i and item j, otherwise A ij = 0.
[0054] 2. Optimize the preference ranking of users for positive and negative samples through the supervised learning paradigm. Initialize the embedding vector E uand E v The user and the item are embedded into a d-dimensional latent space. The graph convolution method based on GCN is used to aggregate the feature information of node neighborhoods. where represents the user-item interaction matrix, I is the number of users, and J is the number of items. is the normalized adjacency matrix and is the degree matrix, and its diagonal elements are the degrees of users and items respectively. In each layer, the embeddings of users and items are updated through neighborhood information. For users and items, their new embedding representations are respectively: and where, z (u) and z (v) represent the information aggregated from neighboring items and users to the central node respectively. To suppress the over-smoothing effect, residual connections are applied during the aggregation phase.
[0055] The final user and item embeddings are obtained by summing the embeddings of each layer: and The final user and item embeddings are obtained by summing the embeddings of all layers and then average pooling: Here, E final is the final user and item embedding. To optimize the user's preference ranking for positive and negative samples, the BPR loss function is used as the core optimization objective of the supervised signal module. Specifically, the formula for the BPR loss is as follows:
[0056]
[0057] where, is the predicted score of user u for the positive sample item v pos , is the predicted score of user u for the negative sample item v neg , and σ(x) is the Sigmoid function. The BPR loss function optimizes the user's preference ranking for items by maximizing the score difference between positive and negative samples. The overall model flow chart is as shown in Figure 1 .
[0058] 3. Data augmentation contrastive learning based on node popularity masking. By analyzing the degree distribution of nodes, an adaptive edge masking matrix is generated to adjust the connection strength of high-popularity nodes in the original adjacency matrix. First, calculate the degree of each node. Specifically, for users and items, the formula for their degrees is: where, A ui represents the interaction between user u and item i in the user interaction matrix, d u and d iDenote the degrees of user \(u\) and item \(i\) respectively. To avoid the situation where the maximum and minimum values of the degrees differ too much, logarithmic transformation is performed on the degrees to smooth their distribution: \(d\) u ' = log2(d u + 1) \(d\) i ' = log2(d i + 1), where \(d'\) u and \(d'\) i are the transformed degrees respectively. Subsequently, calculate the sum of the logarithm of the degrees and the maximum and minimum values in the global nodes (users and items), which are used for normalization calculation in the subsequent adjacency matrix generation process: max ui = max(d u ') + max(d i '), min ui = min(d u ') + min(d i '),
[0059] These values are used to generate a normalized weight matrix to ensure that the elements in the mask matrix are distributed within a reasonable range. Based on the node popularity distribution, an adaptive graph mask graph mask matrix is generated for dynamic adjustment of the user-item interaction graph. This mask is used in combination with the relationship matrix through element-wise multiplication: where is the new adjacency matrix generated after masking, represents the degree to which the interaction between users and items is masked. The closer the value is to 0, the lower the importance of the interaction. The closer the value is to 1, the higher the importance of the interaction. Each element in the mask matrix is calculated as follows: To prevent excessive damage to the original graph structure, a constant term is introduced in the calculation, and the mask value range is restricted to [0, 1] through linear transformation. The finally generated mask matrix can adaptively adjust the interaction weights according to the degree distribution of the nodes, thus balancing the influence of popular and unpopular nodes. The process of generating the mask matrix is shown in Figure 2 , and the connection relationships of nodes with higher popularity are weakened. To ensure that the masked adjacency matrix can be effectively used for information propagation, is normalized. The normalized adjacency matrix is defined as Based on the normalized adjacency matrix, the message passing process is defined as follows: where and represent the updated embeddings of users and items respectively. Specifically, given the original view \(E\) and the enhanced view \(H\), the contrastive learning loss for this task is defined as follows: Among them, s(·) is the cosine similarity function, τ is the temperature hyperparameter, h represents the embedding of the augmented view, and e represents the embedding of the original view.
[0060] 4. Adaptive Perturbation Data Augmentation Contrastive Learning
[0061] This module dynamically adjusts the perturbation augmentation strategy for each node by introducing a learnable perturbation vector. The specific steps are as follows: Among them, E l represents the embedding matrix of the l-th layer, L is the number of layers of the graph convolutional network, and E is the final embedding representation. The multi-layer perceptron MLP is used to learn the user and item embedding means adaptively to obtain the perturbation vector. The calculation formula of the perturbation vector is as follows: p = (e·W + b)⊙sign(e i ), where p represents the perturbation vector adaptively generated according to the characteristics of each node through the MLP, W is the learned weight, b is the bias term, ⊙ represents element-wise multiplication, and sign(e
[0062] Given a node i and its representation e i , by applying the adaptive perturbation to it, two new representations are generated: t′ i = e i + p′ i ·∈, t″ i = e i + p″ i ·∈, where t′ i , t″ i are the node embedding representations in the data augmentation views generated by two independent perturbations respectively, is a small constant used to control the intensity of the perturbation.
[0063] The embeddings of the same node in different views in the two independently generated perturbation views are used for contrastive learning training. Maximize the similarity of embeddings from the same category, while minimize the similarity between embeddings from different categories. The formula for the loss of adaptive perturbation data augmentation contrastive learning is as follows:
[0064] Among them, represents the training batch, τ is the temperature parameter of contrastive learning, sim(·) represents the calculation function, and the dot product is used here; z′ i is the normalized representation of node i, where
[0065] 5. Multi-Task Optimization and Prediction Layer
[0066] The final multi-task loss function combines the BPR loss, the regularization loss, the data augmentation contrastive learning loss based on popularity masking, and the adaptive perturbation contrastive learning loss, and the formula is as follows:
[0067]
[0068] where is the supervised signal module task loss, which is used to optimize the embedding vectors of users and items to ensure that the scores of positive samples are higher than those of negative samples. is the regularization loss, which is used to prevent the model from overfitting. is the data augmentation task contrastive learning loss based on popularity masking. is the adaptive perturbation data augmentation contrastive learning task loss. α and β are hyperparameters, which are used as the weights of the data augmentation contrastive learning loss based on popularity masking and the adaptive perturbation contrastive learning loss respectively. The regularization loss controls the model complexity by performing L2 regularization on the model parameters to prevent overfitting. Its formula is: where, W i represents the parameters of the i-th layer in the model, represents the L2 norm of the parameters, and λ is the hyperparameter of regularization, which is used to adjust the strength of regularization.
[0069] 6. Model training and recommendation generation;
[0070] Train the model parameters based on the gradient descent method to generate the final user and item embedding representations; calculate the predicted scores according to the user and item embedding representations, and generate a recommendation list;
[0071] By combining the above multiple loss terms, the design of the final loss function allows the model to optimize multiple objectives simultaneously during training, thereby improving the accuracy and robustness of the recommendation system.
[0072] After training is completed, the model will use the final embedding vectors for prediction to generate the predicted scores of users for items. During prediction, assume that the embedding vector of user u is u and the embedding vector of item v is v. For user u and item v, calculate their predicted scores through the inner product: where, is the predicted score of user u for item v. After calculating the predicted scores of each user for all items, sort them from high to low to generate a recommendation list for each user. The design of multi-task optimization and the prediction layer improves the recommendation performance of the model and also enhances the generalization ability and robustness of the model. The overall model architecture diagram of the present invention is as Figure 3 shown.
[0073] 7. Use the dataset to evaluate the model. After data preprocessing, we divide the dataset into a training set, a validation set, and a test set in the ratio of 7:1:2. To verify the effectiveness of the present invention, experiments were conducted on the following three public datasets: LastFM: It contains approximately 1,892 users, 17,632 artists, and 92,834 interaction records. This dataset has a high degree of sparsity and is suitable for evaluating the performance of recommendation systems in music recommendations. Gowalla: It originated from a location-based social network platform and contains 38,653 users, 40,981 locations (items), and 213,679 check-in records. Tmall: It comes from a large e-commerce platform and contains 47,939 users, 41,390 items, and 2,357,450 interaction records. This dataset covers rich shopping behaviors and is suitable for evaluating the performance of recommendation systems in an e-commerce environment.
[0074] Select the following state-of-the-art recommendation models as the comparison benchmarks:
[0075] LightGCN: A lightweight graph neural network model that propagates neighborhood information through linear transformation and element-wise addition operations;
[0076] SGL: Utilizes classical data augmentation methods such as node deletion and edge deletion to design self-supervised contrastive learning tasks;
[0077] NCL: Constructs contrastive objectives through structural neighbors and semantic neighbors to enhance the quality of node embedding representations;
[0078] LightGCL: A framework of graph neural networks based on singular value decomposition contrastive learning that captures local and global collaborative relationships;
[0079] Adopt the following two widely used evaluation metrics:
[0080] Recall@K measures the proportion of the model hitting the target item among the top K recommended items and reflects the recall ability of the model.
[0081] NDCG@K takes into account the ranking position of items in the recommendation list, assigns higher weights to correct recommendations that are ranked higher, and comprehensively evaluates the accuracy and ranking ability of the model.
[0082] As shown in Table 1, the experimental results indicate that the present invention has achieved significant performance improvements on all datasets. The specific results are as follows:
[0083] On the LastFM dataset, Recall@20 and NDCG@20 reached 0.2219 and 0.2014 respectively, representing an increase of 18.3% and 15.2% compared to LightGCN.
[0084] On the Gowalla dataset, Recall@20 and NDCG@20 reached 0.2402 and 0.1539 respectively, representing a 6.4% and 6.06% improvement compared to LightGCN.
[0085] On the Tmall dataset, Recall@20 and NDCG@20 reached 0.0640 and 0.0445 respectively, representing a 28.8% and 26.1% improvement compared to LightGCN.
[0086] Table 1 Recommendation results of the present invention on three public datasets
[0087]
Claims
1. A graph contrastive recommendation method based on dual adaptive data augmentation, the method comprising the following steps: Step 1: Data preprocessing, cleaning the obtained user-item interaction data and deleting useless redundant data; Discarding the interaction data with low scores and constructing a user-item interaction graph; Step 2: Construction of the supervision signal module, optimizing the preference ranking of users for positive and negative samples through the supervised learning paradigm; Using the graph convolution method based on GCN to aggregate the node neighborhood feature information and capture the high-order collaborative filtering signals between users and items; Step 3: Data augmentation based on node popularity masking; According to the global node degree distribution characteristics, an adaptive masking strategy is designed to perform a weighting operation on the original adjacency matrix to generate a new masked adjacency matrix, and a graph convolution feature transformation is performed on the new adjacency matrix to generate a first enhanced view for contrastive learning tasks with the node embedding features after graph convolution. The contrastive learning loss is as follows: where is the contrastive learning loss of data augmentation based on node popularity masking, is the batch size, s(·) is the cosine similarity function, τ is the temperature hyperparameter, h represents the embedding of the enhanced view, e represents the embedding of the original view, the subscript i represents the selected node among all nodes, and the subscript j represents the other nodes except i as the contrast node for node i; Step 4: Adaptive perturbation data augmentation; The embedding of each node uses a multi - layer perceptron to generate a perturbation vector. The perturbation vector is applied to the node embedding obtained by base graph propagation training to generate the formula for the second - enhanced view adaptive perturbation data - augmentation contrastive learning loss as follows: Among them, represents the training batch, τ is the temperature parameter of contrastive learning, sim(·) represents the calculation function, and the dot product is used here; t′ i is the normalized representation of the embedding e i of node i after being perturbed once to get e′ i ; t″ i is the normalized representation of the embedding e i of node i after being perturbed twice to get e″ i ; t″ j is the normalized representation of the embedding e j of node j after being perturbed twice to get e″ j , where and e′ i is the vector of node i after being perturbed once, and e″ i represents the vector of node i after being perturbed twice; Step 5: Contrastive learning and multi-task optimization; Performing graph convolution feature transformation on the first augmented view and the second augmented view respectively, and optimizing the model parameters in combination with the contrastive learning mechanism; Using the InfoNCE loss function as the objective function of contrastive learning; The final multi-task loss function combines the BPR loss, the regularization loss, the contrastive learning loss based on popularity masking data augmentation, and the contrastive learning loss of adaptive perturbation, and the formula is as follows: Among them is the task loss of the supervision signal module, is the regularization loss, is the contrastive learning loss of the data augmentation task based on the popularity mask, is the contrastive learning task loss of the adaptive perturbation data augmentation, and α and β are hyperparameters; Step 6: Model training and recommendation generation; Training the model parameters based on the gradient descent method to generate the final user and item embedding representations; Calculating the prediction scores according to the embedding representations of users and items, and generating a recommendation list.
2. The graph contrastive recommendation method based on dual adaptive data augmentation according to claim 1, wherein The specific method of the said step 1 is: converting the interaction records between users and projects into a sparse adjacency matrix A, where A ij = 1 indicates that there is an interaction between user i and project j, otherwise A ij = 0.
3. The graph contrastive recommendation method based on dual adaptive data augmentation according to claim 1, wherein The specific method of step 2 is: Initialize the embedding vector E of user u u and the embedding vector E of user v v , embedding users and items into a d-dimensional latent space; Aggregate the node neighborhood feature information using the graph convolution method based on GCN, where represents the user-item interaction matrix, I is the number of users, and J is the number of items; is the normalized adjacency matrix, and are the diagonal matrices, and their diagonal elements are the degrees of users and items respectively; GCN represents the graph convolutional neural network, D (u) represents the diagonal matrix of users, D (v) represents the diagonal matrix of items, is the adjacency matrix obtained by normalizing the user-item interaction matrix according to the diagonal matrix; In each layer of the GCN neural network, the embeddings of users and items are updated through neighborhood information; for users and items, their new embedding representations are respectively: and The superscript T represents the matrix transpose, that is is the transpose matrix of the matrix; where, z (u) and z (v) represent the information aggregated from neighboring items and users to the central node respectively; to suppress the over-smoothing effect, a residual connection is applied in the aggregation stage. The final user and item embeddings are obtained by summing the embeddings of each layer: and The subscript l represents the number of the item. The final user and item embeddings are obtained by summing the embeddings of all layers and then average pooling: E final are the final user and item embeddings; The BPR loss function is used as the core optimization objective of the supervision signal module, and the formula for the BPR loss is as follows: Among them, is the predicted score of user u for the positive sample item v pos ; is the predicted score of user u for the negative sample item v neg ; σ(x) is the Sigmoid function. The BPR loss function optimizes the preference ranking of users for items by maximizing the score difference between positive and negative samples.
4. The graph contrastive recommendation method based on dual adaptive data augmentation according to claim 1, wherein The specific method of step 3 is: Data augmentation contrastive learning based on node popularity mask generates an adaptive edge mask matrix by analyzing the degree distribution of nodes to adjust the connection strength of high-popularity nodes in the original adjacency matrix. First, calculate the order of each node. Specifically, for users and items, the formula for their degrees is as follows: where A ui represents the interaction between user u and item i in the user interaction matrix, d u and d i represent the degrees of user u and item i respectively, N represents the number of users, and M represents the number of items. Perform a logarithmic transformation on the degrees to smooth their distribution: d u ′ = log2(d u + 1)d i ′ = log2(d i + 1), where d′ u and d′ i are the degrees after transformation respectively. Subsequently, calculate the maximum and minimum values of the logarithm of the degrees in the global nodes of users and items respectively, which are used for normalization calculation in the subsequent adjacency matrix generation process: max ui = max(d u ′) + max(d i ′), min ui = min(d u ′) + min(d i ′) Among them, max ui and min ui represent the maximum and minimum values of the logarithm of degrees in the global nodes of users and projects; Generate an adaptive graph mask based on the node popularity distribution The graph mask matrix is used to dynamically adjust the user-item interaction graph; this mask is used in combination with the relationship matrix through element-wise multiplication: where is the new adjacency matrix generated after masking, represents the degree to which the interaction between the user and the item is masked. The closer the value is to 0, the lower the importance of the interaction; the closer the value is to 1, the higher the importance of the interaction; ⊙ represents element-wise multiplication; Each element in the mask matrix is calculated as follows: Pair is normalized to obtain a normalized adjacency matrix represents the matrix of the i-th row and j-th column The element of is specifically defined as: Among them, represents the element in the i-th row and j-th column of the new adjacency matrix after masking. Based on the normalized adjacency matrix, the message passing process is defined as follows: Among them, and respectively represent the updated embeddings of the user and the project, represents the normalized adjacency matrix 5. The graph contrastive recommendation method based on dual adaptive data augmentation according to claim 1, wherein The specific method of step 4 is: Let where, \(E^{ l}\) l represents the embedding matrix of the \(l\)-th layer, \(L\) is the number of layers of the graph convolutional network, and \(E\) is the final embedding representation; a multi-layer perceptron MLP is used to learn the mean adaptation of user and item embeddings to obtain a perturbation vector, and the calculation formula of the perturbation vector is as follows: \(p=(e^{ i}\cdot W + b)\odot sign(e^{ i})\), where, \(p\) represents the perturbation vector adaptively generated by MLP according to the features of each node itself, \(W\) is the learned weight, \(b\) is the bias term, \(\odot\) represents element-wise multiplication, \(sign(e^{ i})\) is the sign function, which is used to retain the direction information of the original embedding, and \(e^{ i}\) i ·W + b)\odot sign(e^{ i}) i ), where, \(p\) represents the perturbation vector adaptively generated by MLP according to the features of each node itself, \(W\) is the learned weight, \(b\) is the bias term, \(\odot\) represents element-wise multiplication, \(sign(e^{ i})\) is the sign function, which is used to retain the direction information of the original embedding, and \(e^{ i}\) i ) is the sign function, which is used to retain the direction information of the original embedding, and \(e^{ i}\) i represents the node embedding; Given a node i and its representation e i , by applying an adaptive perturbation to it, two new representations are generated: t′ i = e i + p′ i · ε, t″ i = e i + p″ i · ε, where t′ i , t″ i are the node embedding representations in the data augmentation views generated by two independent perturbations respectively, is a small constant used to control the intensity of the perturbation; Performing contrastive learning training on the embeddings of the same node in different views in the two independently generated perturbation views; Maximizing the similarity of the embeddings from the same category, while minimizing the similarity between the embeddings from different categories.
6. The graph contrastive recommendation method based on dual adaptive data augmentation according to claim 1, characterized in that, In the said step 5 where W i represents the parameters of the i-th layer in the model, represents the L2 norm of the parameters, and λ is the hyperparameter of regularization used to adjust the strength of regularization.
7. A graph contrastive recommendation method based on dual adaptive data augmentation according to claim 1, characterized in that After the training in step 6 is completed, the model will use the final embedding vectors for prediction to generate the predicted ratings of users for items. During prediction, assume that the embedding vector of user u is u and the embedding vector of item v is v. For user u and item v, their predicted ratings are calculated through the inner product: where is the predicted rating of user u for item v. After calculating the predicted ratings of each user for all items, they are sorted from high to low according to the ratings to generate the recommendation list for each user.
Citation Information
Cited By
Unsupervised recommendation system anomaly detection method based on multi-scale behavior modeling and boundary perception enhancement
CN121188669A