A Self-Supervised Graph Neural Network-Based Recommendation Method and System
By adopting data enhancement methods based on project popularity and covariant invariant loss function methods in the recommendation system, the problems of inappropriate data enhancement methods, negative sampling and joint learning strategies in the prior art are solved, and more efficient model training and performance improvement are achieved.
Patent Information
- Application Number
- CN202210969801.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-12
AI Technical Summary
The existing self-supervised graph neural network recommendation algorithm has problems with domain-independent data enhancement methods, random negative sampling and joint learning strategies, resulting in data loss, model performance degradation and training difficulty.
Data augmentation is performed using the popular bias reduction method based on project popularity. The popularity and discard probability of each edge are calculated to generate enhanced data from two perspectives, and a loss function including covariant and invariant parts is designed for model training to avoid random negative sampling and joint learning.
It effectively alleviates the problem of popularity deviation in the recommendation system, retains the internal structure of the original graph data, improves the efficiency and performance of model training, and avoids the increase in complexity caused by the collapse of the model and joint learning.
Smart Images

Figure CN115525836B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of recommendation systems, and in particular to a self-supervised graph neural network recommendation method and system. Background Art
[0002] With the rapid development of the Internet, information has exploded. How to filter massive amounts of information has become one of the issues that many researchers are concerned about. Recommendation systems can effectively filter raw information, thereby generating personalized recommendation results for each user and alleviating the problem of information overload. Currently, they are widely used in e-commerce, social networks, smart healthcare, and online education.
[0003] In recent years, the recommendation system based on graph neural network has gradually become a hot topic in the industry. Compared with the classic collaborative filtering algorithm, graph neural network can capture the high-order connection between users and items, so as to improve the upper limit of model performance by using deep collaborative filtering signals during training. At the same time, in order to alleviate the cold start problem of the recommendation system and mine more supervision information from the original data itself, some scholars have proposed to introduce the self-supervised training mode into the recommendation system based on graph. As a branch of unsupervised algorithm, the self-supervised algorithm originated from the field of image processing. Unlike the training of image classification tasks using manually annotated image labels, the typical approach of the self-supervised model is to first generate enhanced images of two perspectives from one image through data enhancement, and then generate two vectors through a parameter-sharing encoder respectively, and use the contrast learning loss function for training. The essence of this loss function is to let the vectors generated from the two perspectives supervise each other. Because its supervision signal comes from the data itself and does not rely on manual annotation, it is named self-supervision. The existing work that combines self-supervised algorithms with graph neural networks for recommendation algorithms mainly involves randomly enhancing the interactive bipartite graph of users and items, and then using the traditional contrast loss function in the image field for training, but it is not optimized based on the background of the recommendation system itself. There are mainly the following problems:
[0004] (1) Domain-independent data augmentation methods. Current self-supervised graph recommendation algorithms usually perform data augmentation based on random edge dropping or random feature masks. This domain-independent data augmentation method has been proven to destroy the internal structure of the original graph data, causing data information loss and generating coupled vector representations.
[0005] (2) Random negative sampling. Contrastive learning loss functions usually adopt a learning strategy of pulling away negative samples to prevent the collapse of the network model, which requires a large amount of high-quality negative sample data. However, existing models usually randomly sample all samples to obtain negative samples, or regard all samples except the target samples in a batch as negative samples. Both of the above methods will mislead the optimization direction of the loss function, thereby reducing the performance of the algorithm.
[0006] (3) Joint learning strategy. Existing methods usually combine self-supervised tasks with the main task of rating prediction in the recommendation system for model learning. Essentially, the self-supervised task is used as an auxiliary task to restrict the distribution of vectors. This joint learning method will increase the complexity of the algorithm, resulting in an increase in the difficulty of model training. Summary of the Invention
[0007] The object of the present invention is to overcome the above-mentioned deficiencies existing in the prior art and provide a self-supervised graph neural network recommendation method and system.
[0008] To achieve the above-mentioned invention object, the present invention provides the following technical solutions:
[0009] A self-supervised graph neural network recommendation method includes the following steps:
[0010] S1. Obtain user and item interaction data, including user data, item data, and user-item interaction record data, to form a bipartite graph G = (X, A), where X is the node attribute matrix, X ∈ R (n+m)×d , A is the adjacency matrix, A ∈ R (n+m)×(n+m) , n represents the number of user nodes, m represents the number of item nodes, and d represents the vector dimension;
[0011] S2. Use the popularity bias reduction method based on item popularity to enhance the data G = (X, A) to generate enhanced data G' = (X, A') and G'' = (X, A'') from two perspectives;
[0012] S3. Encode the enhanced data G' = (X, A') and G'' = (X, A''), and generate vectors Z' and Z'' respectively;
[0013] S4. Use a loss function to train the vectors Z' and Z'', and obtain a graph neural network recommendation model. The loss function includes a covariant part and an invariant part, and the formula of the loss function is as follows:
[0014] L(Z', Z'') = λs(R', R'') + μ[c(Z') + c(Z'')]
[0015] Where λ and μ are hyperparameters, which respectively control the weights of the invariant part loss function and the covariant part loss function, s(R', R'') is the invariant part loss function, and c(Z') and c(Z'') are the covariant part loss functions of the vectors Z' and Z'' respectively;
[0016] S5. Based on the graph neural network recommendation model, push item information of interest to users.
[0017] In the technical solution proposed by the present invention, a graph neural network recommendation model is first trained, and then project information is recommended to users according to the graph neural network recommendation model. The training of the graph neural network recommendation model is to collect original data, and use a data enhancement method related to the field to enhance the original data into data from two perspectives, so that the enhanced data can retain the information required for the recommendation task to the greatest extent. Then, the augmented data is encoded, and an iterative training is performed using a loss function to obtain the graph neural network recommendation model. The loss function can effectively solve the problem of random negative sampling and avoid the joint training of the model.
[0018] Further, step S2 specifically includes the following steps:
[0019] S21. Calculate the popularity of each edge in the bipartite graph G=(X,A). The popularity of each edge is the popularity of the item node it is connected to, that is where S u,i is the popularity of the edge connecting the user node u and the item node i in the bipartite graph, represents the popularity of the item node i;
[0020] S22. Calculate the discard probability of each edge based on the edge popularity. The discard probability calculation formula is as follows:
[0021]
[0022] In the formula, p u,i is the discard probability of the edge, S u,i -s min / s min -s max is the normalization of s u,i , s min and s max are the minimum and maximum values of the popularity of all edges respectively, p e is the global discard probability of each edge defined by the user, p τ is a hyperparameter;
[0023] S23. Based on p u,i augment the data G=(X,A) to obtain the augmented data G'=(X,A') and G''=(X,A''), A' = Mask·A, R'' = Mask·A, where Mask is a mask matrix, Mask∈R (n+m)×(n+m) , and each item m u,i in Mask follows a Bernoulli distribution with a probability of p u,i .
[0024] The data enhancement method of the present invention can dynamically update the discard probability of the edges in the bipartite graph according to the item popularity, so as to enhance the data.
[0025] Furthermore, in step S3, the enhanced data G′ = (X, A′) and G″ = (X, A″) are encoded using the encoder LightGCN to generate the vectors Z′ and Z″ of the users and items of the two perspective-enhanced data, as follows:
[0026] Z′ = LightGCN(X, A), Z″ = LightGCN(X, A″)
[0027] where Z′ = [z′ u1 ,..., z′ un , z′ i1 ,..., z′ im and Z″ = [z″ u1 ,..., z″ un , z″ i1 ,..., z″ im are the sets of vectors of users and items in the two perspectives generated by the encoder, respectively.
[0028] Furthermore, the covariant part loss function is composed of the vectors Z′ and Z″, and the calculation formulas for c(Z′) and c(Z″) are as follows:
[0029]
[0030]
[0031] where d is the dimension of the vector X′ or Z″, C(Z′) is the covariance matrix of the vector Z′, C(Z″) is the covariance matrix of the vector Z″, and the covariance matrix calculation formula is as follows:
[0032]
[0033]
[0034] where, and correspond to the averages of Z′ and Z″ respectively, z′ k is the term of the vector Z′, and z″ k is the term of the vector Z″.
[0035] Through the invariant loss function, the off-diagonal elements in the vector covariance matrix will approach 0 during model training, thereby maintaining the independence between the dimensions of the vector and preventing the collapse of the model.
[0036] Even further, the invariant part loss function is calculated based on the predicted scores, and the invariant part loss function formula is as follows:
[0037]
[0038] where r′k,j is the item of the predicted score R′ corresponding to the vector Z′, R′=Z′ u (Z′ i ) T , r″ k,j is the item of the predicted score R″ corresponding to the vector Z″, R″=Z″ u (Z″ i ) T , R′∈R N×N , R″∈R N×N , N is the size of each batch, that is, the number of training each time, Z′ u and Z′ i Separate the vector Z′ into user vector and item vector, Z′=Z′ u ||Z′ i , Z″ u and Z″ i Separate the vector Z″ into a user vector and an item vector, Z″=Z″ u ||Z″ i The invariant part loss function is based on the idea of contrastive learning, which brings the predicted scores under two perspectives closer and obtains the loss function based on the invariant part of the predicted scores.
[0039] The present invention also provides a self-supervised graph neural network recommendation system, which is used to implement the above-mentioned self-supervised graph neural network recommendation method, and the recommendation system includes:
[0040] Data collection module, used to obtain user and project interaction data;
[0041] A data processing module, used to construct a bipartite graph G = (X, A) from user and project interaction data;
[0042] A data augmentation module is used to enhance the data G = (X, A) to generate enhanced data G′ = (X, A′) and G″ = (X, A″) of two perspectives by using a popular deviation reduction method based on item popularity;
[0043] A data encoding module, used to encode the enhanced data and generate vectors Z′ and Z″ respectively;
[0044] A model training module is used to train vectors Z′ and Z″ using a loss function to obtain a graph neural network recommendation model;
[0045] The project recommendation module is used to push project information of interest to users based on the graph neural network recommendation model.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. The recommendation method of the present invention starts from the popularity bias problem that has long existed in the recommendation system. It uses a domain-related data augmentation method to augment the original data into two perspectives of data. Specifically, it performs augmentation operations on the input data based on the natural popularity indicators in the original input data, so that the augmented data can retain the information required for the recommendation task to the greatest extent. And through experiments, the combined effects of different data augmentation methods are verified, and the optimal data augmentation combination is selected, so as to improve the upper limit of the training of the graph neural network recommendation model, retain the internal structure of the original graph data from the input end, and will not cause data information loss. Different from the random data augmentation method, the popularity bias reduction method can alleviate the popularity bias problem in the recommendation system during the data augmentation stage, thereby improving the performance of the recommendation system.
[0048] 2. The recommendation method of the present invention aims to unify the self-supervised task and the main task of the recommendation system, and designs a new loss function. The loss function includes an invariant part and a covariant part. In the invariant part, the definition of the positive sample is modified from the vector level to the score level, so as to combine the self-supervised task with the scoring task in the recommendation system, enabling the training of the graph neural network recommendation model to break out of the joint learning mode, avoiding the loss of vector information in the training of different tasks, simplifying the model structure, reducing the model complexity, improving the training efficiency of the model, and realizing the end-to-end unity of the model training; in the covariant part, by constraining the covariance matrix of the vectors, the non-diagonal elements of the matrix approach 0, so as to ensure the independence between the dimensions of the vectors and prevent the collapse of the model. Based on the loss function proposed by the present invention, the random negative sampling problem can be effectively solved, and the joint training of the model can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is the flowchart of the self-supervised graph neural network recommendation method of the present invention;
[0050] Figure 2 is the basic framework of the self-supervised graph neural network recommendation model in Embodiment 1;
[0051] Figure 3 is the flowchart of the popularity bias reduction method based on item popularity in Embodiment 1;
[0052] Figure 4 is the schematic diagram of using different data augmentation methods in Embodiment 1;
[0053] Figure 5 is the performance result graph of the recommendation method under different data augmentation methods in Embodiment 1. DETAILED DESCRIPTION OF THE INVENTION
[0054] The present invention will be further described in detail below in conjunction with test examples and specific embodiments. However, it should not be understood that the scope of the above-mentioned subject matter of the present invention is limited to the following embodiments, and all technologies implemented based on the content of the present invention belong to the scope of the present invention.
[0055] Embodiment 1
[0056] This embodiment provides a self-supervised graph neural network-based recommendation method, as Figure 1 shown, including the following steps:
[0057] S1. Obtain user and item interaction data, including user data, item data, and user-item interaction record data, to form a bipartite graph G=(X, A), where X is the node attribute matrix, X∈R (n+m)×d , and A is the adjacency matrix, A∈R (n+m)×(n+m) ;
[0058] X represents the set of attribute values of all nodes in the graph. The historical interactions between users and items can naturally form a bipartite graph G=(V, ε), where V is the node, V = U∪I, consisting of all user nodes U and item nodes I, and ε is the edge, representing the historical interactions between users and items, ε∈O + . For the adjacency matrix A, if nodes i and j have had historical interactions, that is, (v i , v j )∈ε, then A i,j = 1, otherwise A i,j = 0, n represents the number of user nodes, m represents the number of item nodes, and d represents the vector dimension. Table 1 below shows the data to be analyzed in this embodiment. In the three data sets obtained in this embodiment, the items are goods.
[0059] Table 1 Details of the public data set
[0060]
[0061]
[0062] S2. Use the popularity bias reduction method based on item popularity to enhance the data G=(X, A) to generate enhanced data G'=(X, A') and G''=(X, A'') from two perspectives, as Figure 2 shown;
[0063] Recommendation systems have been proven to have a serious popularity bias problem, that is, recommendation systems tend to recommend high-popularity items to users, thus ignoring a large number of unpopular items. Over time, this will exacerbate the homogeneity of the recommended items, resulting in information cocoons and long-tail problems. Therefore, a new data augmentation method, the popularity bias reduction method, is proposed. This data augmentation method enhances the original data based on the popularity of the items, reducing the popularity difference of the items to be recommended during the data augmentation stage, and can effectively alleviate the popularity bias problem in the recommendation system.
[0064] Specifically, for the popularity bias reduction method, first, the popularity of each edge in the bipartite graph is defined based on the historical interaction frequencies of each user and item, and the discard probability of each edge in the bipartite graph is dynamically updated according to the popularity of the edge, so that edges with higher popularity have a greater probability of being discarded. This method will balance the popularity differences of the nodes in the newly generated enhanced graph, thus alleviating the popularity bias of the recommendation system. The flowchart is as Figure 3 shown, and step S2 specifically includes the following steps:
[0065] S21. Calculate the popularity of each edge in the bipartite graph G=(X,A). The popularity of each edge is the popularity of the item node it connects, that is where s u,i is the popularity of the edge connecting user node u and item node i in the bipartite graph, represents the popularity of item node i;
[0066] For each item, its popularity (popularity level) is determined by the historical interaction frequency of the item. In the bipartite graph, the degree of a node is defined as the number of edges connected to the node. Therefore, in the bipartite graph composed of users and items, the popularity level of the item can be defined by the degree of each item node. Define the popularity of each edge in the bipartite graph as the popularity of the item node it connects, that is where s u,i is the popularity of the edge connecting user node u and item node i in the bipartite graph, represents the popularity of item node i. For example Figure 1 in, when calculating the popularity of the edge between user node u1 and item node i1, u1 has 1 edge connected to it, and the degree of u1 is 1, then the popularity of the edge between user node u1 and item node i1 is 1.
[0067] S22. Calculate the discard probability of each edge based on the popularity of the edge. The discard probability calculation formula is as follows:
[0068]
[0069] In the formula, p u,i is the discard probability of the edge, su,i -s min / s min -s max is to normalize s u,i where s min and s max are the minimum and maximum values of the popularity of all edges respectively, and pe is the global discard probability of each edge defined by the user; the global discard probability will be dynamically updated based on the normalized s u,i to ensure that items with high popularity have a greater chance of being discarded, reducing the popularity difference between nodes in the enhanced graph; p τ is a hyperparameter, and p τ is to prevent the discard probability from being too large and thus destroying the information of the original bipartite graph, where p τ < 1.
[0070] S23. Based on p u,i the data G = (X, A) is enhanced to obtain enhanced data G′ = (X, A′) and G″ = (X, A″), where A′ = Mask·A and A″ = Mask·A, and Mask is a mask matrix, Mask ∈ R (n+m)×(n+m) and each item m u,i in Mask follows a Bernoulli distribution with a probability of p u,i .
[0071] Different from the domain - independent random enhancement method in the existing algorithm, PBR starts from the background of the recommendation system, aims to solve the popularity bias problem of the recommendation system, and dynamically updates the discard probability of edges based on the popularity of each item. This makes the popularity difference between nodes in the bipartite graph generated after data enhancement smaller. While generating two different bipartite graphs for contrastive learning, it maximally retains the key information affecting the performance of the recommendation system in the bipartite graph. It alleviates the popularity bias problem of the recommendation system at the data input end, thereby improving the performance of the recommendation system.
[0072] S3. Encode the enhanced data G′ = (X, A′) and G″ = (X, A″) to generate vectors Z′ and Z″ respectively.
[0073] In this embodiment, a weight - shared graph neural network - based encoder LightGCN is used to encode the enhanced data G′ = (X, A′) and G″ = (X, A″). LightGCN is a graph neural network encoder for the recommendation system, which simplifies the feature transformation and non - linear activation modules in the original graph neural network, making it a lighter and better - performing encoder suitable for the recommendation system. Thus, vectors Z′ and Z″ of users and items of the two - perspective enhanced data are generated respectively, as follows:
[0074] Z′ = LightGCN(X, A′), Z″ = LightGCN(X, A″)
[0075] where Z′ = [z′ u1 ,..., z′ un , z′ i1 ,..., z′ im and Z″ = [z″ u1 ,..., z″ un , z″ i1 ,..., z″ im are the sets of user and item vectors from two perspectives generated by the encoder respectively.
[0076] S4. Use the loss function to train the vectors Z′ and Z″ to obtain the graph neural network recommendation model. The loss function includes a covariant part and an invariant part. The formula of the loss function is as follows:
[0077] L(Z′, Z″) = λs(R′, R″) + μ[c(Z′) + c(Z″)]
[0078] where λ and μ are hyperparameters that control the weights of the invariant part loss function and the covariant part loss function respectively, s(R′, R″) is the invariant part loss function, and c(Z′) and c(Z″) are the covariant part loss functions of the vectors Z′ and Z″ respectively.
[0079] s(R′, R″) is the invariant part loss function, which is calculated based on the predicted scores. The formula of the invariant part loss function is as follows:
[0080]
[0081] In the formula, r′ k,j is the term of the predicted score R′ corresponding to the vector Z′, R′ = Z′ u (Z′ i ) T , r″ k,j is the term of the predicted score R″ corresponding to the vector Z″, R″ = Z″ u (Z″ i ) T , R′ ∈ R N×N , R″ ∈ R N×N , N is the size of each batch, that is, the number of training each time, Z′ u and Z′ i are the vector Z′ separated into user vector and item vector, Z′ = Z′ u ||Z′ i , Z″ u and Z″ i are the vector Z″ separated into user vector and item vector, Z″ = Z″u ||Z″ i . The separation of vectors Z′ and Z″ into user and item vectors is based on the indexes of users and items in the original graph, and vectors Z′ and Z″ are separated. The invariant part loss function is based on the idea of contrastive learning, which brings the predicted scores under two perspectives closer and obtains a loss function based on the invariant part of the predicted score.
[0082] The covariant part loss function is composed of vector Z′ and vector Z″. The calculation formulas of c(Z′) and c(Z″) are as follows:
[0083]
[0084]
[0085] Where d is the dimension of vector Z′ or Z″, C(Z′) is the covariance matrix of vector Z′, and C(Z″) is the covariance matrix of vector Z″. The formula for calculating the covariance matrix is as follows:
[0086]
[0087]
[0088] In the formula, and Corresponding to the average values of Z′ and Z″, z′ k is the term of vector Z′, z″ k is the term of vector Z″. Through the covariant partial loss function, the off-diagonal elements in the vector covariance matrix will approach 0 during model training, thereby maintaining the independence between the dimensions of the vector and preventing the collapse of the model. The vectors Z′ and Z″ are iteratively optimized using the loss function to realize the training of the graph neural network recommendation model.
[0089] Existing methods train models based on traditional contrastive learning loss functions. Contrastive learning considers different vectors of the same node from two perspectives as positive samples, and closes positive samples through constraints in the loss function. Considering that simply closing positive samples will cause the collapse of the model, contrastive learning will introduce a large number of negative samples and separate the distance between positive and negative samples. However, in existing contrastive learning, all samples except the target sample in a batch are often considered negative samples. This definition method will increase the distance between many essentially similar samples in the vector space, reducing the performance of the recommendation system. In addition, existing self-supervised recommendation algorithms usually regard the self-supervised model as an auxiliary task to constrain the spatial distribution of vectors, and ultimately need to be trained in conjunction with the main task of rating prediction of the recommendation system. This joint learning model increases the complexity of the loss function and increases the burden of model training.
[0090] Different from defining vectors Z' and Z'' as positive samples in traditional contrastive learning loss functions, in this application, by calculating the predicted scores of users for items, the score matrices from two perspectives are directly regarded as positive samples. We call this kind of positive sample task-related positive samples, which is also the origin of task oriented in the loss function. Based on the definition of this positive sample, we directly connect the self-supervised task with the main task of score prediction in the recommendation system. Therefore, no additional joint learning strategy is needed. By transferring the definition object of the positive sample from the vector level to the score level, the self-supervised task and the main task of the recommendation system are unified, enabling the graph neural network recommendation model to train accurate user and item vectors through a single loss function, improving the efficiency of model training and reducing the hardware burden. Additionally, by calculating the covariance matrix of vectors and making the distance between this covariance matrix and the identity matrix closer in the loss function, the independence between each dimension of the vectors is well maintained by constraining the vector covariance matrix, avoiding the collapse of the model, thus avoiding the uncertainty of random sampling and improving the performance of the recommendation model.
[0091] S5. Push information of items that the user is interested in based on the graph neural network recommendation model.
[0092] To illustrate the effects of the popularity bias reduction method and the loss function proposed in the present invention, ablation experiments were respectively conducted for the popularity bias reduction method and the loss function. For the popularity bias reduction method, on the basis of keeping other parts of the model unchanged, by using four typical data augmentation methods to replace the popularity bias reduction method to verify the role of the popularity bias reduction method in the present invention. These four data augmentation methods are as Figure 3 shown, specifically as follows
[0093] (a) Node-dropping (N): Randomly drop a certain proportion of nodes in the bipartite graph. In Figure 4 (a), the dotted circles represent the randomly dropped nodes.
[0094] (b) Edge dropping (E): Randomly drop a certain proportion of edges in the bipartite graph. In Figure 4 (b), the dotted lines represent the randomly dropped edges.
[0095] (c) Feature masking (F): Randomly cover some dimensions of the user and item feature vectors. In Figure 4 (c), the "×" represents the covered dimensions.
[0096] (d) Adaptive (A): Drop a certain proportion of edges in the bipartite graph based on the degrees of nodes. In Figure 4In (d), the dashed straight lines represent the edges discarded based on the degrees of the nodes.
[0097] Applying the above four data augmentation methods to the ML-100k dataset for the recommendation system, the experimental results are as Figure 5 shown in (a). The recall rate of the popularity bias reduction method (P) proposed by the present invention is higher than the other four, which can prove the effectiveness of the popularity bias reduction method. In addition, different data augmentation methods are combined, and the experimental results are as Figure 5 shown in (b). The optimal combination scheme is P&P. Secondly, the combination of PBR and other data also achieves good results, once again proving the superiority of the popularity bias reduction method proposed by the present invention.
[0098] For the loss function, while keeping other parts of the model unchanged, the present invention and three other loss functions are applied to the ML-100k dataset for the recommendation system. The three other loss functions are InfoNCE, Barlow Twins, and ICT-variants. ICT-variants is a variant method of the loss function (ICT), that is, keeping the covariant part unchanged and modifying the definition of positive samples in the invariant part to the vector level. The experimental results are shown in Table 2. From the data in the table, it can be seen that whether jointly optimizing the self-supervised task and the main task of the recommendation system or only optimizing the self-supervised task, the loss function is superior to other loss functions; in the case of only training the self-supervised task, it is far superior to other loss functions, even the loss function variants. This result confirms that the modification of the definition of positive samples in the loss function of the present invention helps to unify the scoring tasks of the self-supervised task and the recommendation system. This is mainly because in this experiment, the self-supervised task is changed from the embedding level to the scoring level, so the model performance does not need to rely on the scoring task of the recommendation system.
[0099] Table 2 Performance of recommendation algorithms under different loss functions
[0100] Self-supervised task Joint learning Loss function Recall Recall InfoNCE 0.1145 0.2014 BT 0.0955 0.1954 ICT-variants 0.1254 0.2145 Loss function 0.2226 0.2237
[0101] To further verify the effect of the present invention, multiple groups of comparative experiments are carried out on three public datasets using the present invention and five other methods. The five methods are as follows:
[0102] NGCF: A recommendation system based on graph neural networks;
[0103] LightGCN: Based on NGCF, two network architectures that are not suitable for the recommendation system, namely non-linear transformation and activation function, are removed, and the efficiency of the recommendation algorithm is improved on the basis of simplifying the model;
[0104] SGL: Introduced self-supervised algorithms into the recommendation system and adopted three random augmentation methods for data augmentation (SGL-ND, SGL-ED, SGL-RW);
[0105] SelfCF: Introduced an asymmetric network into the recommendation system, improving the efficiency of model training;
[0106] GCA-DE: Dynamically updated the probability of discarding edges according to the node degrees in the graph, avoiding excessive damage to the original graph caused by data augmentation.
[0107] The parameters of the comparative experiments are shown in Table 3, and the results of multiple comparative experiments on three datasets are shown in Table 4.
[0108] Table 3 Parameter settings of comparative experiments
[0109]
[0110]
[0111] Table 4 Comparative chart of recommendation performance of different models
[0112]
[0113] As shown in Table 4, compared with the existing baseline models, the method proposed in the present invention has made certain progress in two indicators: recall and normalized discounted cumulative gain (ndcg). On the one hand, relying on the existence of the data augmentation mechanism, unsupervised training can help the model expand the range of input data, thereby increasing the number of model training times and improving model performance. On the other hand, the PBR data augmentation method and the ICT loss function proposed in the present invention are both designed specifically for the characteristics of the recommendation system. PBR can alleviate the popularity bias problem existing in the recommendation system at the data input end, while ICT can combine the self-supervised task and the scoring main task of the recommendation system, thereby reducing the training difficulty of the model and improving the training efficiency of the model.
[0114] Embodiment 2
[0115] This embodiment provides a self-supervised graph neural network recommendation system that implements the recommendation method of Embodiment 1. The system includes:
[0116] A data collection module for obtaining user and item interaction data;
[0117] A data processing module for constructing a bipartite graph G=(X, A) from user and item interaction data;
[0118] A data augmentation module, which is used to generate augmented data G'=(X, A') and G''=(X, A'') from two perspectives by enhancing the data G=(X, A) through a popularity deviation reduction method based on item popularity;
[0119] A data encoding module, which is used to encode the augmented data to generate vectors Z' and Z'' respectively;
[0120] A model training module, which is used to train the vectors Z' and Z'' using a loss function to obtain a graph neural network recommendation model;
[0121] An item recommendation module, which is used to push item information of interest to users based on the graph neural network recommendation model.
[0122] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A self-supervised graph neural network-based recommendation method, characterized in that It includes the following steps: S1. Obtain user and project interaction data, including user data, project data, and user-project interaction record data, to form a bipartite graph , is the node attribute matrix, , is the adjacency matrix, , represents the number of user nodes, represents the number of project nodes, represents the vector dimension; S2. Augment the data by the popularity deviation reduction method based on project popularity to generate augmented data from two perspectives and , which specifically includes the following steps: S21. Calculate the popularity of each edge in the bipartite graph The popularity of each edge is the popularity of the item node it connects, that is , where is the popularity of the edge connecting the user node and the item node in the bipartite graph, represents the popularity of the item node ; S22. Calculate the discard probability of each edge based on the popularity of the edge. The discard probability calculation formula is as follows: where is the edge discard probability, is to perform normalization on and are the minimum and maximum values of all edge popularities respectively, is the globally defined discard probability for each edge, is a hyperparameter; S23. Based on the data is enhanced to obtain enhanced data and , , , where is a mask matrix, , each item in obeys a Bernoulli distribution with a probability of ; S3. Encode the enhanced data and to generate vectors and respectively; S4. Use a loss function for the vectors and to perform training and obtain a graph neural network recommendation model. The loss function includes a covariant part and an invariant part, and the formula of the loss function is as follows: where and are hyperparameters, is the invariant part loss function, are the covariant part loss functions of vectors and vector respectively; S5. Push information of items that the user is interested in based on the graph neural network recommendation model.
2. The self-supervised graph neural network-based recommendation method according to claim 1, wherein In step S3, the enhanced data and are encoded using the encoder LightGCN to generate the vectors of users and items for the two-view enhanced data and , as follows: wherein and are respectively the sets of user and item vectors from two perspectives generated by the encoder.
3. The self-supervised graph neural network-based recommendation method according to claim 1, wherein The covariant partial loss function is a vector and a vector which are composed of and The calculation formulas of where is the dimension of the vector or the vector . is the covariance matrix of the vector , and is the covariance matrix of the vector . The calculation formula of the covariance matrix is as follows: wherein, and correspond to the average values of and respectively, is the term of the vector , is the term of the vector .
4. The self-supervised graph neural network-based recommendation method according to claim 3, wherein The invariant part loss function is calculated based on the predicted score. The invariant part loss function formula is as follows: where is a vector corresponding predicted score of the term, , is a vector corresponding predicted score of the term, , , , is the size of each batch, and are vectors separated into user vectors and item vectors, , and are vectors separated into user vectors and item vectors, .
5. A self-supervised graph neural network recommendation system for implementing the self-supervised graph neural network recommendation method according to any one of claims 1-4, characterized in that, The recommendation system includes: A data acquisition module for obtaining user and item interaction data; A data processing module for constructing a bipartite graph from user and project interaction data ; A data augmentation module for augmenting data by generating augmented data of two perspectives through a popularity deviation reduction method based on item popularity and ; A data encoding module, which is used to encode the enhanced data and generate vectors and ; A model training module for training a vector using a loss function and to obtain a graph neural network recommendation model; An item recommendation module for pushing information of items that the user is interested in based on the graph neural network recommendation model.
Citation Information
Patent Citations
MOOC recommendation method based on graph convolutional neural network
CN114154070A
Recommendation method and system based on adaptive noise reduction training, electronic equipment and medium
CN114254187A