Recommendation method for lightweight data enhancement graph contrast learning
New views are constructed by approximate singular value decomposition and embedded spatial structure correlation, combined with multi-layer graph convolution network and joint training strategy, the noise amplification and training cost problems in graph comparison learning are solved, and the accuracy and robustness of the recommended model are improved.
Patent Information
- Application Number
- CN202510580506.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The noise signal amplification effect and training time cost problems in the existing graph comparison learning method lead to insufficient accuracy and robustness of the recommended model.
Approximate singular value decomposition and embedded spatial structure correlation are used to enhance data, build a new view, and learn representations through multi-layer graph convolution networks, combining the joint training strategy of main task and contrast learning to share graph convolution parameters.
It effectively alleviates the noise problem, reduces training time, and improves the accuracy and robustness of the model.
Smart Images

Figure CN120492729A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of recommendation systems, and in particular relates to a recommendation method based on lightweight data-enhanced graph contrast learning. Background Art
[0002] With the exponential growth of information overload on the internet, recommendation systems have become a key technology for addressing user personalization needs. Current mainstream approaches have shifted from traditional matrix factorization collaborative filtering to architectural innovations based on graph neural networks (GNNs). By constructing a user-item interaction graph and leveraging a multi-layer message passing mechanism to aggregate neighborhood features, this theoretically overcomes the limitations of traditional methods in modeling first-order relationships. Graph collaborative filtering models, such as LightGCN, extend the coverage radius of user preference representation to higher-order relationships beyond three hops through learnable neighborhood sampling weights and feature propagation rules, demonstrating significant accuracy improvements in benchmark tests.
[0003] To further enhance model accuracy and robustness, graph contrastive learning (GCL) technology has been introduced into the recommendation field, forming a new paradigm of "feature enhancement-view contrast." Typical approaches include the SGL model, which uses a random edge deletion strategy to construct contrast views, generating differentiated graph structures by randomly removing interacting edges; and SimGCL, which perturbs features by adding uniformly distributed noise to the embedding space. These methods further enhance recommendation accuracy beyond the existing GNN model by maximizing the consistency of front view representations.
[0004] However, existing data augmentation strategies have two key problems: one is the amplification effect of noise signals. Since user behavior data in real-world scenarios naturally contain noise interactions, traditional random augmentation operations will produce a coupling effect with the inherent noise: during edge deletion, key preference edges may be mistakenly removed, while noise edges are retained; during feature perturbations, the direction of adding uniform noise deviates from the true semantic space. More seriously, the multi-hop propagation characteristics of graph neural networks will amplify the initial noise exponentially; the second is the problem of training time and cost. The addition of graph contrastive learning to the network architecture has increased the time for data training, and the existing contrastive learning framework needs to regenerate the augmented view in each training cycle (epoch). Therefore, the existing graph contrastive learning increases the training cost several times compared to traditional graph neural networks. Summary of the Invention
[0005] The present invention aims to address these issues by proposing a lightweight data-augmented recommendation method based on graph contrastive learning. This method employs both approximate singular value decomposition and embedded spatial structure correlation for data augmentation, effectively mitigating the noise problem. Furthermore, the data augmentation strategy only needs to be executed once, significantly reducing time overhead and ultimately improving the accuracy and robustness of recommendations.
[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0007] S1. Data preprocessing: Using the MovieLens 10M dataset, extract key information and construct an adjacency matrix;
[0008] S2, lightweight data enhancement: use approximate singular value decomposition and embedding spatial structure correlation to perform data enhancement and construct two new views;
[0009] S3. Embedded representation learning: Using multi-layer graph convolutional networks to learn representations of different views;
[0010] S4. Model training: Use a joint training strategy of the main task and contrastive learning to share graph convolution parameters.
[0011] Further as a preferred technical solution of the present invention, step S1 includes the following steps:
[0012] S101. Collect 10M of open-source MovieLens data, clean it, retain the user ID, movie ID, and rating data, and then divide it into training, validation, and test sets according to the ratio of 8:1:1.
[0013] S102, construct the user ID, movie ID and rating data in the training set data into an adjacency matrix A, the user set U={u1,u2,...,u I}, (|U|=I), movie set V={v1,v2,...,v J},(|V|=J), A i,j The value of is the rating data, or 0 if there is no rating.
[0014] Further as a preferred technical solution of the present invention, step S2 includes the following steps:
[0015] S201, normalize the adjacency matrix A:
[0016]
[0017] in, is the user degree matrix, whose diagonal elements represent the number of interactions of each user, is the movie degree matrix, whose diagonal elements represent the number of interactions for each movie;
[0018] S202. Perform approximate singular value decomposition on the normalized adjacency matrix A:
[0019]
[0020] Where q is the number of singular values retained (a hyperparameter that controls the accuracy of the low-rank approximation), is the user's low-dimensional latent feature vector (each row represents a user's q-dimensional feature), is the singular value matrix (diagonal elements are arranged in descending order, indicating the importance of latent dimensions), is the low-dimensional latent feature vector of the movie (each column represents the q-dimensional feature of a movie), It is to retain the first q largest singular values and generate an approximate matrix of rank q, which is the generated new view Figure 1 ;
[0021] S203, embed the normalized adjacency matrix into the d-dimensional latent space:
[0022]
[0023] in, is the user embedding matrix, is the movie embedding matrix, is the embedding space structural feature vector of user i (also the initial embedding vector of user i), is the embedding space structural feature vector of movie j (also the initial embedding vector of movie j);
[0024] S204, calculation of the correlation degree of the embedded space structure:
[0025]
[0026] Where Π(·) is a binary indicator function that returns 1 when the condition is met and 0 otherwise.
[0027] This formula is used to calculate the structural correlation between users and movies in the embedding space. If their correlation is high (greater than the threshold β), the interaction is retained and the calculated r u,v The value is written into the new adjacency matrix as the weight; if the correlation is low (less than or equal to the threshold β), the value is 0, that is, the interaction is deleted. After calculating all interactions, a new adjacency matrix is constructed and a new view is formed. Figure 2 .
[0028] Further as a preferred technical solution of the present invention, step S3 includes the following steps:
[0029] S301. Use multi-layer graph convolution for message passing. The embedding of users and movies at layer l is updated as follows:
[0030]
[0031] Among them, the user embedding of the lth layer is the aggregate information of the current layer and the embedding of the previous layer The same is true for the movie embedding at layer l;
[0032] S302: Cross-layer embedding summation, adding up the embeddings of each layer to generate the final embedding representations of users and movies respectively:
[0033]
[0034] S303, preference prediction, using the inner product to calculate the user's predicted score for the movie:
[0035]
[0036] in, represents the predicted score of user i for movie j.
[0037] Further as a preferred technical solution of the present invention, step S4 includes the following steps:
[0038] S401, main task training, using BPR loss function:
[0039]
[0040] Among them, O represents all triplets in the training set, represents the predicted score of the interacting user u and movie i, represents the predicted score of user u and movie j without interaction;
[0041] S402, self-supervised contrastive learning, using InfoNCE loss function:
[0042]
[0043] Among them, {u, u′∈U, u≠u′}, {v, v′∈U, v≠v′}, (h u ′,h u ″) is the embedding representation of the same user u in two new views, which is considered to be positive, (h u ′,h u ″ ′ ) are the embedding representations of different users in two new views, which are regarded as negative pairs. The same is true for movies. The final self-supervised contrastive learning loss is This formula shortens the positive distance and increases the negative distance to ensure the recommendation performance;
[0044] S403, joint training, shared graph convolution parameters:
[0045]
[0046] Here, θ is the parameter of the graph convolutional network, and λ1 and λ2 are the hyperparameters controlling self-supervised contrastive learning and L2 regularization, respectively.
[0047] The recommended method for lightweight data-enhanced graph contrast learning described in the present invention, using the above technical solution, has the following technical effects compared with the existing technology:
[0048] (1) The present invention uses approximate singular value decomposition and embedded spatial structure correlation to perform data enhancement, which alleviates data noise interaction and improves data quality.
[0049] (2) The data enhancement strategy of the present invention only needs to be executed once before training, which greatly reduces the training time and alleviates the training overhead.
[0050] (3) The present invention uses the main recommendation task and the graph contrast learning task for joint training and shares graph convolution parameters, which not only reduces the training overhead but also improves the accuracy and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a flow chart of the steps of the present invention.
[0052] Figure 2 This is a model framework diagram of the present invention. DETAILED DESCRIPTION
[0053] The present invention will be further explained below in detail with reference to the accompanying drawings so that those skilled in the art can have a deeper understanding of the present invention and be able to implement it. However, the following reference examples are only used to explain the present invention and are not intended to limit the present invention.
[0054] according to Figure 1 and Figure 2 As shown in Figure 1, a recommended method for lightweight data augmentation graph contrastive learning includes the following steps:
[0055] S1. Data preprocessing: Using the MovieLens 10M dataset, extract key information and construct an adjacency matrix;
[0056] S2, lightweight data enhancement: use approximate singular value decomposition and embedding spatial structure correlation to perform data enhancement and construct two new views;
[0057] S3. Embedded representation learning: Using multi-layer graph convolutional networks to learn representations of different views;
[0058] S4. Model training: Use a joint training strategy of the main task and contrastive learning to share graph convolution parameters.
[0059] Step S1 includes the following steps:
[0060] S101. Collect 10M of open-source MovieLens data, clean it, retain the user ID, movie ID, and rating data, and then divide it into training, validation, and test sets according to the ratio of 8:1:1.
[0061] S102, construct the user ID, movie ID and rating data in the training set data into an adjacency matrix A, the user set U={u1,u2,...,u I}, (|U|=I), movie set V={v1,v2,...,v J},(|V|=J), A i,j The value of is the rating data, or 0 if there is no rating.
[0062] Step S2 includes the following steps:
[0063] S201, normalize the adjacency matrix A:
[0064]
[0065] in, is the user degree matrix, whose diagonal elements represent the number of interactions of each user, is the movie degree matrix, whose diagonal elements represent the number of interactions for each movie;
[0066] S202. Perform approximate singular value decomposition on the normalized adjacency matrix A:
[0067]
[0068] Where q is the number of singular values retained (a hyperparameter that controls the accuracy of the low-rank approximation. The values of q range from {2, 5, 8, 10}, and the recall rate is highest when q is 5). is the user's low-dimensional latent feature vector (each row represents a user's q-dimensional feature), is the singular value matrix (diagonal elements are arranged in descending order, indicating the importance of latent dimensions), is the low-dimensional latent feature vector of the movie (each column represents the q-dimensional feature of a movie), It is to retain the first q largest singular values and generate an approximate matrix of rank q, which is the generated new view Figure 1 ;
[0069] S203, embed the normalized adjacency matrix into the d-dimensional latent space:
[0070]
[0071] in, is the user embedding matrix, is the movie embedding matrix, is the embedding space structural feature vector of user i (also the initial embedding vector of user i), is the embedding space structural feature vector of movie j (also the initial embedding vector of movie j);
[0072] S204, calculation of the correlation degree of the embedded space structure:
[0073]
[0074] Where Π(·) is a binary indicator function that returns 1 when the condition is met and 0 otherwise. β takes values of {0.02, 0.04, ..., 0.1}. When β is 0.02, the recall rate is the highest.
[0075] This formula is used to calculate the structural correlation between users and movies in the embedding space. If their correlation is high (greater than the threshold β), the interaction is retained and the calculated r u,v The value is written into the new adjacency matrix as the weight; if the correlation is low (less than or equal to the threshold β), the value is 0, that is, the interaction is deleted. After calculating all interactions, a new adjacency matrix is constructed and a new view is formed. Figure 2 .
[0076] Step S3 includes the following steps:
[0077] S301. Use multi-layer graph convolution for message passing. The embedding of users and movies at layer l is updated as follows:
[0078]
[0079] Among them, the user embedding of the lth layer is the aggregate information of the current layer and the embedding of the previous layer The sum of the movie embeddings in the lth layer is similar, and the number of graph convolution layers is 2;
[0080] S302: Cross-layer embedding summation, adding up the embeddings of each layer to generate the final embedding representations of users and movies respectively:
[0081]
[0082] S303, preference prediction, using the inner product to calculate the user's predicted score for the movie:
[0083]
[0084] in, represents the predicted score of user i for movie j.
[0085] Step S4 includes the following steps:
[0086] S401, main task training, using BPR loss function:
[0087]
[0088] Among them, O represents all triplets in the training set, represents the predicted score of the interacting user u and movie i, represents the predicted score of user u and movie j without interaction;
[0089] S402, self-supervised contrastive learning, using InfoNCE loss function:
[0090]
[0091] Among them, {u, u′∈U, u≠u′}, {v, v′∈U, v≠v′}, (h u ′,h u ″) is the embedding representation of the same user u in two new views, which is considered to be positive, (h u ′,h u ″ ′ ) are the embedding representations of different users in two new views, which are regarded as negative pairs. The same is true for movies. The final self-supervised contrastive learning loss is This formula shortens the positive distance and increases the negative distance to ensure the recommendation performance;
[0092] S403, joint training, shared graph convolution parameters:
[0093]
[0094] Among them, θ is the parameter of the graph convolutional network, λ1 and λ2 are the hyperparameters controlling self-supervised contrastive learning and L2 regularization respectively, the learning rate is [1e-5, 1e-3], the learning rate is 5e-4, the λ1 value range is {1e-4, 1e-3, 1e-2, 1e-1, 1}, and the effect is better when λ1 is 1e-1, the λ2 value range is {1e-6, 1e-5, ..., 1e-2}, and the effect is better when λ2 is 1e-4.
[0095] This paper proposes a lightweight data-enhanced graph contrastive learning recommendation method, which uses approximate singular value decomposition and embedded spatial structure correlation for data enhancement, reduces noise interaction, improves data quality, reduces training time, and improves model accuracy and robustness.
[0096] The specific implementation scheme described above further illustrates in detail the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above is only a specific implementation scheme of the present invention and is not intended to limit the scope of the present invention. Any equivalent changes and modifications made by any technician in this field without departing from the concept and principle of the present invention should fall within the scope of protection of the present invention.
Claims
1. A lightweight data-enhanced graph contrast learning recommendation method, characterized by: The following steps are involved: S1. Data preprocessing: Using the MovieLens 10M dataset, extract key information and construct an adjacency matrix; S2, lightweight data enhancement: use approximate singular value decomposition and embedding spatial structure correlation to perform data enhancement and construct two new views; S3. Embedded representation learning: Using multi-layer graph convolutional networks to learn representations of different views; S4. Model training: Use a joint training strategy of the main task and contrastive learning to share graph convolution parameters.
2. A lightweight data-enhanced graph contrast learning recommendation method according to claim 1, characterized in that: The step S1 comprises the following steps: S101. Collect 10M of open-source MovieLens data, clean it, retain the user ID, movie ID, and rating data, and then divide it into training, validation, and test sets according to the ratio of 8:1:
1. S102, construct the user ID, movie ID and rating data in the training set data into an adjacency matrix A, the user set U={u1,u2,...,u I }, (|U|=I), movie set V={v1,v2,...,v J },(|V|=J), A i,j The value of is the rating data, or 0 if there is no rating.
3. A lightweight data-enhanced graph contrast learning recommendation method according to claim 2, characterized in that: The step S2 comprises the following steps: S201, normalize the adjacency matrix A: in, is the user degree matrix, whose diagonal elements represent the number of interactions of each user, is the movie degree matrix, whose diagonal elements represent the number of interactions for each movie; S202. Perform approximate singular value decomposition on the normalized adjacency matrix A: Among them, q is the number of singular values retained, which is a hyperparameter that controls the accuracy of the low-rank approximation. is the user's low-dimensional latent feature vector, each row represents a user's q-dimensional feature, is a singular value matrix, with diagonal elements arranged in descending order, indicating the importance of latent dimensions, It is the low-dimensional latent feature vector of the movie, and each column represents the q-dimensional feature of a movie. It is to retain the first q largest singular values and generate an approximate matrix of rank q, which is the generated new view 1; S203, embed the normalized adjacency matrix into the d-dimensional latent space: in, is the user embedding matrix, is the movie embedding matrix, is the embedding space structural feature vector of user i, and is also the initial embedding vector of user i. is the embedding space structural feature vector of movie j, and is also the initial embedding vector of movie j; S204, calculation of the correlation degree of the embedded space structure: Where Π(·) is a binary indicator function that returns 1 when the condition is met and 0 otherwise; This formula is used to calculate the structural correlation between users and movies in the embedding space. If their correlation is high and greater than the threshold β, the interaction is retained and the calculated r u,v The value is written into the new adjacency matrix as the weight; if the correlation is low, less than or equal to the threshold β, the value is 0, that is, the interaction is deleted; after calculating all interactions, a new adjacency matrix is constructed and a new view 2 is formed.
4. A lightweight data-enhanced graph contrast learning recommendation method according to claim 3, characterized in that: The step S3 comprises the following steps: S301. Use multi-layer graph convolution for message passing. The embedding of users and movies at layer l is updated as follows: Among them, the user embedding of the lth layer is the aggregate information of the current layer and the embedding of the previous layer The same is true for the movie embedding at level l; S302: Cross-layer embedding summation, adding up the embeddings of each layer to generate the final embedding representations of users and movies respectively: S303, preference prediction, using the inner product to calculate the user's predicted score for the movie: in, represents the predicted score of user i for movie j.
5. The lightweight data-enhanced graph contrast learning recommendation method according to claim 4, characterized in that: The step S4 comprises the following steps: S401, main task training, using BPR loss function: Among them, O represents all triplets in the training set, represents the predicted score of the interacting user u and movie i, represents the predicted score of user u and movie j without interaction; S402, self-supervised contrastive learning, using InfoNCE loss function: Among them, {u, u′∈U, u≠u′}, {v, v′∈U, v≠v′}, (h′ u ,h″ u ) is the embedding representation of the same user u in two new views, which is considered to be positive, (h′ u ,h″ u′ ) are the embedding representations of different users in two new views, which are regarded as negative pairs, and the same is true for movies; the final self-supervised contrastive learning loss is This formula shortens the positive distance and increases the negative distance to ensure the recommendation performance; S403, joint training, shared graph convolution parameters: Here, θ is the parameter of the graph convolutional network, and λ1 and λ2 are the hyperparameters controlling self-supervised contrastive learning and L2 regularization, respectively.