Recommendation system noise pruning and long tail enhancement method based on two-stage graph optimization

Through a two-stage graph optimization method, noisy edges are dynamically pruned and long-tail user links are enhanced, which solves the problems of noise interaction and data sparsity in implicit feedback recommendation systems, improves the performance and fairness of the recommendation system, and enhances its adaptability and robustness.

CN120745810APending Publication Date: 2025-10-03NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510837243.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Noisy interactions and long-tail user data sparsity in implicit feedback recommendation systems lead to degraded recommendation performance and unfairness. Existing methods find it difficult to collaboratively optimize the interaction graph structure.

Method used

A two-stage graph optimization method is adopted to evaluate the reliability of interaction edges through node similarity indicators, dynamically prune noisy edges and generate denoised subgraphs, combine probabilistic sampling mechanism to enhance long-tail user links, and optimize the embedding representation of graph convolutional network models.

Benefits of technology

It significantly improves the overall performance and fairness of the recommendation system, dynamically adapts to changes in graph structure during training, and ensures the stability of the training process and the interpretability of recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745810A_ABST
    Figure CN120745810A_ABST
Patent Text Reader

Abstract

The invention discloses a recommendation system noise trimming and long tail enhancement method based on two-stage graph optimization. In order to solve the problems of noise interaction and long-tail user data sparsity in an implicit feedback recommendation system, the method comprises the following steps: firstly, constructing a user-article interaction bipartite graph, and initializing a graph convolutional network model to generate a preliminary embedded representation; in the first stage, the reliability of an interaction edge is evaluated through a node similarity index (Nsim), a noise edge is trimmed in combination with a dynamic threshold strategy, and a de-noised subgraph is generated to improve the embedding quality. And in the second stage, for the long-tail user, a probability sampling mechanism is adopted to add a high-confidence potential interaction edge, and an enhanced sub-graph is generated to improve the long-tail recommendation effect. Finally, Bayesian personalized ranking (BPR) loss is optimized through iterative training, and an accurate personalized recommendation result is generated. The accuracy and fairness of the recommendation system are remarkably improved, and the method is suitable for application scenes such as e-commerce, social media and content recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of recommendation systems, and in particular to a method based on two-stage graph optimization, which optimizes the interaction graph structure in implicit feedback recommendation systems by dynamically pruning noisy interaction edges and enhancing long-tail user links, so as to improve recommendation performance and fairness. Background Art

[0002] With the popularity of internet applications, implicit feedback recommendation systems have been widely used in e-commerce, social media, and content recommendation platforms. These systems provide personalized recommendations to users by analyzing user behavior data such as clicks and browsing. In recent years, the graph-based collaborative filtering (GCF) method has attracted attention due to its ability to capture high-order dependencies in user-item interactions. GCF models user-item interactions as a bipartite graph and uses a graph convolutional network (GCN) to aggregate multi-layer neighborhood information to generate embedding representations of users and items. However, the ubiquitous noisy interactions in implicit feedback data and the sparsity of long-tail user data pose significant challenges to the performance of GCF.

[0003] Noisy interactions, such as users' accidental clicks or interactions that are not genuinely interesting, are considered positive examples in the graph. Due to the multi-hop propagation mechanism of GCN, these noises will be amplified to high-order neighborhoods, distorting the embedding representations of users and items, resulting in a decrease in recommendation accuracy. At the same time, long-tail users, that is, users with fewer interactions, have poor embedding representation quality due to insufficient neighborhood information. The recommendation results are often biased towards popular items, further exacerbating the unfairness of the recommendation system. Existing technologies usually propose solutions to the noise or long-tail problems separately. For example, some methods reduce the impact of noise by adjusting the interaction weights, or increase the samples of long-tail users by resampling. However, these methods often deal with a single problem independently, ignoring the interaction between noise and data sparsity, and it is difficult to achieve joint optimization of the interaction graph structure. In addition, existing methods have limited ability to dynamically adapt to changes in the graph structure during training, and the optimization effect is restricted.

[0004] Therefore, to address the dual challenges of noisy interactions and sparsity of long-tail user data in implicit feedback recommendation systems, there is an urgent need for a method that can collaboratively optimize the interaction graph structure and improve recommendation performance and fairness. Summary of the Invention

[0005] Purpose of the invention: In response to the problems existing in the above-mentioned background technology, the present invention proposes a noise pruning and long-tail enhancement method for recommendation systems based on two-stage graph optimization, aiming to solve the performance bottlenecks caused by noise interaction and sparsity of long-tail user data in implicit feedback recommendation systems, and improve the overall accuracy and fairness of the recommendation system.

[0006] Technical solution: The present invention provides a method for noise pruning and long-tail enhancement in a recommendation system based on a two-stage graph optimization, comprising the following steps:

[0007] Step S1: Input implicit feedback interaction data, construct the initial user-item interaction bipartite graph, and initialize the graph convolutional network model to generate preliminary embedding representations of users and items.

[0008] Step S2: In the first stage, the reliability of the interaction edges is evaluated by the node similarity index (Nsim), and a dynamic threshold strategy is adopted to prune the noise edges with low reliability to generate a denoised subgraph, and the embedding representation is updated based on the subgraph.

[0009] Step S3: In the second stage, for long-tail users (i.e., users with a small number of interactions), a probabilistic sampling model is used to selectively add high-confidence potential interaction edges to generate an enhanced subgraph and further optimize the embedding representation.

[0010] Step S4: Combine the subgraphs generated in steps S2 and S3, perform iterative training on the graph convolutional network model, optimize the Bayesian personalized ranking (BPR) loss, and update the embedding representations of users and items.

[0011] Step S5: Apply the trained model to the test set to generate user personalized recommendation results.

[0012] Preferably, in step S2, the node similarity index (Nsim) evaluates the interaction reliability by the neighborhood similarity of users and items. For each interaction edge (u, i), the Nsim value s u,i Calculated as user two-hop embedding One-hop embedding with items Inner product and two-hop embedding of items One-hop embedding with users The average value of the inner product of is

[0013]

[0014] Dynamic threshold θ t It decays exponentially with the number of training rounds t, and the calculation formula is

[0015]

[0016] where θ initis the initial threshold, β is the attenuation factor, and T1 is the total number of rounds in the first stage. By retaining u,i ≥θ t The edge generation, the adjacency matrix is ​​defined as if s u,i ≥θ t , A denoise (u, i) = 1; otherwise 0.

[0017] Preferably, in step S3, long-tail user enhancement is achieved through a probabilistic sampling mechanism. Long-tail users are defined as the set of users with the last 80% of the number of interactions. For each long-tail user u, a candidate edge set ε with a ratio α is sampled from the uninteracted items. u , calculate the Nsim value s of the candidate edge (u, j) uj , and normalized to the enhanced probability, the formula is

[0018]

[0019] Where τ is the temperature parameter. Generated by Bernoulli sampling, the adjacency matrix is ​​defined as A augment (u, j) with probability p u,j 1 if sampled; 0 otherwise.

[0020] Preferably, in step S4, the training process optimizes the BPR loss, which is:

[0021]

[0022] in is the set of interactive items of user u, and λ is the regularization parameter. The embedding smoothing strategy is expressed as Ensure phase switching stability, where γ is a smoothing factor.

[0023] Beneficial effects:

[0024] This paper uses a two-stage graph optimization framework to collaboratively prune noisy edges and enhance long-tail user links, effectively addressing the dual challenges of noise interaction and data sparsity in implicit feedback recommendation systems, significantly improving the overall performance and fairness of the recommendation system. A dynamic threshold and probability enhancement mechanism enhance the adaptability and robustness of the method, enabling it to dynamically adapt to changes in graph structure during training. Furthermore, through phased optimization and embedded smoothing strategies, the method ensures the stability of the training process and the interpretability of recommendation results, providing an efficient and fair solution for recommendation systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1This is a flowchart of the two-stage graph optimization method proposed in this invention, which shows the collaborative optimization process of dynamic noise edge pruning and long-tail user link enhancement, including the overall process of interaction graph construction, denoising subgraph generation, enhanced subgraph generation and model training.

[0026] Figure 2 It is the algorithm framework diagram of the method of the present invention during the training process, which describes in detail the specific calculation steps of Nsim calculation, dynamic threshold adjustment, probability sampling and embedding update. DETAILED DESCRIPTION

[0027] The technical solution of the present invention is further described below with reference to the accompanying drawings. It is apparent that the embodiments described herein are only some embodiments of the present invention, not all embodiments. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0028] The present invention provides a method for noise pruning and long-tail enhancement of recommendation systems based on two-stage graph optimization. The specific process is as follows: Figure 1 As shown in the figure, by optimizing the interaction graph structure in stages, the problems of noisy interactions and long-tail user data sparsity are solved collaboratively. The method includes the following steps:

[0029] Step S1: Input implicit feedback interaction data, construct the initial user-item interaction bipartite graph, and initialize the graph convolutional network model to generate a preliminary embedding representation. The interaction data comes from user behavior records (such as clicks, purchases), including user sets, item sets, and interaction edge sets, formalized as a bipartite graph. in Represents the user node, Denotes an item node, and ε denotes an interaction edge. The model is initialized using the LightGCN framework, generating embeddings for users and items through multi-layer neighborhood aggregation. Each layer of embeddings is updated through weighted aggregation of neighboring nodes, with weights normalized based on node degrees to capture high-order dependencies in interactions. This initial embedding provides a foundation for subsequent optimization, ensuring the model can initially learn interaction patterns between users and items.

[0030] Step S2: In the first stage, the reliability of the interaction edge is evaluated by the node similarity index (Nsim), and the low-reliability noise edge is pruned using a dynamic threshold strategy to generate a denoised subgraph, and the embedding representation is updated based on the subgraph. Nsim quantifies the confidence of the interaction by analyzing the neighborhood similarity of users and items. For each interaction edge (u, i), the similarity of the user perspective is calculated by the user two-hop embedding One-hop embedding with items The inner product calculation of the item perspective is performed by embedding the item two hops One-hop embedding with users The inner product calculation of Nsim is the mean of the two, and the formula is

[0031]

[0032] To alleviate the heterogeneity of the embedding space. Dynamic threshold θ t Designed by exponential decay function, the form is

[0033]

[0034] where θ init is the initial threshold, β is the decay factor (usually 0.5), and T1 is the total number of rounds in the first stage. The high threshold in the early stage retains more edges to maintain the connectivity of the graph, and the threshold is gradually reduced in the later stage to focus on high-quality interactions. By retaining the Nsim value s u,i ≥θ t The edge generation, the adjacency matrix is ​​defined as if s u,i ≥θ t , A denoise (u, i) = 1; otherwise 0. LightGCN is trained on the denoised subgraph, optimizing the BPR loss to generate a robust embedding representation, laying the foundation for subsequent long-tail enhancement.

[0035] Step S3: In the second stage, for long-tail users, high-confidence potential interaction edges are added through the probabilistic sampling model to generate enhanced subgraphs and further optimize the embedding representation. Long-tail users are defined as the set of users with the last 80% of the interaction number. By counting the number of interactions and sorting them in descending order, we can determine the last 80%. Randomly sample a certain proportion (usually 10%) of the candidate edge set ε from the uninteracted items u The Nsim value s of the candidate edge (u, j) u,j Obtained by the same calculation method as step S2 and normalized to the enhancement probability, the formula is

[0036]

[0037] Where τ is a temperature parameter (usually 0.1) that controls the smoothness of the probability distribution. Generated by Bernoulli sampling, the adjacency matrix of the newly added edge is defined as A augment (u, j) with probability p u,j The sampled value is 1; otherwise, it is 0. This probability enhancement mechanism prioritizes edges with high Nsim values ​​to ensure reliability, while also avoiding biased recommendations towards popular items through random sampling. LightGCN is trained on the enhanced subgraph, optimizing the BPR loss to further improve the representation quality of long-tail users.

[0038] Step S4: Combine the denoised subgraph and the enhanced subgraph for iterative training, optimize the BPR loss, and update the embedding representation of users and items. The BPR loss is defined by maximizing the ranking difference between observed interactions and unobserved interactions, in the form of

[0039]

[0040] in is the set of interactive items of user u, and λ is the regularization parameter. In the first stage, the denoising subgraph is used Training, focusing on learning robust embeddings; the second stage switches to enhancing subgraphs Optimize the representation of long-tail users. Phase switching is controlled by the preset round T1, which is usually 60% of the total round T. To avoid embedding mutations caused by phase switching, the embedding smoothing strategy is implemented by Calculate, where γ is a smoothing factor (usually 0.8). Training is repeated for a preset number of rounds, and the optimal model parameters are saved.

[0041] Step S5: Apply the trained model to the test set and generate personalized recommendation results based on the final embedding representation of users and items. Test set recommendations are made by calculating the inner product of user embedding and item embedding. Output a list of recommendations sorted by score.

[0042] The foregoing is merely a preferred embodiment of the present invention. Persons skilled in the art may, without departing from the principles of the present invention, make various improvements and adjustments to the embodiment, such as adjusting the attenuation factor of the dynamic threshold, the sampling ratio of the enhanced edge, or the value of the smoothing factor. Such improvements and adjustments shall be deemed to be within the scope of protection of the present invention.

Claims

1. A noise pruning and long-tail enhancement method for recommendation systems based on two-stage graph optimization, characterized by: The method comprises the following steps: Step S1: Input implicit feedback interaction data, construct a bipartite graph of user-item interactions, initialize the graph convolutional network model, and generate a preliminary embedding representation; Step S2: In the first stage, the reliability of the interaction edges is evaluated by the node similarity index (Nsim), and the dynamic threshold strategy is used to prune the noisy edges, generate the denoised subgraph, and update the embedding representation; Step S3: In the second stage, high-confidence potential interaction edges are added for long-tail users through probabilistic sampling to generate enhanced subgraphs and optimize the embedding representation; Step S4: Combine the denoised subgraph and the enhanced subgraph, iteratively train the graph convolutional network, optimize the Bayesian personalized ranking (BPR) loss, and update the embedding representation; Step S5: Apply the trained model to the test set to generate personalized recommendation results.

2. The method for noise pruning and long-tail enhancement of recommendation systems based on dual-stage graph optimization according to claim 1, characterized in that: In step S2, the node similarity index (Nsim) evaluates the interaction reliability by the neighborhood similarity of users and items; specifically, for each interaction edge (u, i), the Nsim value s u,i Calculated as user two-hop embedding One-hop embedding with items Inner product and two-hop embedding of items One-hop embedding with users The average value of the inner product of is Dynamic threshold θ t It decays exponentially with the number of training rounds t, and the calculation formula is where θ init is the initial threshold, β is the attenuation factor, T1 is the total number of rounds in the first stage; denoising subgraph By retaining u,i ≥θ t The edge is generated.

3. The method for noise pruning and long-tail enhancement of recommendation systems based on dual-stage graph optimization according to claim 2, characterized in that: In step S3, the long tail users are defined as the set of users with the last 80% of the interaction number. For each long-tail user Sample candidate edge sets ε with a ratio α from uninteracted items u , calculate the Nsim value s of the candidate edge (u, j) u,j , and normalized to the enhanced probability, the formula is Where τ is the temperature parameter; the enhanced subgraph Generated by Bernoulli sampling, the adjacency matrix is ​​defined as A augment (u,j) with probability p u,j 1 if sampled, 0 otherwise.

Citation Information

Cited By

  • Confidence-guided adaptive graph representation reinforcement learning method, equipment and medium

    CN121145970A

  • A two-stage graph recommendation method and system with enhanced distribution robustness

    CN122388268A