Dynamic retrieval enhancement recommendation method based on graph reasoning
By building a dynamic reasoning graph structure of user behavior and items, combined with graph neural networks and reinforcement learning, the problems of user interest drift and insufficient item exposure in the recommendation system are solved, the performance of the recommendation system and user experience are improved, and explainable recommendation decisions are provided.
Patent Information
- Application Number
- CN202510745512.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
AI Technical Summary
Existing recommendation systems have difficulty capturing dynamic changes in users' interests and have poor item recommendation effects, and the recommendation results are not interpretable enough.
By constructing a dynamic reasoning graph structure of user behavior, memory, and candidate items, combining graph neural networks and reinforcement learning mechanisms, a dynamic memory retrieval module is designed, and a graph attention network is used for dynamic classification and weight allocation to optimize the graph reasoning strategy.
It effectively solves the problems of user interest drift and insufficient item exposure, improves the performance of the recommendation system and user experience, and provides explainable transparency of recommendation decisions.
Smart Images

Figure CN120632215A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a dynamic retrieval enhancement recommendation method based on graph reasoning. Background Art
[0002] In today's digital age, recommendation systems have become an indispensable component of internet applications, widely used in a variety of fields, including e-commerce, social media, online video, and music platforms. Their goal is to analyze users' historical behavior and preferences to provide personalized content or product recommendations, thereby improving user experience and the commercial value of the platform.
[0003] Traditional recommendation methods mainly rely on collaborative filtering, content-based recommendation, and hybrid recommendation methods. Collaborative filtering makes recommendations by analyzing the similarity between users or items, but has limitations when dealing with cold start problems and data sparsity. Content-based recommendation focuses on analyzing item features and user preferences, but has difficulty capturing the dynamic changes in user interests. In recent years, with the development of deep learning technology, neural network-based recommendation methods have gradually emerged. Recurrent neural networks and their variants are used to model user behavior sequences to capture users' short-term and long-term interests. However, these methods still face challenges in dealing with complex user behavior patterns and dynamic interest drift.
[0004] Furthermore, existing recommendation systems lack the interpretability of their recommendations. Users often struggle to understand the rationale behind recommendations, which not only impacts their trust but also limits further optimization of the recommendation system. To improve recommendation effectiveness, some research has explored the use of technologies such as knowledge graphs and graph neural networks to enhance the performance of recommendation systems. Knowledge graphs provide rich semantic information for recommendations by constructing relationship graphs between users, items, and context, but they have limitations in terms of dynamic updates and real-time recommendations. Graph neural networks, by learning representations of nodes and edges, can better capture the complex relationships between users and items, but how to effectively integrate users' dynamic behavior and memory patterns remains an unresolved issue. Summary of the Invention
[0005] In response to the problems that existing recommendation systems have difficulty in capturing dynamic changes in user interests and poor item recommendation results, the present invention proposes a dynamic retrieval-enhanced recommendation method based on graph reasoning. The present invention implements recommendations by constructing a dynamic reasoning graph structure of user behavior, memory, and candidate items, and combining it with a graph neural network. Specifically, the present invention designs a dynamic memory retrieval module to extract key patterns from user historical behavior to construct a memory library; constructs the user's current state, retrieval memory, and candidate items into a heterogeneous reasoning graph, and dynamically classifies and assigns weights to memory nodes through a graph attention network; at the same time, introduces a reinforcement learning mechanism, uses recommendation effect indicators as reward signals, and optimizes graph reasoning strategies end-to-end. This invention effectively solves the problems of user interest drift and insufficient item exposure in recommendation systems, improves the transparency of recommendation decisions through an explainable graph reasoning path, and enhances recommendation performance and user experience.
[0006] The present invention provides a dynamic retrieval enhancement recommendation method based on graph reasoning, comprising the following steps:
[0007] S1: Given the interaction history of any user u, it is represented as an ordered sequence of item IDs [i1,i2,...,i T ], where each item i t Represents the item index of the user interaction, T represents the maximum sequence length, and the discrete ID sequence needs to be converted into a continuous vector representation through the embedding layer; a trainable item embedding matrix F∈R ∣V∣×d , where |V| represents the total capacity of the item library and d represents the embedding dimension. Through the table lookup operation, each item ID in the sequence is converted into a corresponding d-dimensional dense vector:
[0008] v t =F[i t ,:]
[0009] In this formula, v t Represents the vector of interactive items at time t, and obtains the user history sequence matrix X=[v1,v2,...,v T ] T ∈R T×d To capture temporal patterns and long-term dependencies in sequences, a Transformer-based encoder architecture is used. The core of the encoder is a multi-head self-attention mechanism that first linearly projects the input sequence into query, key, and value spaces:
[0010] Q=XW Q
[0011] K=XW K
[0012] V=XW V
[0013] In this formula, X represents the matrix of user history sequence, and the projection matrix Map the input to a dimension d k The feature space of , Q, K, V represent the query projection, key projection, and value projection in the self-attention mechanism, and the attention weight is calculated using the scaled dot product form:
[0014]
[0015] In this formula, d k Represents the feature dimension of each attention head. To ensure the correctness of the timing, a lower triangular mask matrix is introduced before the softmax to prevent information leakage. The output of the multi-head attention is concatenated and linearly transformed, and then residually connected with the original input and layer normalized:
[0016] H = LayerNorm(X + MultiHead(X))
[0017] In this formula, X represents the matrix of the user's historical sequence, and the representation capability is further enhanced through a feed-forward network consisting of two linear transformations and a ReLU activation:
[0018] FFN(X)=max(0,XW1+b1)W2+b2
[0019] In this formula, and represents the learnable parameter, d ff Set to 4d, b1 and b2 represent bias terms, and X represents the matrix of the user's historical sequence; after stacking L layers of Transformer blocks, the hidden state of the last position is taken as the final user representation:
[0020] h T =Encoder(X)[-1,:]
[0021] In this formula, h T The user state vector output by the encoder not only contains the user's historical behavior characteristics, but also encodes rich temporal dependencies;
[0022] S2: Based on the encoded user state vector, similar historical behavior patterns are retrieved from the constructed memory library, and an approximate nearest neighbor search algorithm is used to achieve efficient retrieval;
[0023] S3: The current user state, retrieved memory patterns, and candidate items are constructed into a ternary heterogeneous reasoning graph, where nodes include user nodes, memory nodes, and item nodes, and edge weights are dynamically calculated through an attention mechanism;
[0024] S4: A multi-layer graph neural network is used to process the constructed reasoning graph, enabling information exchange between nodes through a message passing mechanism. The importance classification results of each memory node are output to complete the screening of memory patterns.
[0025] S5: Based on the node weights output by the graph neural network, the candidate items are weighted and scored to generate the final recommendation list, while retaining the complete graph reasoning path as an interpretable basis for the recommendation decision;
[0026] S6: Establish a reinforcement learning framework with recommendation effect indicators as reward functions, and optimize the parameters of memory retrieval and graph reasoning modules end-to-end through the policy gradient method.
[0027] According to a specific implementation of an embodiment of the present invention, the specific steps of S2 are:
[0028] S2, based on the encoded user state vector, retrieves similar historical behavior patterns from the constructed memory library and uses the approximate nearest neighbor search algorithm to achieve efficient retrieval;
[0029] S3: The current user state, retrieved memory patterns, and candidate items are constructed into a ternary heterogeneous reasoning graph, where nodes include user nodes, memory nodes, and item nodes, and edge weights are dynamically calculated through an attention mechanism;
[0030] S4: A multi-layer graph neural network is used to process the constructed reasoning graph, enabling information exchange between nodes through a message passing mechanism. The importance classification results of each memory node are output to complete the screening of memory patterns.
[0031] S5: Based on the node weights output by the graph neural network, the candidate items are weighted and scored to generate the final recommendation list, while retaining the complete graph reasoning path as an interpretable basis for the recommendation decision;
[0032] S6: Establish a reinforcement learning framework with recommendation effect indicators as reward functions, and optimize the parameters of memory retrieval and graph reasoning modules end-to-end through the policy gradient method.
[0033] A dynamic retrieval enhancement recommendation method based on graph reasoning, wherein the specific method of step S2 is as follows:
[0034] S2, dynamic memory retrieval uses efficient approximate nearest neighbor search technology to retrieve relevant historical behavior patterns from the memory library and convert the user state vector h T Normalized to a unit vector:
[0035]
[0036] In this formula, q represents the normalized user query vector, the memory bank M stores a large number of historical behavior patterns, and each memory item contains a key vector k i and the corresponding value vector v i , where the key vector has been L2-normalized; the memory bank can be represented as:
[0037] M={(k1,v1),(k2,v2),...,(k N ,v N )}
[0038] In this formula, N represents the total number of memory items stored in the memory bank. To improve retrieval efficiency, an improved hierarchical navigable small-world algorithm is used to establish an index structure. During retrieval, the similarity score between the query vector and the memory key vector is first calculated:
[0039] s i =q T k i
[0040] In this formula, k i Represents the index key vector of the i-th memory item, s i Represents the original similarity between the query and the i-th memory key, and finds the top-K most relevant memory items through maximum inner product search:
[0041] R={(k j ,v j )|s j ∈Top-K({s1,...,s N},K)}
[0042] In this formula, R is the memory item set of Top-K search results, s j represents the original similarity between the query and the jth memory key, N represents the total number of memory items stored in the memory library, and K is the number of retrievals, with a value of 5-20. To enhance the retrieval quality, a hybrid similarity metric is used:
[0043] s' i =αq T k i +(1-α)cosine(f q ,f i )
[0044] In this formula, α∈[0,1] represents the balance factor, k i Represents the index key vector of the i-th memory item, s' i represents the final score after the mixed similarity calculation, cosine represents the cosine similarity between auxiliary features, and f q and f iThey represent auxiliary feature vectors extracted from the user's current behavior and memory items respectively. The final retrieval results are sorted by score to form a memory set:
[0045]
[0046] In this formula, v j Represents the content vector of the jth memory item, M ret Represents the final search results, K represents the number of searches, and the weight coefficient W j Calculated by the softmax function:
[0047]
[0048] In this formula, s' i represents the final score after the hybrid similarity calculation, K is the number of retrievals, the temperature parameter τ controls the sharpness of the weight distribution, and the memory set M ret Each memory item is associated with the original user behavior fragment and context information.
[0049] According to a specific implementation of an embodiment of the present invention, the specific steps of S3 are:
[0050] S3. Design a dynamic graph structure to model the complex relationship between users, memories and items based on the user state vector h. T and retrieval memory set M ret , initialize the graph node set:
[0051]
[0052] n u =LayerNorm(h T )
[0053]
[0054] In this formula, n u Represents the user status node, The node representation of the kth candidate item, Y represents the number of items to be recommended, represents the jth retrieval memory node, v j represents the content vector of the jth memory item, Represents the set of candidate item nodes, which is obtained through the item embedding matrix E:
[0055]
[0056] k=1,...,Y
[0057] In this formula, i krepresents the item index of the user interaction, Y represents the number of items to be recommended, and the edge construction adopts a multi-type attention mechanism. The edge weight between the user node and the memory node is calculated as follows:
[0058]
[0059] In this formula, W a ∈R d×d represents the trainable parameter matrix, e u,j Indicates the connection strength between the user node and the jth memory node, n u represents the user status node, K represents the number of retrievals, and the edge between the memory node and the item node is calculated by bilinear transformation:
[0060]
[0061] In this formula, W b ∈R d×d represents the weight matrix, b represents the bias term, σ represents the sigmoid activation function, e j,k Represents the connection strength between the memory node and the item node. The final constructed graph structure is expressed as:
[0062] G=(V,E)
[0063] In this formula, G represents the constructed heterogeneous reasoning graph, and the set E contains all connections with non-zero weights. To enhance the graph's expressiveness, the following auxiliary connections are introduced:
[0064] Direct connection between user node and Top-N candidate items:
[0065]
[0066] In this formula, The node representation of the kth candidate item, n u Represents the user status node, e u,k Represents the direct connection weight between the user and the candidate item;
[0067] Memory similarity connections between nodes:
[0068] e j,l =|(sim(v j ,v l )>θ)
[0069] In this formula, | represents the indicator function, θ represents the similarity threshold, and v j Represents the content vector of the j-th memory item. The dynamic graph structure completely captures the explicit and implicit relationships between users, memories, and items.
[0070] According to a specific implementation of an embodiment of the present invention, the specific steps of S4 are:
[0071] S4. Use a multi-layer graph attention network to extract features from the constructed heterogeneous reasoning graph. For the heterogeneous reasoning graph G = (V, E), initialize the hidden state of each node:
[0072]
[0073] In this formula, Represents different specific projection matrices, represents the initial hidden state of the node, d h represents the hidden layer dimension of the graph neural network. The information transmission process adopts an improved multi-head graph attention mechanism. In the calculation of the lth layer, node v aggregates information from neighboring nodes:
[0074]
[0075] In this formula, N(v) represents the set of neighbor nodes of node v. Represents the transformation matrix of the l-th layer head h, represents the number of attention heads, and the attention coefficient Calculated as follows:
[0076]
[0077] In this formula, a h Represents the parameter vector of the h-th attention head, / / represents the vector splicing operation. To enhance the long-distance dependency capability, the network also includes cross-layer connections:
[0078]
[0079] In this formula, represents the node representation after layer l, and β represents the dynamic weight of the cross-layer connection, which is dynamically calculated through the gating mechanism:
[0080]
[0081] In this formula, a h represents the parameter vector of the h-th attention head, represents the node representation after layer l, w g and b g Represents the calculation parameters of the skip connection coefficient. After L layers of graph convolution, the importance of memory nodes is classified:
[0082]
[0083] In this formula, W c and b cRepresents the classification parameter of memory node importance, P j represents the probability distribution of the processing strategy of memory node j, Represents the representation vector of the j-th node after passing through the l-layer graph neural network, Mapping node representations to a 4-dimensional classification space, corresponding to four processing strategies for memory nodes: strengthening, retaining, weakening, and ignoring;
[0084] The memory node weights output by graph reasoning are calculated as:
[0085]
[0086] In this formula, Represents the adjusted importance weight of the memory node.
[0087] According to a specific implementation of an embodiment of the present invention, the specific steps of S5 are:
[0088] S5. Based on memory node weight and the representation of each node Calculate candidate item i k Recommendation score:
[0089]
[0090] In this formula, C k represents the recommendation score of the candidate item, Represents the representation of the user node, Represents item node i k The representation of represents the importance of the adjusted memory node, the weight c∈[0,1] represents the memory fusion coefficient, and K represents the number of retrievals. To enhance the diversity of recommendations, a category balance term is introduced:
[0091]
[0092] In this formula, C k represents the recommendation score of the candidate item, C' k represents the revised recommendation score of the candidate item, Indicates item i k The number of occurrences of the category in the current recommendation batch. f represents the diversity control parameter. After the score calculation is completed, it is converted into a probability distribution through the softmax function:
[0093]
[0094] In this formula, C' k represents the corrected recommendation score of the candidate item, p krepresents the recommendation probability of item k, Y represents the number of items to be recommended, and parameter g adjusts the concentration of the recommendation list. The final recommendation list is generated in descending order of probability:
[0095] R k ={(i k ,p k )|p k ∈TopA{p1,...,p Y},A}
[0096] In this formula, A represents the length of the recommendation list. To provide interpretability, the complete reasoning path is recorded:
[0097]
[0098] In this formula, J k Indicates that item i k The indexes of the top three memory nodes that contribute the most to the score.
[0099] According to a specific implementation of an embodiment of the present invention, the specific step of S6 is:
[0100] S6. Define the state space as S, where each state g t Contains the current user representation Memory node status and candidate item representation The action space B corresponds to the generated recommendation list R k , the reward function is designed as:
[0101] r t =η1·NDCG@10+η2·Diversity-η3·Redundancy
[0102] In this formula, η1, η2, η3 represent weight coefficients, NDCG@10 represents the ranking quality of the top 10 recommended items, and r t represents the instant reward at time t, Diversity calculates the category entropy of the recommendation list, Redundancy measures the repetition with the previous recommendation, and the policy network π θ The parameter update adopts PPO algorithm:
[0103]
[0104] In this formula, ∈ represents the clipping range of the policy update, represents the policy parameters before update, π θ represents the parameterized model for generating recommendation strategies, a represents the recommended action, and the advantage function A t Calculated by generalized odds estimate:
[0105]
[0106] δ t =r t +nV φ (s t+1 )-V φ (g t )
[0107] In this formula, n represents the discount factor, δ t represents the value difference of adjacent states, z represents the GAE parameter, and the value function network V φ The parameter update target is:
[0108]
[0109] In this formula, R t Represents the cumulative return. To stabilize the training process, a double buffer mechanism is used to alternately update the policy network and the value network, and independent learning rates are set. The sample batches for each update are randomly sampled from the experience replay pool D:
[0110]
[0111] In this formula, a i represents the recommended action performed by the i-th sample, r i represents the immediate reward obtained by the i-th sample, g i represents the immediate reward obtained by the i-th sample, the batch size x is dynamically adjusted according to the resources, and the policy exploration is maintained through the regularization term:
[0112]
[0113] In this formula, a represents the recommended action, L ent represents the entropy of the policy probability distribution, π θ Represents a parameterized model for generating recommendation strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0114] Figure 1 Flowchart of this method;
[0115] Figure 2 This is the architectural diagram of this method. DETAILED DESCRIPTION
[0116] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below with reference to examples and drawings.
[0117] As attached Figure 1 and attached Figure 2As shown in the figure, a dynamic retrieval enhancement recommendation method based on graph reasoning includes the following steps:
[0118] Step 1: Given the interaction history of any user u, represent it as an ordered sequence of item IDs [i1,i2,...,i T ], where each item i t Represents the item index of the user interaction, T represents the maximum sequence length, and the discrete ID sequence needs to be converted into a continuous vector representation through the embedding layer; a trainable item embedding matrix F∈R ∣V∣×d , where |V| represents the total capacity of the item library and d represents the embedding dimension. Through the table lookup operation, each item ID in the sequence is converted into a corresponding d-dimensional dense vector:
[0119] v t =F[i t ,:]
[0120] In this formula, v t Represents the vector of interactive items at time t, and obtains the user history sequence matrix X=[v1,v2,...,v T ] T ∈R T×d To capture temporal patterns and long-term dependencies in sequences, a Transformer-based encoder architecture is used. The core of the encoder is a multi-head self-attention mechanism that first linearly projects the input sequence into query, key, and value spaces:
[0121] Q=XW Q
[0122] K=XW K
[0123] V=XW V
[0124] In this formula, X represents the matrix of user history sequence, and the projection matrix Map the input to a dimension d k The feature space of , Q, K, V represent the query projection, key projection, and value projection in the self-attention mechanism, and the attention weight is calculated using the scaled dot product form:
[0125]
[0126] In this formula, d k Represents the feature dimension of each attention head. To ensure the correctness of the timing, a lower triangular mask matrix is introduced before the softmax to prevent information leakage. The output of the multi-head attention is concatenated and linearly transformed, and then residually connected with the original input and layer normalized:
[0127] H = LayerNorm(X + MultiHead(X))
[0128] In this formula, X represents the matrix of the user's historical sequence, and the representation capability is further enhanced through a feed-forward network consisting of two linear transformations and a ReLU activation:
[0129] FFN(X)=max(0,XW1+b1)W2+b2
[0130] In this formula, and represents the learnable parameter, d ff Set to 4d, b1 and b2 represent bias terms, and X represents the matrix of the user's historical sequence; after stacking L layers of Transformer blocks, the hidden state of the last position is taken as the final user representation:
[0131] h T =Encoder(X)[-1,:]
[0132] In this formula, h T The user state vector output by the encoder not only contains the user's historical behavior characteristics, but also encodes rich temporal dependencies;
[0133] Step 2: Dynamic memory retrieval uses efficient approximate nearest neighbor search technology to retrieve relevant historical behavior patterns from the memory bank and convert the user state vector h T Normalized to a unit vector:
[0134]
[0135] In this formula, q represents the normalized user query vector, the memory bank M stores a large number of historical behavior patterns, and each memory item contains a key vector k i and the corresponding value vector v i , where the key vector has been L2-normalized; the memory bank can be represented as:
[0136] M={(k1,v1),(k2,v2),...,(k N ,v N )}
[0137] In this formula, N represents the total number of memory items stored in the memory bank. To improve retrieval efficiency, an improved hierarchical navigable small-world algorithm is used to establish an index structure. During retrieval, the similarity score between the query vector and the memory key vector is first calculated:
[0138] s i =q T k i
[0139] In this formula, k i Represents the index key vector of the i-th memory item, s i Represents the original similarity between the query and the i-th memory key, and finds the top-K most relevant memory items through maximum inner product search:
[0140] R={(k j ,v j )|s j ∈Top-K({s1,...,s N},K)}
[0141] In this formula, R is the memory item set of Top-K search results, s j represents the original similarity between the query and the jth memory key, N represents the total number of memory items stored in the memory library, and K is the number of retrievals, with a value of 5-20. To enhance the retrieval quality, a hybrid similarity metric is used:
[0142] s' i =αq T k i +(1-α)cosine(f q ,f i )
[0143] In this formula, α∈[0,1] represents the balance factor, k i Represents the index key vector of the i-th memory item, s' i represents the final score after the mixed similarity calculation, cosine represents the cosine similarity between auxiliary features, and f q and f i They represent auxiliary feature vectors extracted from the user's current behavior and memory items respectively. The final retrieval results are sorted by score to form a memory set:
[0144]
[0145] In this formula, v j Represents the content vector of the jth memory item, M ret Represents the final search results, K represents the number of searches, and the weight coefficient W j Calculated by the softmax function:
[0146]
[0147] In this formula, s' i represents the final score after the hybrid similarity calculation, K is the number of retrievals, the temperature parameter τ controls the sharpness of the weight distribution, and the memory set M ret Each memory item is associated with the original user behavior fragment and context information;
[0148] Step 3: Design a dynamic graph structure to model the complex relationships between users, memories, and items based on the user state vector h T and retrieval memory set M ret , initialize the graph node set:
[0149]
[0150] n u =LayerNorm(h T )
[0151]
[0152] In this formula, n u Represents the user status node, The node representation of the kth candidate item, Y represents the number of items to be recommended, represents the jth retrieval memory node, v j represents the content vector of the jth memory item, Represents the set of candidate item nodes, which is obtained through the item embedding matrix E:
[0153]
[0154] k=1,...,Y
[0155] In this formula, i k represents the item index of the user interaction, Y represents the number of items to be recommended, and the edge construction adopts a multi-type attention mechanism. The edge weight between the user node and the memory node is calculated as follows:
[0156]
[0157] In this formula, W a ∈R d×d represents the trainable parameter matrix, e u,j Indicates the connection strength between the user node and the jth memory node, n u represents the user status node, K represents the number of retrievals, and the edge between the memory node and the item node is calculated by bilinear transformation:
[0158]
[0159] In this formula, W b ∈R d×d represents the weight matrix, b represents the bias term, σ represents the sigmoid activation function, e j,k Represents the connection strength between the memory node and the item node. The final constructed graph structure is expressed as:
[0160] G=(V,E)
[0161] In this formula, G represents the constructed heterogeneous reasoning graph, and the set E contains all connections with non-zero weights. To enhance the graph's expressiveness, the following auxiliary connections are introduced:
[0162] Direct connection between user node and Top-N candidate items:
[0163]
[0164] In this formula, The node representation of the kth candidate item, n u Represents the user status node, e u,k Represents the direct connection weight between the user and the candidate item;
[0165] Memory similarity connections between nodes:
[0166] e j,l =|(sim(v j ,v l )>θ)
[0167] In this formula, | represents the indicator function, θ represents the similarity threshold, and v j Represents the content vector of the jth memory item. The dynamic graph structure fully captures the explicit and implicit relationships between users, memories, and items;
[0168] Step 4: Use a multi-layer graph attention network to extract features from the constructed heterogeneous reasoning graph. For the heterogeneous reasoning graph G = (V, E), initialize the hidden state of each node:
[0169]
[0170] In this formula, Represents different specific projection matrices, represents the initial hidden state of the node, d h represents the hidden layer dimension of the graph neural network. The information transmission process adopts an improved multi-head graph attention mechanism. In the calculation of the lth layer, node v aggregates information from neighboring nodes:
[0171]
[0172] In this formula, N(v) represents the set of neighbor nodes of node v. Represents the transformation matrix of the l-th layer head h, represents the number of attention heads, and the attention coefficient Calculated as follows:
[0173]
[0174] In this formula, a h Represents the parameter vector of the h-th attention head, / / represents the vector splicing operation. To enhance the long-distance dependency capability, the network also includes cross-layer connections:
[0175]
[0176] In this formula, represents the node representation after layer l, and β represents the dynamic weight of the cross-layer connection, which is dynamically calculated through the gating mechanism:
[0177]
[0178] In this formula, a h represents the parameter vector of the h-th attention head, represents the node representation after layer l, w g and b g Represents the calculation parameters of the skip connection coefficient. After L layers of graph convolution, the importance of memory nodes is classified:
[0179]
[0180] In this formula, W c and b c Represents the classification parameter of memory node importance, P j represents the probability distribution of the processing strategy of memory node j, Represents the representation vector of the j-th node after passing through the l-layer graph neural network, Mapping node representations to a 4-dimensional classification space, corresponding to four processing strategies for memory nodes: strengthening, retaining, weakening, and ignoring;
[0181] The memory node weights output by graph reasoning are calculated as:
[0182]
[0183] In this formula, Represents the adjusted importance weight of the memory node;
[0184] Step 5: Based on memory node weight and the representation of each node Calculate candidate item i k Recommendation score:
[0185]
[0186] In this formula, C k represents the recommendation score of the candidate item, Represents the representation of the user node, Represents item node ik The representation of represents the importance of the adjusted memory node, the weight c∈[0,1] represents the memory fusion coefficient, and K represents the number of retrievals. To enhance the diversity of recommendations, a category balance term is introduced:
[0187]
[0188] In this formula, C k represents the recommendation score of the candidate item, C' k represents the revised recommendation score of the candidate item, Indicates item i k The number of occurrences of the category in the current recommendation batch. f represents the diversity control parameter. After the score calculation is completed, it is converted into a probability distribution through the softmax function:
[0189]
[0190] In this formula, C' k represents the corrected recommendation score of the candidate item, p k represents the recommendation probability of item k, Y represents the number of items to be recommended, and parameter g adjusts the concentration of the recommendation list. The final recommendation list is generated in descending order of probability:
[0191] R k ={(i k ,p k )|p k ∈TopA{p1,...,p Y},A}
[0192] In this formula, A represents the length of the recommendation list. To provide interpretability, the complete reasoning path is recorded:
[0193]
[0194] In this formula, J k Indicates that item i k The indexes of the top three memory nodes with the largest score contribution;
[0195] Step 6: Define the state space as S, where each state g t Contains the current user representation Memory node status and candidate item representation The action space B corresponds to the generated recommendation list R k , the reward function is designed as:
[0196] r t=η1·NDCG@10+η2·Diversity-η3·Redundancy
[0197] In this formula, η1, η2, η3 represent weight coefficients, NDCG@10 represents the ranking quality of the top 10 recommended items, and r t represents the instant reward at time t, Diversity calculates the category entropy of the recommendation list, Redundancy measures the repetition with the previous recommendation, and the policy network π θ The parameter update adopts PPO algorithm:
[0198]
[0199] In this formula, ∈ represents the clipping range of the policy update, represents the policy parameters before update, π θ represents the parameterized model for generating recommendation strategies, a represents the recommended action, and the advantage function A t Calculated by generalized odds estimate:
[0200]
[0201] δ t =r t +nV φ (s t+1 )-V φ (g t )
[0202] In this formula, n represents the discount factor, δ t represents the value difference of adjacent states, z represents the GAE parameter, and the value function network V φ The parameter update target is:
[0203]
[0204] In this formula, R t Represents the cumulative return. To stabilize the training process, a double buffer mechanism is used to alternately update the policy network and the value network, and independent learning rates are set. The sample batches for each update are randomly sampled from the experience replay pool D:
[0205]
[0206] In this formula, a i represents the recommended action performed by the i-th sample, r i represents the immediate reward obtained by the i-th sample, g i represents the immediate reward obtained by the i-th sample, the batch size x is dynamically adjusted according to the resources, and the policy exploration is maintained through the regularization term:
[0207]
[0208] In this formula, a represents the recommended action, L ent represents the entropy of the policy probability distribution, π θ Represents a parameterized model for generating recommendation strategies.
Claims
1. A dynamic retrieval enhancement recommendation method based on graph reasoning, characterized by The following steps are involved: S1: Given the interaction history of any user u, it is represented as an ordered sequence of item IDs [i1,i2,...,i T ], where each item i t Represents the item index of the user interaction, T represents the maximum sequence length, and the discrete ID sequence needs to be converted into a continuous vector representation through the embedding layer; a trainable item embedding matrix F∈R ∣V∣×d , where |V| represents the total capacity of the item library and d represents the embedding dimension. Through the table lookup operation, each item ID in the sequence is converted into a corresponding d-dimensional dense vector: v t =F[i t ,:] In this formula, v t Represents the vector of interactive items at time t, and obtains the user history sequence matrix X=[v1,v2,...,v T ] T ∈R T×d To capture temporal patterns and long-term dependencies in sequences, a Transformer-based encoder architecture is used. The core of the encoder is a multi-head self-attention mechanism that first linearly projects the input sequence into query, key, and value spaces: Q=XW Q K=XW K V=XW V In this formula, X represents the matrix of user history sequence, and the projection matrix Map the input to a dimension d k The feature space of , Q, K, V represent the query projection, key projection, and value projection in the self-attention mechanism, and the attention weight is calculated using the scaled dot product form: In this formula, d k Represents the feature dimension of each attention head. To ensure the correctness of the timing, a lower triangular mask matrix is introduced before the softmax to prevent information leakage. The output of the multi-head attention is concatenated and linearly transformed, and then residually connected with the original input and layer normalized: H = LayerNorm(X + MultiHead(X)) In this formula, X represents the matrix of the user's historical sequence, and the representation capability is further enhanced by a feed-forward network consisting of two linear transformations and a ReLU activation: FFN(X)=max(0,XW1+b1)W2+b2 In this formula, and represents the learnable parameter, d ff Set to 4d, b1 and b2 represent bias terms, and X represents the matrix of the user's historical sequence; after stacking L layers of Transformer blocks, the hidden state of the last position is taken as the final user representation: h T =Encoder(X)[-1,:] In this formula, h T The user state vector output by the encoder not only contains the user's historical behavior characteristics, but also encodes rich temporal dependencies; S2: Based on the encoded user state vector, similar historical behavior patterns are retrieved from the constructed memory library, and an approximate nearest neighbor search algorithm is used to achieve efficient retrieval; S3: The current user state, retrieved memory patterns, and candidate items are constructed into a ternary heterogeneous reasoning graph, where nodes include user nodes, memory nodes, and item nodes, and edge weights are dynamically calculated through an attention mechanism; S4: A multi-layer graph neural network is used to process the constructed reasoning graph, enabling information exchange between nodes through a message passing mechanism. The importance classification results of each memory node are output to complete the screening of memory patterns. S5: Based on the node weights output by the graph neural network, the candidate items are weighted and scored to generate the final recommendation list, while retaining the complete graph reasoning path as an interpretable basis for the recommendation decision; S6: Establish a reinforcement learning framework with recommendation effect indicators as reward functions, and optimize the parameters of memory retrieval and graph reasoning modules end-to-end through the policy gradient method.
2. A dynamic retrieval enhancement recommendation method based on graph reasoning according to claim 1, characterized in that The specific method of step S2 is: S2, dynamic memory retrieval uses efficient approximate nearest neighbor search technology to retrieve relevant historical behavior patterns from the memory library and convert the user state vector h T Normalized to a unit vector: In this formula, q represents the normalized user query vector, the memory bank M stores a large number of historical behavior patterns, and each memory item contains a key vector k i and the corresponding value vector v i , where the key vector has been L2 normalized; the memory bank can be expressed as: M={(k1,v1),(k2,v2),...,(k N ,v N )} In this formula, N represents the total number of memory items stored in the memory bank. To improve retrieval efficiency, an improved hierarchical navigable small-world algorithm is used to establish an index structure. During retrieval, the similarity score between the query vector and the memory key vector is first calculated: s i =q T k i In this formula, k i Represents the index key vector of the i-th memory item, s i Represents the original similarity between the query and the i-th memory key, and finds the top-K most relevant memory items through maximum inner product search: R={(k j ,v j )∣s j ∈Top-K({s1,...,s N },K)} In this formula, R is the memory item set of Top-K search results, s j represents the original similarity between the query and the jth memory key, N represents the total number of memory items stored in the memory library, and K is the number of retrievals, with a value of 5-20. To enhance the retrieval quality, a hybrid similarity metric is used: s' i =αq T k i +(1-α)cosine(f q ,f i ) In this formula, α∈[0,1] represents the balance factor, k i Represents the index key vector of the i-th memory item, s' i represents the final score after the mixed similarity calculation, cosine represents the cosine similarity between auxiliary features, and f q and f i They represent auxiliary feature vectors extracted from the user's current behavior and memory items respectively. The final retrieval results are sorted by score to form a memory set: In this formula, v j Represents the content vector of the jth memory item, M ret Represents the final search results, K represents the number of searches, and the weight coefficient W j Calculated by the softmax function: In this formula, s' i represents the final score after the hybrid similarity calculation, K is the number of retrievals, the temperature parameter τ controls the sharpness of the weight distribution, and the memory set M ret Each memory item is associated with the original user behavior fragment and context information.
3. A dynamic retrieval enhancement recommendation method based on graph reasoning according to claim 1, characterized in that The specific method in step S3 is: S3. Design a dynamic graph structure to model the complex relationship between users, memories and items based on the user state vector h. T and retrieval memory set M ret , initialize the graph node set: n u =LayerNorm(h T ) In this formula, n u Represents the user status node, The node representation of the kth candidate item, Y represents the number of items to be recommended, represents the jth retrieval memory node, v j represents the content vector of the jth memory item, Represents the set of candidate item nodes, which is obtained through the item embedding matrix E: k=1,...,Y In this formula, i k Represents the index of the item the user interacted with, Y represents the number of items to be recommended, and the edge is constructed using a multi-type attention mechanism. The edge weight between the user node and the memory node is calculated as follows: In this formula, W a ∈R d×d represents the trainable parameter matrix, e u,j Indicates the connection strength between the user node and the jth memory node, n u represents the user status node, K represents the number of retrievals, and the edge between the memory node and the item node is calculated by bilinear transformation: In this formula, W b ∈R d×d represents the weight matrix, b represents the bias term, σ represents the sigmoid activation function, e j,k Represents the connection strength between the memory node and the item node. The final constructed graph structure is expressed as: G=(V,E) In this formula, G represents the constructed heterogeneous reasoning graph, and the set E contains all connections with non-zero weights. To enhance the graph's expressiveness, the following auxiliary connections are introduced: Direct connection between user node and Top-N candidate items: In this formula, The node representation of the kth candidate item, n u Represents the user status node, e u,k Represents the direct connection weight between the user and the candidate item; Memory similarity connections between nodes: yes j,l =|(sim(v j ,v l )>θ) In this formula, | represents the indicator function, θ represents the similarity threshold, and v j Represents the content vector of the j-th memory item. The dynamic graph structure completely captures the explicit and implicit relationships between users, memories, and items.
4. The method for dynamic retrieval enhancement recommendation based on graph reasoning according to claim 1 is characterized in that The specific steps in step S4 are: S4. Use a multi-layer graph attention network to extract features from the constructed heterogeneous reasoning graph. For the heterogeneous reasoning graph G = (V, E), initialize the hidden state of each node: In this formula, Represents different specific projection matrices, represents the initial hidden state of the node, d h represents the hidden layer dimension of the graph neural network. The information transmission process adopts an improved multi-head graph attention mechanism. In the calculation of the lth layer, node v aggregates information from neighboring nodes: In this formula, N(v) represents the set of neighbor nodes of node v. Represents the transformation matrix of the l-th layer head h, represents the number of attention heads, and the attention coefficient Calculated as follows: In this formula, a h Represents the parameter vector of the h-th attention head, / / represents the vector splicing operation. To enhance the long-distance dependency capability, the network also includes cross-layer connections: In this formula, represents the node representation after layer l, and β represents the dynamic weight of the cross-layer connection, which is dynamically calculated through the gating mechanism: In this formula, a h represents the parameter vector of the h-th attention head, represents the node representation after layer l, w g and b g Represents the calculation parameters of the skip connection coefficient. After L layers of graph convolution, the importance of memory nodes is classified: In this formula, W c and b c Represents the classification parameter of memory node importance, P j represents the probability distribution of the processing strategy of memory node j, Represents the representation vector of the j-th node after passing through the l-layer graph neural network, Mapping node representations to a 4-dimensional classification space, corresponding to four processing strategies for memory nodes: strengthening, retaining, weakening, and ignoring; The memory node weights output by graph reasoning are calculated as: In this formula, Represents the adjusted importance weight of the memory node.
5. The method for dynamic retrieval enhancement recommendation based on graph reasoning according to claim 1 is characterized in that The specific steps in step S5 are: S5. Based on memory node weight and the representation of each node Calculate candidate item i k Recommendation score: In this formula, C k represents the recommendation score of the candidate item, Represents the representation of the user node, Represents item node i k The representation of represents the importance of the adjusted memory node, the weight c∈[0,1] represents the memory fusion coefficient, and K represents the number of retrievals. To enhance the diversity of recommendations, a category balance term is introduced: In this formula, C k represents the recommendation score of the candidate item, C' k represents the revised recommendation score of the candidate item, Indicates item i k The number of times the category appears in the current recommendation batch. f represents the diversity control parameter. After the score calculation is completed, it is converted into a probability distribution through the softmax function: In this formula, C' k represents the corrected recommendation score of the candidate item, p k represents the recommendation probability of item k, Y represents the number of items to be recommended, and parameter g adjusts the concentration of the recommendation list. The final recommendation list is generated in descending order of probability: R k ={(i k ,p k )∣p k ∈TopA{p1,...,p Y },A} In this formula, A represents the length of the recommendation list. To provide interpretability, the complete reasoning path is recorded: In this formula, J k Indicates that item i k The indexes of the top three memory nodes that contribute the most to the score.
6. A dynamic retrieval enhancement recommendation method based on graph reasoning according to claim 1, characterized in that The specific steps in step S5 are: S6. Define the state space as S, where each state g t Contains the current user representation Memory node status and candidate item representation The action space B corresponds to the generated recommendation list R k , the reward function is designed as: r t =η1·NDCG@10+η2·Diversity-η3·Redundancy In this formula, η1, η2, η3 represent weight coefficients, NDCG@10 represents the ranking quality of the top 10 recommended items, and r t represents the instant reward at time t, Diversity calculates the category entropy of the recommendation list, Redundancy measures the repetition with the previous recommendation, and the policy network π θ The parameter update adopts PPO algorithm: In this formula, ∈ represents the clipping range of the policy update, represents the policy parameters before update, π θ represents the parameterized model for generating recommendation strategies, a represents the recommended action, and the advantage function A t Calculated by generalized odds estimate: δ t =r t +nV φ (s t+1 )-V φ (g t ) In this formula, n represents the discount factor, δ t represents the value difference of adjacent states, z represents the GAE parameter, and the value function network V φ The parameter update target is: In this formula, R t Represents the cumulative return. To stabilize the training process, a double buffer mechanism is used to alternately update the policy network and the value network, and independent learning rates are set. The sample batches for each update are randomly sampled from the experience replay pool D: In this formula, a i represents the recommended action performed by the i-th sample, r i represents the immediate reward obtained by the i-th sample, g i represents the immediate reward obtained by the i-th sample, the batch size x is dynamically adjusted according to the resources, and the policy exploration is maintained through the regularization term: In this formula, a represents the recommended action, L ent represents the entropy of the policy probability distribution, π θ Represents a parameterized model for generating recommendation strategies.
Citation Information
Cited By
Intelligent control method and system based on state space
CN121091757A
Full-scene adaptive intelligent recommendation method and system based on graph neural network
CN121479067A