Self-adaptive recommendation method based on dynamic strategy optimization
Through the combination of temporal behavior coding, dynamic memory storage and strategy control modules, the multi-objective optimization conflicts of existing recommendation systems in dynamic scenarios are resolved, accurate and diversified personalized recommendations are achieved, and the real-time adaptability and efficiency of the recommendation system are improved.
Patent Information
- Application Number
- CN202510729364.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
Existing recommendation systems have shortcomings in terms of long-tail effect and diversity dilemma, real-time modeling of dynamic scenarios, multi-objective collaborative optimization and privacy protection, making it difficult to provide accurate, diverse and real-time personalized recommendations.
It adopts temporal behavior encoding module, dynamic memory storage library, policy control module and heterogeneous fusion module, and realizes dynamic capture of user interests and efficient fusion of multi-source information through unidirectional attention mechanism, dual-channel memory library and dual-way cross attention mechanism, and generates adaptive retrieval strategy.
It improves the accuracy and diversity of recommendations, solves the problems of user interest drift and long-tail recommendations, and improves the efficiency of online updates, providing users with more accurate and diverse recommendation results.
Smart Images

Figure CN120632212A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to adaptive recommendation based on dynamic strategy optimization, which solves the problem of insufficient accuracy of existing recommendations by designing temporal behavior coding and dynamic memory storage. Background Art
[0002] With the growth of internet information resources, users are increasingly facing information overload as they grapple with massive amounts of data. Recommender systems, as information filtering technologies, offer personalized services by analyzing user behavior and content features. They have become critical infrastructure for e-commerce, social media, and content platforms. Traditional recommendation methods primarily focus on content-based recommendations, collaborative filtering, and hybrid recommendations. Content-based methods rely on matching user and item features, generating recommendations through metadata analysis such as text and tags. While advantageous in avoiding the cold-start problem, they suffer from difficulties in feature extraction and the homogeneity of recommendation results. Collaborative filtering, which mines implicit associations within the user-item interaction matrix for prediction, has become a widely used recommendation paradigm. However, its performance is limited by data sparsity, the cold-start problem, and its difficulty handling heterogeneous information from multiple sources. Hybrid recommendations attempt to combine the advantages of content and collaborative filtering, but they still face bottlenecks in multimodal data processing and adaptability to dynamic scenarios.
[0003] In recent years, the introduction of deep learning technology has enhanced the expressive power of recommendation systems. Models based on deep neural networks, convolutional neural networks, and graph neural networks, through end-to-end learning of latent representations of users and items, can extract high-level features from multi-source data such as review text and images, and alleviate sparsity issues. The integration of attention mechanisms and Transformer architectures enables models to dynamically capture temporal changes in user interests, while graph neural networks, by modeling the topological structure of the user-item interaction graph, can mine implicit connections among long-tail items. Despite this, existing methods still face challenges such as insufficient dynamic adaptability, conflicts in multi-objective optimization, and the conflict between privacy protection and recommendation effectiveness.
[0004] Existing technologies have the following problems: 1) The long-tail effect and the diversity dilemma. Over-optimization of accuracy may lead to convergence of recommendation results and trigger the information cocoon effect; 2) Real-time modeling of dynamic scenarios. User interest drift and changes in item popularity require the model to have online update capabilities. Existing methods mostly rely on offline training, which makes it difficult to balance efficiency and effectiveness; 3) Multi-objective collaborative optimization. Requirements such as privacy protection, explainability, and fairness are inherently in conflict with recommendation performance. Existing solutions mostly adopt a single-objective priority strategy and lack a systematic balancing mechanism. Summary of the Invention
[0005] In response to the limitations and challenges of existing solutions, the present invention proposes an adaptive recommendation method based on dynamic strategy optimization. The present invention realizes accurate and real-time optimized personalized recommendations by designing a temporal behavior encoding module, a dynamic memory repository, a policy control module and a heterogeneous fusion module. Specifically, the temporal behavior encoding module of the present invention uses a unidirectional attention mechanism combined with a time decay characteristic to process the user's historical interaction sequence, generates a user state vector with time decay, and captures the dynamic changes of user interests; the dynamic memory repository automatically updates historical records based on a hybrid scoring mechanism to ensure the timeliness and representativeness of the memory library; the policy control module outputs the learnable retrieval quantity, weight coefficient and fusion parameter to form an adaptive retrieval strategy; the heterogeneous fusion module realizes cross-channel feature enhancement through a dual-path attention mechanism, further improving the accuracy and diversity of recommendations; the invention effectively solves the problems of preference drift and long-tail recommendations, while improving the efficiency of online updates, providing users with more accurate and diverse recommendation results.
[0006] The present invention provides an adaptive recommendation method based on dynamic strategy optimization, comprising the following steps:
[0007] S1: The temporal encoding process of user behavior is realized through the designed temporal attention network, which uses a unidirectional attention mechanism combined with the time decay feature to process the user's historical interaction sequence. Given a user interaction sequence X of length T = [x1, x2, ..., x T ], where the feature vector x at each time step t ∈R d Contains multiple information such as item embedding and interaction timestamp. T is the total length of the user's historical behavior sequence, d is the dimension of the feature vector, and the query vector, key vector, and value vector are generated through linear transformation:
[0008] Q t =x t W q
[0009] K t =x t W k
[0010] V t =x t W v
[0011] Among them, W q ,W k ,W v ∈R d×d represents the trainable parameter matrix, Q t ,K t ,V tRepresents the query, key, and value vector at the tth time step. To capture the time decay law in the behavior sequence, an exponential decay weight coefficient γ is designed. t-i Acting on historical key-value pairs, where γ∈(0,1] is an adjustable decay rate, i represents the historical time step index, and the attention score is:
[0012]
[0013] Among them, α t,i represents the attention weight of the t-th time step to the i-th historical step, ti represents the number of time interval steps, d represents the dimension of the feature vector, K i represents the key vector at historical time step i, It represents the dot product of the query and the key, reflecting the correlation between t and i. This makes the longer the historical behavior, the less impact it has on the current state. This is consistent with the characteristics of user interest drift in real scenarios. The time series representation is obtained by weighted aggregation of historical value vectors:
[0014]
[0015] Among them, h t Represents the time series representation vector of time step t, V i Represents the value vector of historical step i, α t,i Represents the attention weight of the t-th time step to the i-th historical step, and introduces residual connection and layer normalization operations in the output layer:
[0016]
[0017] Among them, h t represents the time series representation vector of time step t, represents the user state vector, x t Represents the user behavior feature vector at the tth time step, and the resulting user state vector It contains both the real-time information of the current interaction and the temporal dependencies of historical behaviors. Parameters are optimized through end-to-end training to ensure the adaptability of the encoding results to downstream tasks.
[0018] S2: Build a dual-channel memory library to store user history and associated item information. A time-based similarity hybrid scoring mechanism automatically eliminates low-value memory items, enabling real-time updates of the memory library.
[0019] S3: Analyze the current user status and memory metadata through the policy control network, output the adjustable number of retrievals, original representation weights, and cross-channel fusion coefficients, and form an adaptive retrieval strategy;
[0020] S4: A dual-path attention mechanism is used to interactively calculate user behavior memory and item memory, and dynamically weighted fusion is performed based on strategy parameters to generate enhanced user representation.
[0021] S5: Calculate candidate item matching scores based on the enhanced representation, generate personalized recommendation lists, and collect user feedback data for continuous optimization of the policy network;
[0022] S6: Utilize the grouped relative policy gradient algorithm to independently optimize each policy parameter group based on the recommendation effect feedback to achieve improvements during operation.
[0023] According to a specific implementation of an embodiment of the present invention, the specific steps of S2 are:
[0024] S2, dynamic memory storage and update mechanism realizes efficient historical pattern management through dual-channel ring buffer architecture, which consists of behavioral memory channel M behavior and target item channel M item The two channels adopt a synchronous update strategy to ensure data consistency. The current time step is t current , the time decay weight of the record is calculated by the exponential decay function:
[0025]
[0026] Where β represents the decay rate coefficient, t j represents the storage timestamp of the jth record in the memory, ΔT represents the time normalization factor, represents the time decay weight of the jth record. For any record j in the memory, the semantic activity is measured by calculating the cosine similarity with the most recent K newly added records:
[0027]
[0028] in, represents the semantic activity score of the jth record, N represents the total number of records stored in the current memory, K represents the size of the most recently added record window referenced when calculating semantic activity, and h j , h k Represents the user state vector of the j-th and k-th records in the memory bank. The comprehensive value evaluation function of the memory entry combines the two dimensions of time decay and semantic activity:
[0029]
[0030] Where Ψ(j) represents the comprehensive value score of the jth record, and α∈[0,1] represents the balance coefficient, which is dynamically adjusted by the trainable parameter matrix A:
[0031]
[0032] Where σ(·) represents the sigmoid function, h avg Represents the average vector of all current behaviors in the memory bank, h current Represents the current user state vector. When the memory reaches the preset upper threshold C max When , the memory compression operation is performed and the retention probability of all entries is calculated:
[0033]
[0034] in, represents the probability of retaining the jth record, Ψ(j) represents the comprehensive value score of the jth record, N represents the total number of records currently stored in the memory bank, and τ represents the temperature coefficient, which controls the strictness of the elimination strategy. The actual number of entries retained is determined by the adaptive threshold:
[0035]
[0036] Among them, N retain Indicates the number of entries actually retained, N indicates the total number of records stored in the memory, and the writing of new records adopts a fusion strategy with weight distribution. new With existing record h j When the similarity exceeds the threshold θ, a weighted update is performed:
[0037]
[0038] Among them, γ∈(0,1) represents the fusion coefficient, and the time decay weight Positive correlation, memory metadata maintenance includes updating the most recent access timestamp, recalculating the average similarity feature, and adjusting the time decay counter of each entry. It can automatically identify and strengthen frequently occurring user patterns, while gradually eliminating outdated behavioral features to maintain the timeliness and representativeness of the memory.
[0039] According to a specific implementation of an embodiment of the present invention, the specific steps of S3 are:
[0040] S3. The dynamic generation mechanism of retrieval strategy realizes parameterized decision-making through deep strategy network, taking the user’s current state vector and memory metadata features m t As the joint input, the intermediate representation layer of the policy network is constructed, and the two types of inputs are fused through the gated feature interaction mechanism:
[0041]
[0042] Among them, W g , W z represents the trainable weight matrix, bg ,b z represents the bias term, ⊙ represents element-by-element multiplication, z t represents the intermediate representation, g t Represents the intermediate gated feature vector of the policy network. The policy network generates three decision branches in parallel, and the retrieval quantity branch outputs the probability distribution of the discrete action space:
[0043] p k =softmax(W k z t +b k )
[0044] Among them, p k represents the probability distribution of the number of retrievals k, z t represents the intermediate representation, W k , b k Represents the weight matrix and bias term of the retrieval quantity branch; Gumbel-Softmax reparameterization technique is used to achieve differentiable sampling:
[0045] k t =argmax(logp k +G)
[0046] Among them, k t represents the number of retrievals finally selected, G is the independent and identically distributed Gumbel noise, and the weight coefficient branch generates continuous decisions through sigmoid activation:
[0047]
[0048] Among them, ρ t , η t Represents the weight coefficient and fusion coefficient, w ρ , w η Represents the weight vector of the weight coefficient and fusion coefficient branch, b ρ , b η Represents the bias term of the weight coefficient and fusion coefficient branch, z t Represents the intermediate representation. To enhance the exploration ability of the strategy, Gaussian noise is injected into the continuous decision during the training phase:
[0049]
[0050] Among them, ∈ ρ ,∈ η represents the Gaussian noise injected during the training phase, represents the final action parameter with noise, ρ t , η tRepresents the weight coefficient and fusion coefficient. The policy network implements nonlinear decision-making through a multi-layer perceptron structure, and the hidden layer uses the LeakyReLU activation function:
[0051]
[0052] in, represents the output of the lth hidden layer of the policy network, W (l) , b (l) ) represents the weight matrix and bias term of the lth layer, and the output layer uses the temperature coefficient τ to control the sharpness of the decision distribution:
[0053]
[0054] Among them, p final represents the final action probability distribution, It represents the output of the hidden layer of the policy network, and automatically learns the complex mapping relationship between user status and optimal retrieval strategy through end-to-end training, realizing the decision-making ability of dynamically adjusting the retrieval scope and intensity according to different user scenarios.
[0055] According to a specific implementation of an embodiment of the present invention, the specific steps of S4 are:
[0056] S4, realize the integration of multi-source information through the dual-path cross attention mechanism, and transform the dynamic parameter k t , ρ t and η t As a regulatory factor, given the Top-k retrieved from the memory bank t Behavioral Memory Collection and target item set Constructing behavioral memory cross-attention pathways:
[0057]
[0058] in, Represents the query, key, and value projection matrix of the behavioral memory attention mechanism, A behavior represents the behavior memory cross attention weight matrix, c behavior represents the weighted aggregation result of behavioral memory, d is the dimension of the feature vector, and the cross-attention path of the target item is constructed using asymmetric projection:
[0059]
[0060] Among them, A item represents the item cross attention weight matrix, Represents the query, key, and value projection matrix of the item attention mechanism, c item Represents the weighted aggregation result of items, M maskrepresents the attention mask matrix, V t Represents the target item set, attention mask matrix M mask Dynamically generated based on the time distance of memory entries:
[0061] M mask [j]=-∞·|(t current -t j >ΔT max )
[0062] Among them, t j Represents the storage timestamp of the jth memory record, t current represents the current time step, ΔT max Represents the maximum allowed time interval threshold, and the dual-path features are adaptively fused through a gating mechanism:
[0063] c fusion =g t ⊙c behavior +(1-g t )⊙c item
[0064] Among them, g t represents the intermediate gating feature vector of the policy network, c item Represents the weighted aggregation result of items, c behavior represents the weighted aggregation result of behavioral memory, c fusion Represents the result of dual-path feature fusion. The enhanced representation is combined with the original user state vector through residual connection, and the retention ratio is determined by the strategy parameter ρ t Precise control:
[0065]
[0066] Among them, ρ t represents the original representation retention weight of the policy network output, c fusion Represents the dual-path feature fusion result, Representation enhanced representation, through the differentiable attention mechanism to achieve selective fusion of memory information, to ensure that important historical patterns are strengthened while suppressing noise interference, the output enhanced representation not only retains the core features of the user's current state, but also combines the most relevant historical behavior information in the memory bank.
[0067] According to a specific implementation of an embodiment of the present invention, the specific steps of S5 are:
[0068] S5. Use a deep matching network to achieve accurate ranking of candidate items and build a multi-view matching function that considers the global similarity and local interaction characteristics between users and items:
[0069]
[0070] Among them, v j represents the feature vector of item j, represents the global matching score of item j, represents the enhanced representation, W g Represents the global matching weight matrix, which is used to capture the overall correlation between users and items. Local interaction features are extracted through element-by-element product and cascade operations:
[0071]
[0072] in, represents the local interaction feature vector of item j, v j Represents the feature vector of item j. The final calculation of the matching score uses a residual connection structure:
[0073]
[0074] Among them, s j represents the matching score of item j, w i The weight vector representing the local interaction feature. To eliminate the recommendation bias, an adaptive correction term based on item popularity is introduced:
[0075]
[0076] in, represents the corrected matching score, N j represents the historical exposure times of item j, s j represents the matching score of item j, μ N and σ N are the mean and standard deviation of logarithmic popularity, ζ is the correction strength coefficient, and the recommendation list generation adopts an entropy-based regularized sampling strategy:
[0077]
[0078] Where H(·) represents entropy, ε is the entropy regularization coefficient, and p -j represents the current recommendation probability distribution of other items, H(p -j ) represents the recommendation distribution after excluding item j, ∈ represents the entropy regularization coefficient, and the user feedback processing module designs a time-varying importance weighting mechanism:
[0079] w t =exp(-ν|tt feedback ∣)
[0080]
[0081] Among them, r t represents the user feedback signal, w trepresents the time decay weight, ν represents the decay rate coefficient, t feedback Indicates the timestamp of user feedback, t impression Indicates the timestamp of the first exposure of the item to the user. Represents the weighted feedback signal, r t Represents the original feedback signal of the user at time t. The feedback data is stored in the experience buffer after characterization processing:
[0082]
[0083] Among them, B represents the experience playback buffer, V cand represents the current candidate item set, s represents the item matching score vector, r t Represents the user feedback signal. To deal with the sparse feedback problem of cold-start items, a generalized reward estimator based on item features is designed:
[0084]
[0085] Among them, U and V represent the trainable parameter matrices, represents the estimated generalization reward for the cold-start item j, u represents the trainable vector, The concatenated vector representing item features and enhanced user representation is used to generate recommendation results and process feedback data through multi-level features and dynamic bias correction.
[0086] S6. For the number of retrievals k t , advantage function A k Adopting the temporal difference error based on state-action pairs:
[0087] A k (s t ,a t )=r t +γV φ (s t+1 )-V φ (s t )
[0088] Among them, s t+1 Represents the state vector of the next time step, s t represents the current state vector, a t Represents the specific value of the discrete action, V φ represents the state value function network, γ is the discount factor, and the discrete actions are sampled with a baseline when the policy gradient is updated:
[0089]
[0090] in, represents the gradient operator for the policy network parameters θ, Jk (θ) represents the objective function of the policy network, π θ represents the policy network, Represents the old policy network, E[·] represents the expected value, and for the weight coefficient ρ t and η t For these two consecutive actions, the advantage function introduces an action-dependent baseline:
[0091]
[0092] Among them, A c represents the advantage function of continuous actions, Q ψ represents the action-value function network, is the action sampled by the current policy, M represents the number of sampled actions, and the policy gradient of consecutive actions is calculated using the reparameterization technique:
[0093]
[0094] Among them, J c (θ) represents the objective function of the continuous action, J c (θ) represents the objective function of the continuous action, μ θ (s t ) represents the deterministic action value output by the policy network;
[0095] To stabilize the training process, design the strategy update constraints:
[0096]
[0097] Among them, D KL represents the Kullback-Leibler divergence, Represents the old policy network, δ represents the maximum allowed KL divergence threshold, and the constraint is implemented by the adaptive KL penalty coefficient β:
[0098]
[0099] Among them, α is the adjustment factor, β is the adaptive KL penalty coefficient, and the experience replay buffer adopts a priority sampling mechanism:
[0100] P(i)∝|δ i | ω +∈
[0101] Among them, δ i represents the temporal difference error of sample i, ω represents the priority index, ε represents a small constant, P(i) represents the priority sampling probability of sample i, and the network parameter update adopts the soft target update strategy:
[0102] φ'←τφ'+(1-τ)φ
[0103] ψ'←τψ'+(1-τ)ψ
[0104] Among them, φ' and ψ' represent the target network parameters, τ represents the temperature coefficient, and φ and ψ represent the value function parameters of the main network. Stable updates of strategy parameters are achieved through group optimization to ensure that the system maintains a balanced strategy during continuous learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0105] Figure 1 Flowchart of this method;
[0106] Figure 2 This is the architectural diagram of this method. DETAILED DESCRIPTION
[0107] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below with reference to examples and drawings.
[0108] As attached Figure 1 and attached Figure 2 As shown in FIG, an adaptive recommendation method based on dynamic strategy optimization includes the following steps:
[0109] Step 1: The temporal encoding process of user behavior is realized through the designed temporal attention network. The unidirectional attention mechanism is combined with the time decay feature to process the user's historical interaction sequence. Given a user interaction sequence X=[x1,x2,...,x T ], where the feature vector x at each time step t ∈R d Contains multiple information such as item embedding and interaction timestamp. T is the total length of the user's historical behavior sequence, d is the dimension of the feature vector, and the query vector, key vector, and value vector are generated through linear transformation:
[0110] Q t =x t W q
[0111] K t =x t W k
[0112] V t =x t W v
[0113] Among them, W q ,W k ,W v ∈R d×d represents the trainable parameter matrix, Q t ,K t ,Vt Represents the query, key, and value vector at the tth time step. To capture the time decay law in the behavior sequence, an exponential decay weight coefficient γ is designed. t-i Acting on historical key-value pairs, where γ∈(0,1] is an adjustable decay rate, i represents the historical time step index, and the attention score is:
[0114]
[0115] Among them, α t,i represents the attention weight of the t-th time step to the i-th historical step, ti represents the number of time interval steps, d represents the dimension of the feature vector, K i represents the key vector at historical time step i, It represents the dot product of the query and the key, reflecting the correlation between t and i. This makes the longer the historical behavior, the less impact it has on the current state. This is consistent with the characteristics of user interest drift in real scenarios. The time series representation is obtained by weighted aggregation of historical value vectors:
[0116]
[0117] Among them, h t Represents the time series representation vector of time step t, V i Represents the value vector of historical step i, α t,i Represents the attention weight of the t-th time step to the i-th historical step, and introduces residual connection and layer normalization operations in the output layer:
[0118]
[0119] Among them, h t represents the time series representation vector of time step t, represents the user state vector, x t Represents the user behavior feature vector at the tth time step, and the resulting user state vector It contains both the real-time information of the current interaction and the temporal dependencies of historical behaviors. Through end-to-end training and parameter optimization, it ensures the coordinated adaptability of the encoding results and downstream tasks.
[0120] Step 2: Dynamic memory storage and update mechanism realizes efficient historical pattern management through dual-channel ring buffer architecture, which consists of behavior memory channel M behavior and target item channel M item The two channels adopt a synchronous update strategy to ensure data consistency. The current time step is t current , the time decay weight of the record is calculated by the exponential decay function:
[0121]
[0122] Where β represents the decay rate coefficient, t j represents the storage timestamp of the jth record in the memory, ΔT represents the time normalization factor, represents the time decay weight of the jth record. For any record j in the memory, the semantic activity is measured by calculating the cosine similarity with the most recent K newly added records:
[0123]
[0124] in, represents the semantic activity score of the jth record, N represents the total number of records stored in the current memory, K represents the size of the most recently added record window referenced when calculating semantic activity, and h j , h k Represents the user state vector of the j-th and k-th records in the memory bank. The comprehensive value evaluation function of the memory entry combines the two dimensions of time decay and semantic activity:
[0125]
[0126] Where Ψ(j) represents the comprehensive value score of the jth record, and α∈[0,1] represents the balance coefficient, which is dynamically adjusted by the trainable parameter matrix A:
[0127]
[0128] Where σ(·) represents the sigmoid function, h avg Represents the average vector of all current behaviors in the memory bank, h current Represents the current user state vector. When the memory reaches the preset upper threshold C max When , the memory compression operation is performed and the retention probability of all entries is calculated:
[0129]
[0130] in, represents the probability of retaining the jth record, Ψ(j) represents the comprehensive value score of the jth record, N represents the total number of records currently stored in the memory bank, and τ represents the temperature coefficient, which controls the strictness of the elimination strategy. The actual number of entries retained is determined by the adaptive threshold:
[0131]
[0132] Among them, N retain Indicates the number of entries actually retained, N indicates the total number of records stored in the memory, and the writing of new records adopts a fusion strategy with weight distribution. new With existing record h jWhen the similarity exceeds the threshold θ, a weighted update is performed:
[0133]
[0134] Among them, γ∈(0,1) represents the fusion coefficient, and the time decay weight Positive correlation: Memory metadata maintenance includes updating the most recent access timestamp, recalculating average similarity features, and adjusting the time decay counter of each entry. This automatically identifies and enhances frequently occurring user patterns while gradually eliminating outdated behavioral features, maintaining the timeliness and representativeness of the memory.
[0135] Step 3: The dynamic generation mechanism of the retrieval strategy realizes parameterized decision-making through the deep policy network, taking the user's current state vector and memory metadata features m t As the joint input, the intermediate representation layer of the policy network is constructed, and the two types of inputs are fused through the gated feature interaction mechanism:
[0136]
[0137] Among them, W g , W z represents the trainable weight matrix, b g ,b z represents the bias term, ⊙ represents element-by-element multiplication, z t represents the intermediate representation, g t Represents the intermediate gated feature vector of the policy network. The policy network generates three decision branches in parallel, and the retrieval quantity branch outputs the probability distribution of the discrete action space:
[0138] p k =softmax(W k z t +b k )
[0139] Among them, p k represents the probability distribution of the number of retrievals k, z t represents the intermediate representation, W k , b k Represents the weight matrix and bias term of the retrieval quantity branch; Gumbel-Softmax reparameterization technique is used to achieve differentiable sampling:
[0140] k t =argmax(logp k +G)
[0141] Among them, k trepresents the number of retrievals finally selected, G is the independent and identically distributed Gumbel noise, and the weight coefficient branch generates continuous decisions through sigmoid activation:
[0142]
[0143] Among them, ρ t , η t Represents the weight coefficient and fusion coefficient, w ρ , w η Represents the weight vector of the weight coefficient and fusion coefficient branch, b ρ , b η Represents the bias term of the weight coefficient and fusion coefficient branch, z t Represents the intermediate representation. To enhance the exploration ability of the strategy, Gaussian noise is injected into the continuous decision during the training phase:
[0144]
[0145] Among them, ∈ ρ ,∈ η represents the Gaussian noise injected during the training phase, represents the final action parameter with noise, ρ t , η t Represents the weight coefficient and fusion coefficient. The policy network implements nonlinear decision-making through a multi-layer perceptron structure, and the hidden layer uses the LeakyReLU activation function:
[0146]
[0147] in, represents the output of the lth hidden layer of the policy network, W (l) , b (l) ) represents the weight matrix and bias term of the lth layer, and the output layer uses the temperature coefficient τ to control the sharpness of the decision distribution:
[0148]
[0149] Among them, p final represents the final action probability distribution, The output of the hidden layer of the policy network automatically learns the complex mapping relationship between user status and optimal retrieval strategy through end-to-end training, and realizes the decision-making ability to dynamically adjust the retrieval scope and intensity according to different user scenarios;
[0150] Step 4: Integrate multi-source information through a two-way cross attention mechanism and transform the dynamic parameter k t , ρ t and η t As a regulatory factor, given the Top-k retrieved from the memory bank tBehavioral Memory Collection and target item set Constructing behavioral memory cross-attention pathways:
[0151]
[0152] in, Represents the query, key, and value projection matrix of the behavioral memory attention mechanism, A behavior represents the behavior memory cross attention weight matrix, c behavior represents the weighted aggregation result of behavioral memory, d is the dimension of the feature vector, and the cross-attention path of the target item is constructed using asymmetric projection:
[0153]
[0154] Among them, A item represents the item cross attention weight matrix, Represents the query, key, and value projection matrix of the item attention mechanism, c item Represents the weighted aggregation result of items, M mask represents the attention mask matrix, V t Represents the target item set, attention mask matrix M mask Dynamically generated based on the time distance of memory entries:
[0155] M mask [j]=-∞·|(t current -t j >ΔT max )
[0156] Among them, t j Represents the storage timestamp of the jth memory record, t current represents the current time step, ΔT max Represents the maximum allowed time interval threshold, and the dual-path features are adaptively fused through a gating mechanism:
[0157] c fusion =g t ⊙c behavior +(1-g t )⊙c item
[0158] Among them, g t represents the intermediate gating feature vector of the policy network, c item Represents the weighted aggregation result of items, c behavior represents the weighted aggregation result of behavioral memory, c fusion Represents the result of dual-path feature fusion. The enhanced representation is combined with the original user state vector through residual connection, and the retention ratio is determined by the strategy parameter ρ t Precise control:
[0159]
[0160] Among them, ρ t represents the original representation retention weight of the policy network output, c fusion Represents the dual-path feature fusion result, Enhanced representation: This uses a differentiable attention mechanism to selectively integrate memory information, ensuring that important historical patterns are reinforced while suppressing noise interference. The output enhanced representation retains the core features of the user's current state while incorporating the most relevant historical behavior information from the memory bank.
[0161] Step 5: Use a deep matching network to achieve accurate ranking of candidate items and construct a multi-view matching function that considers the global similarity and local interaction characteristics between users and items:
[0162]
[0163] Among them, v j represents the feature vector of item j, represents the global matching score of item j, represents the enhanced representation, W g Represents the global matching weight matrix, which is used to capture the overall correlation between users and items. Local interaction features are extracted through element-by-element product and cascade operations:
[0164]
[0165] in, represents the local interaction feature vector of item j, v j Represents the feature vector of item j. The final calculation of the matching score uses a residual connection structure:
[0166]
[0167] Among them, s j represents the matching score of item j, w i The weight vector representing the local interaction feature. To eliminate the recommendation bias, an adaptive correction term based on item popularity is introduced:
[0168]
[0169] in, represents the corrected matching score, N j represents the historical exposure times of item j, s j represents the matching score of item j, μ N and σ Nare the mean and standard deviation of logarithmic popularity, ζ is the correction strength coefficient, and the recommendation list generation adopts an entropy-based regularized sampling strategy:
[0170]
[0171] Where H(·) represents entropy, ε is the entropy regularization coefficient, and p -j represents the current recommendation probability distribution of other items, H(p -j ) represents the recommendation distribution after excluding item j, ∈ represents the entropy regularization coefficient, and the user feedback processing module designs a time-varying importance weighting mechanism:
[0172] w t =exp(-ν|tt feedback ∣)
[0173]
[0174] Among them, r t represents the user feedback signal, w t represents the time decay weight, ν represents the decay rate coefficient, t feedback Indicates the timestamp of user feedback, t impression Indicates the timestamp of the first exposure of the item to the user. Represents the weighted feedback signal, r t Represents the original feedback signal of the user at time t. The feedback data is stored in the experience buffer after characterization processing:
[0175]
[0176] Among them, B represents the experience playback buffer, V cand represents the current candidate item set, s represents the item matching score vector, r t Represents the user feedback signal. To deal with the sparse feedback problem of cold-start items, a generalized reward estimator based on item features is designed:
[0177]
[0178] Among them, U and V represent the trainable parameter matrices, represents the estimated generalization reward for the cold-start item j, u represents the trainable vector, The concatenated vector representing item features and enhanced user representations is used to generate recommendation results and process feedback data through multi-level features and dynamic bias correction.
[0179] Step 6: For the number of retrievals k t , advantage function A k Adopting the temporal difference error based on state-action pairs:
[0180] A k (s t ,a t )=r t +γV φ (s t+1 )-V φ (s t )
[0181] Among them, s t+1 Represents the state vector of the next time step, s t represents the current state vector, a t Represents the specific value of the discrete action, V φ represents the state value function network, γ is the discount factor, and the discrete actions are sampled with a baseline when the policy gradient is updated:
[0182]
[0183] in, represents the gradient operator for the policy network parameters θ, J k (θ) represents the objective function of the policy network, π θ represents the policy network, Represents the old policy network, E[·] represents the expected value, and for the weight coefficient ρ t and η t For these two consecutive actions, the advantage function introduces an action-dependent baseline:
[0184]
[0185] Among them, A c represents the advantage function of continuous actions, Q ψ represents the action-value function network, is the action sampled by the current policy, M represents the number of sampled actions, and the policy gradient of consecutive actions is calculated using the reparameterization technique:
[0186]
[0187] Among them, J c (θ) represents the objective function of the continuous action, J c (θ) represents the objective function of the continuous action, μ θ (s t ) represents the deterministic action value output by the policy network;
[0188] To stabilize the training process, design the strategy update constraints:
[0189]
[0190] Among them, D KLrepresents the Kullback-Leibler divergence, Represents the old policy network, δ represents the maximum allowed KL divergence threshold, and the constraint is implemented by the adaptive KL penalty coefficient β:
[0191]
[0192] Among them, α is the adjustment factor, β is the adaptive KL penalty coefficient, and the experience replay buffer adopts a priority sampling mechanism:
[0193] P(i)∝|δ i | ω +∈
[0194] Among them, δ i represents the temporal difference error of sample i, ω represents the priority index, ε represents a small constant, P(i) represents the priority sampling probability of sample i, and the network parameter update adopts the soft target update strategy:
[0195] φ'←τφ'+(1-τ)φ
[0196] ψ'←τψ'+(1-τ)ψ
[0197] Among them, φ' and ψ' represent the target network parameters, τ represents the temperature coefficient, and φ and ψ represent the value function parameters of the main network. Stable updates of strategy parameters are achieved through group optimization to ensure that the system maintains a balanced strategy during continuous learning.
Claims
1. An adaptive recommendation method based on dynamic strategy optimization, characterized by The following steps are involved: S1: The temporal encoding process of user behavior is realized through the designed temporal attention network, which uses a unidirectional attention mechanism combined with the time decay feature to process the user's historical interaction sequence. Given a user interaction sequence X of length T = [x1, x2, ..., x T ], where the feature vector x at each time step t ∈R d Contains multiple information such as item embedding and interaction timestamp. T is the total length of the user's historical behavior sequence, d is the dimension of the feature vector, and the query vector, key vector, and value vector are generated through linear transformation: Q t =x t W q K t =x t W k V t =x t W v Among them, W q ,W k ,W v ∈R d×d represents the trainable parameter matrix, Q t ,K t ,V t Represents the query, key, and value vector at the tth time step. To capture the time decay law in the behavior sequence, an exponential decay weight coefficient γ is designed. t-i Acting on historical key-value pairs, where γ∈(0,1] is an adjustable decay rate, i represents the historical time step index, and the attention score is: Among them, α t,i represents the attention weight of the t-th time step to the i-th historical step, ti represents the number of time interval steps, d represents the dimension of the feature vector, K i represents the key vector at historical time step i, It represents the dot product of the query and the key, reflecting the correlation between t and i. This makes the longer the historical behavior, the less impact it has on the current state. This is consistent with the characteristics of user interest drift in real scenarios. The time series representation is obtained by weighted aggregation of historical value vectors: Among them, h t Represents the time series representation vector of time step t, V i Represents the value vector of historical step i, α t,i Represents the attention weight of the t-th time step to the i-th historical step, and introduces residual connection and layer normalization operations in the output layer: Among them, h t represents the time series representation vector of time step t, represents the user state vector, x t Represents the user behavior feature vector at the tth time step, and the resulting user state vector It contains both the real-time information of the current interaction and the temporal dependencies of historical behaviors. Through end-to-end training and parameter optimization, it ensures the coordinated adaptability of the encoding results and downstream tasks. S2: Build a dual-channel memory library to store user history and associated item information. A time-based similarity hybrid scoring mechanism automatically eliminates low-value memory items, enabling real-time updates of the memory library. S3: Analyze the current user status and memory metadata through the policy control network, output the adjustable number of retrievals, original representation weights, and cross-channel fusion coefficients, and form an adaptive retrieval strategy; S4: A dual-path attention mechanism is used to interactively calculate user behavior memory and item memory, and dynamically weighted fusion is performed based on strategy parameters to generate enhanced user representation. S5: Calculate candidate item matching scores based on the enhanced representation, generate personalized recommendation lists, and collect user feedback data for continuous optimization of the policy network; S6: Utilize the grouped relative policy gradient algorithm to independently optimize each policy parameter group based on the recommendation effect feedback to achieve improvements during operation.
2. The adaptive recommendation method based on dynamic strategy optimization according to claim 1, characterized in that The specific method of step S2 is: S2, dynamic memory storage and update mechanism realizes efficient historical pattern management through dual-channel ring buffer architecture, which consists of behavioral memory channel M behavior and target item channel M item The two channels adopt a synchronous update strategy to ensure data consistency. The current time step is t current , the time decay weight of the record is calculated by the exponential decay function: Where β represents the decay rate coefficient, t j represents the storage timestamp of the jth record in the memory, ΔT represents the time normalization factor, represents the time decay weight of the jth record. For any record j in the memory, the semantic activity is measured by calculating the cosine similarity with the most recent K newly added records: in, represents the semantic activity score of the jth record, N represents the total number of records stored in the current memory, K represents the size of the most recently added record window referenced when calculating semantic activity, and h j , h k Represents the user state vector of the j-th and k-th records in the memory bank. The comprehensive value evaluation function of the memory entry combines the two dimensions of time decay and semantic activity: Where Ψ(j) represents the comprehensive value score of the jth record, and α∈[0,1] represents the balance coefficient, which is dynamically adjusted by the trainable parameter matrix A: Where σ(·) represents the sigmoid function, h avg Represents the average vector of all current behaviors in the memory bank, h current Represents the current user state vector. When the memory reaches the preset upper threshold C max When , the memory compression operation is performed and the retention probability of all entries is calculated: in, represents the probability of retaining the jth record, Ψ(j) represents the comprehensive value score of the jth record, N represents the total number of records currently stored in the memory bank, and τ represents the temperature coefficient, which controls the strictness of the elimination strategy. The actual number of entries retained is determined by the adaptive threshold: Among them, N retain Indicates the number of entries actually retained, N indicates the total number of records stored in the memory, and the writing of new records adopts a fusion strategy with weight distribution. new With existing record h j When the similarity exceeds the threshold θ, a weighted update is performed: Among them, γ∈(0,1) represents the fusion coefficient, and the time decay weight Positive correlation, memory metadata maintenance includes updating the most recent access timestamp, recalculating the average similarity feature, and adjusting the time decay counter of each entry. It can automatically identify and strengthen frequently occurring user patterns, while gradually eliminating outdated behavioral features to maintain the timeliness and representativeness of the memory.
3. The adaptive recommendation method based on dynamic strategy optimization according to claim 1, characterized in that The specific method in step S3 is: S3. The dynamic generation mechanism of retrieval strategy realizes parameterized decision-making through deep strategy network, taking the user’s current state vector and memory metadata features m t As the joint input, the intermediate representation layer of the policy network is constructed, and the two types of inputs are fused through the gated feature interaction mechanism: Among them, W g , W z represents the trainable weight matrix, b g ,b z represents the bias term, ⊙ represents element-by-element multiplication, z t represents the intermediate representation, g t Represents the intermediate gated feature vector of the policy network. The policy network generates three decision branches in parallel, and the retrieval quantity branch outputs the probability distribution of the discrete action space: p k =softmax(W k z t +b k ) Among them, p k represents the probability distribution of the number of retrievals k, z t represents the intermediate representation, W k , b k Represents the weight matrix and bias term of the retrieval quantity branch; Gumbel-Softmax reparameterization technique is used to achieve differentiable sampling: k t =argmax(logp k +G) Among them, k t represents the number of retrievals finally selected, G is the independent and identically distributed Gumbel noise, and the weight coefficient branch generates continuous decisions through sigmoid activation: Among them, ρ t , η t Represents the weight coefficient and fusion coefficient, w ρ , w η Represents the weight vector of the weight coefficient and fusion coefficient branch, b ρ , b η Represents the bias term of the weight coefficient and fusion coefficient branch, z t Represents the intermediate representation. To enhance the exploration ability of the strategy, Gaussian noise is injected into the continuous decision during the training phase: Among them, ∈ ρ ,∈ η represents the Gaussian noise injected during the training phase, represents the final action parameter with noise, ρ t , η t Represents the weight coefficient and fusion coefficient. The policy network implements nonlinear decision-making through a multi-layer perceptron structure, and the hidden layer uses the LeakyReLU activation function: in, represents the output of the lth hidden layer of the policy network, W (l) , b (l) ) represents the weight matrix and bias term of the lth layer, and the output layer uses the temperature coefficient τ to control the sharpness of the decision distribution: Among them, p final represents the final action probability distribution, It represents the output of the hidden layer of the policy network, and automatically learns the complex mapping relationship between user status and optimal retrieval strategy through end-to-end training, realizing the decision-making ability of dynamically adjusting the retrieval scope and intensity according to different user scenarios.
4. The adaptive recommendation method based on dynamic strategy optimization according to claim 1, characterized in that The specific steps in step S4 are: S4, realize the integration of multi-source information through the dual-path cross attention mechanism, and transform the dynamic parameter k t , ρ t and η t As a regulatory factor, given the Top-k retrieved from the memory bank t Behavioral Memory Collection and target item set Constructing behavioral memory cross-attention pathways: in, Represents the query, key, and value projection matrix of the behavioral memory attention mechanism, A behavior represents the behavior memory cross attention weight matrix, c behavior represents the weighted aggregation result of behavioral memory, d is the dimension of the feature vector, and the cross-attention path of the target item is constructed using asymmetric projection: Among them, A item represents the item cross attention weight matrix, Represents the query, key, and value projection matrix of the item attention mechanism, c item Represents the weighted aggregation result of items, M mask represents the attention mask matrix, V t Represents the target item set, attention mask matrix M mask Dynamically generated based on the time distance of memory entries: M mask [j]=-∞·|(t current -t j >ΔT max ) Among them, t j Represents the storage timestamp of the jth memory record, t current represents the current time step, ΔT max Represents the maximum allowed time interval threshold, and the dual-path features are adaptively fused through a gating mechanism: c fusion =g t ⊙c behavior +(1-g t )⊙c item Among them, g t represents the intermediate gating feature vector of the policy network, c item Represents the weighted aggregation result of items, c behavior represents the weighted aggregation result of behavioral memory, c fusion Represents the result of dual-path feature fusion. The enhanced representation is combined with the original user state vector through residual connection, and the retention ratio is determined by the strategy parameter ρ t Precise control: Among them, ρ t represents the original representation retention weight of the policy network output, c fusion Represents the dual-path feature fusion result, Representation enhanced representation, through the differentiable attention mechanism to achieve selective fusion of memory information, to ensure that important historical patterns are strengthened while suppressing noise interference, the output enhanced representation not only retains the core features of the user's current state, but also combines the most relevant historical behavior information in the memory bank.
5. The adaptive recommendation method based on dynamic strategy optimization according to claim 1, characterized in that The specific steps in step S5 are: S5. Use a deep matching network to achieve accurate ranking of candidate items and build a multi-view matching function that considers the global similarity and local interaction characteristics between users and items: Among them, v j represents the feature vector of item j, represents the global matching score of item j, represents the enhanced representation, W g Represents the global matching weight matrix, which is used to capture the overall correlation between users and items. Local interaction features are extracted through element-by-element product and cascade operations: in, represents the local interaction feature vector of item j, v j Represents the feature vector of item j. The final calculation of the matching score uses a residual connection structure: Among them, s j represents the matching score of item j, w i The weight vector representing the local interaction feature. To eliminate the recommendation bias, an adaptive correction term based on item popularity is introduced: in, represents the corrected matching score, N j represents the historical exposure times of item j, s j represents the matching score of item j, μ N and σ N are the mean and standard deviation of logarithmic popularity, ζ is the correction strength coefficient, and the recommendation list generation adopts an entropy-based regularized sampling strategy: Where H(·) represents entropy, ε is the entropy regularization coefficient, and p -j represents the current recommendation probability distribution of other items, H(p -j ) represents the recommendation distribution after excluding item j, ∈ represents the entropy regularization coefficient, and the user feedback processing module is designed with a time-varying importance weighting mechanism: w t =exp(-ν∣t-t feedback ∣) Among them, r t represents the user feedback signal, w t represents the time decay weight, ν represents the decay rate coefficient, t feedback Indicates the timestamp of user feedback, t impression Indicates the timestamp of the first exposure of the item to the user. Represents the weighted feedback signal, r t Represents the original feedback signal of the user at time t. The feedback data is stored in the experience buffer after characterization processing: Among them, B represents the experience playback buffer, V cand represents the current candidate item set, s represents the item matching score vector, r t Represents the user feedback signal. To deal with the sparse feedback problem of cold-start items, a generalized reward estimator based on item features is designed: Among them, U and V represent the trainable parameter matrices, represents the estimated generalization reward for the cold-start item j, u represents the trainable vector, The concatenated vector representing item features and enhanced user representation is used to generate recommendation results and process feedback data through multi-level features and dynamic bias correction.
6. The adaptive recommendation method based on dynamic strategy optimization according to claim 1, characterized in that The specific steps in step S5 are: S6. For the number of retrievals k t , advantage function A k Adopting the temporal difference error based on state-action pairs: From k (s t ,a t )=r t +γV φ (s t+1 )-V φ (s t ) Among them, s t+1 represents the state vector of the next time step, s t represents the current state vector, a t Represents the specific value of the discrete action, V φ represents the state value function network, γ is the discount factor, and the discrete actions are sampled with a baseline when the policy gradient is updated: in, represents the gradient operator for the policy network parameter θ, J k (θ) represents the objective function of the policy network, π θ represents the policy network, Represents the old policy network, E[·] represents the expected value, and for the weight coefficient ρ t and η t For these two consecutive actions, the advantage function introduces an action-dependent baseline: Among them, A c represents the advantage function of continuous actions, Q ψ represents the action-value function network, is the action sampled by the current policy, M represents the number of sampled actions, and the policy gradient of consecutive actions is calculated using the reparameterization technique: Among them, J c (θ) represents the objective function of the continuous action, J c (θ) represents the objective function of the continuous action, μ θ (s t ) represents the deterministic action value output by the policy network; To stabilize the training process, design the strategy update constraints: Among them, D KL represents the Kullback-Leibler divergence, Represents the old policy network, δ represents the maximum allowed KL divergence threshold, and the constraint is implemented by the adaptive KL penalty coefficient β: Among them, α is the adjustment factor, β is the adaptive KL penalty coefficient, and the experience replay buffer adopts a priority sampling mechanism: P(i)∝|δ i | ω +∈ Among them, δ i represents the temporal difference error of sample i, ω represents the priority index, ε represents a small constant, P(i) represents the priority sampling probability of sample i, and the network parameter update adopts the soft target update strategy: φ'←τφ'+(1-τ)φ ψ'←τψ'+(1-τ)ψ Among them, φ' and ψ' represent the target network parameters, τ represents the temperature coefficient, and φ and ψ represent the value function parameters of the main network. Stable updates of strategy parameters are achieved through group optimization to ensure that the system maintains a balanced strategy during continuous learning.
Citation Information
Cited By
Instrument full life cycle management method and system based on Internet of Things
CN120996793A
Lightweight cross-domain recommendation method and system based on user alignment Agent drive
CN121030097A
Strategy optimization method and device and storage medium
CN121094170A
Personalized article recommendation method and device, electronic equipment and storage medium
CN121434501A
Full-scene adaptive intelligent recommendation method and system based on graph neural network
CN121479067A