A method and system for optimizing recommendation strategies based on user behavior models
By building a user behavior model based on the generative adversarial imitation learning algorithm and combining the reinforcement learning algorithm PPO, the video recommendation strategy is optimized, and the problem of medium- and long-term feedback indicator optimization in the existing technology is solved, the trial and error cost is reduced, and the realization of real-time and long-term indicator improvement and strategy interpretability is achieved.
Patent Information
- Application Number
- CN202210537164.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The existing video recommendation strategy optimization methods are difficult to effectively optimize long-term feedback indicators, and the trial and error cost of reinforcement learning in real recommendation systems is high, resulting in the failure to fully realize the potential.
Based on the generative adversarial imitation learning algorithm (GAIL), the user behavior model is constructed, combined with the reinforcement learning algorithm PPO, and the recommendation strategy is optimized from offline interactive data between the user and the video recommendation system, reducing trial and error costs, and improving real-time and long-term interaction indicators.
It significantly improves the instant interaction indicators and long-term interaction indicators of the recommendation strategy, improves the interpretability and optimization sustainability of the recommendation strategy, and reduces the unpleasant experience and loss risks brought by the suboptimal strategy.
Smart Images

Figure CN114911969B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for optimizing a recommendation strategy based on a user behavior model, and belongs to the technical field of system data processing. Background Art
[0002] Short videos are becoming increasingly closely linked to modern people's lives. A mainstream short video recommendation system needs to make tens of billions of decisions every day to meet the personalized needs of hundreds of millions of users. Whether in terms of user scale, video quantity, or technical depth, the recommendation system faces greater challenges. Modern short video recommendation systems usually screen out videos that match the user's interests from a large number of background video pools through a series of processes such as recall, rough ranking, fine ranking, re-ranking, and manual rule filtering, and recommend them to users. Existing video recommendation strategies mainly predict the user's immediate feedback metrics (such as viewing time, click-through rate, etc.) through deep learning models, and optimize these immediate feedback metrics using supervised learning methods. Existing video recommendation strategy optimization methods usually indirectly optimize long-term feedback metrics (such as user satisfaction and retention rate, etc.) by optimizing short-term feedback metrics. However, due to the sparsity and latency of long-term feedback metrics, it is often difficult to accurately measure the correlation between short-term feedback metrics and long-term feedback metrics. Therefore, this indirect optimization method faces bottlenecks.
[0003] Compared with supervised learning that focuses on the prediction ability of the model, reinforcement learning pays more attention to the sequential decision-making ability of the model. By describing the continuous interaction between the recommendation system and the user as a sequential decision-making process and maximizing the cumulative reward in the entire decision-making process, reinforcement learning technology can directly optimize such long-term feedback metrics with sparsity and latency. However, the reinforcement learning training process relies on a large number of trials and errors, and the cost of conducting trial-and-error experiments on a real recommendation system is extremely high, making it difficult to meet business requirements. Therefore, existing recommendation strategy optimization schemes based on reinforcement learning technology usually adopt degraded or compromised methods, resulting in the potential of reinforcement learning being difficult to be fully exploited and exerted. Summary of the Invention
[0004] Objective of the Invention: Aiming at the problems and deficiencies in the prior art, the present invention provides a method and system for constructing a user behavior model from the offline interaction data between users and video recommendation systems and optimizing short-video recommendation strategies based on the user behavior model. Different from the existing methods that directly optimize short-video recommendation strategies from offline data to improve the immediate interaction metrics and long-term interaction metrics of the recommendation strategies, a recommendation strategy with good generalization and stronger robustness is trained on a user behavior model with a low-cost trial-and-error cost based on reinforcement learning technology, effectively reducing the poor experience and loss risk caused by users directly interacting with sub-optimal recommendation strategies. At the same time, the constructed user behavior model can be used for the metric evaluation of the recommendation strategy, improving the interpretability and optimization sustainability of the recommendation strategy.
[0005] Technical Solution: A method for optimizing a recommendation strategy based on a user behavior model, which constructs a user behavior model that can reflect user behavior preferences from the offline interaction data between users and video recommendation systems based on the Generative Adversarial Imitation Learning algorithm (GAIL). By allowing a reinforcement learning agent to interact with the user behavior model to collect data and optimizing the relevant metrics of the video recommendation strategy based on the Proximal Policy Optimization (PPO) algorithm to obtain an optimal recommendation strategy, the cost of directly trial-and-error of reinforcement learning on the recommendation system is reduced, and the immediate interaction metrics (such as viewing time, click-through rate, like rate, etc.) and long-term interaction metrics (such as user satisfaction and retention rate, etc.) of the recommendation strategy are significantly improved. The optimal recommendation strategy is deployed to a real recommendation system for online evaluation. If the relevant metrics do not meet the system requirements, new interaction data is continuously collected and the user simulator construction process and recommendation strategy optimization process are repeated until the relevant metrics of the recommendation strategy meet the system requirements.
[0006] The specific process is as follows:
[0007] 1) Generate an offline user-recommendation system interaction dataset. Retrieve the interaction data of users within a period of time from the log system of the recommendation system. Each piece of data includes: current timestamp, user ID, user's viewing history video list, recommended video information, and user click feedback information. After sorting the corresponding interaction data of the same user ID according to the timestamp, the interaction trajectory data between the user and the recommendation system is obtained, and the interaction trajectory data constitutes the user-recommendation system interaction dataset;
[0008] 2) Train the user behavior model using the Generative Adversarial Imitation Learning algorithm (GAIL); by simultaneously modeling the user's explicit feedback such as click behavior and implicit feedback such as the interval time of the user's next request, a more accurate prediction result of the user behavior feedback is obtained. The training steps are as follows:
[0009] Step 1: Initialize the user behavior network, recommendation strategy network, and discriminator D.
[0010] Step 2: Sample a batch of data from the user-recommendation system interaction dataset. Each piece of data is the starting point of a trajectory, which contains the corresponding timestamp information, the user's click history list, and the user's click feedback information at the previous moment. The user's click history list is converted into the corresponding user click history state through the Embedding Model.
[0011] Step 3: Input the timestamp information, the user click history state, and the user's click feedback at the previous moment into the recommendation policy network to generate the corresponding candidate video information.
[0012] Step 4: Input the user click history state and the candidate videos into the user behavior network to obtain the user's click feedback information on the candidate videos and the interval time information for the next request.
[0013] Step 5: Adding the current timestamp to the interval time information for the next request can obtain the timestamp information for the next request. Iterate steps 3 to 5 to generate a batch of trajectory data D of the interaction between the user and the recommendation policy network g ;
[0014] Step 6: Update the discriminator parameters. Sample a batch of real interaction trajectory data D of the user and the recommendation system from the historical offline interaction dataset of the real user and the video recommendation system r , and input the generated trajectory data D g into the discriminator D at the same time, and optimize the following objective loss function:
[0015]
[0016] where τ represents the trajectory in the dataset, log represents taking the natural logarithm with base e, and the discriminator D maximizes the discriminator score under the real interaction trajectory data D of the user and the recommendation system r , and at the same time minimizes the discriminator score under the generated interaction trajectory data D of the user and the recommendation policy network g . Eventually, the discriminator D can distinguish as much as possible whether the trajectory comes from the real interaction trajectory data D of the user and the recommendation system r or from the interaction trajectory data D of the user and the recommendation policy network generated by the interaction between the user behavior network and the recommendation policy network g .
[0017] Step 7: Update the user behavior network parameters. Iterate steps 2 to 5 to generate a batch of trajectory datasets {τ 1 , τ 2 , …, τ N} generated by the interaction between the user behavior network and the recommendation policy network, and the optimization objective is to generate the discounted cumulative reward on the interaction trajectory data:
[0018]
[0019] Among them, is the discount coefficient of the reward, usually set as a real number between (0, 1]. The reward at time t r t = logD(τ t ) is set as the natural logarithm transformation value of the score output by the discriminator for the trajectory τ t .
[0020] Step 8: Repeat Steps 2 - 7 until the loss function of the discriminator D converges or reaches a given number of training times.
[0021] Step 9: Output the final user behavior network as the user behavior model, and the training process ends.
[0022] 3) Train the recommendation strategy. Initialize a recommendation strategy, interact with the above - trained user behavior model to collect data, and use the reinforcement learning algorithm PPO to optimize the relevant metrics of the recommendation strategy until convergence or reaching a given number of training times, and output the optimal recommendation strategy.
[0023] 4) Deploy and evaluate the optimal recommendation strategy. Deploy the above - mentioned optimal recommendation strategy to a real - world recommendation system, and use online data to evaluate whether the interaction metrics of the recommendation strategy meet the requirements of the system.
[0024] 5) If the result of the above online evaluation does not meet the requirements of the system, continue to collect new user - recommendation system interaction data, and repeat processes 1) - 4) until the relevant metrics of the recommendation strategy meet the requirements of the system.
[0025] In the method of the present invention, the video recommendation strategy directly interacts with the above - trained user behavior model to collect data, and uses the reinforcement learning algorithm PPO to update the recommendation strategy until the algorithm converges or reaches a given number of training times, and outputs the optimal recommendation strategy on the user behavior model.
[0026] In the method of the present invention, the above - mentioned optimal recommendation strategy is directly deployed to a real - world video recommendation system, and the interaction metrics of the recommendation strategy are evaluated through online data to determine whether they meet the requirements of the system. If not, continue to collect new user - recommendation system interaction data. The new interaction data can be used as offline data for further training of the user behavior model and the optimal recommendation strategy, iteratively improving the strategy performance until the relevant metrics of the trained recommendation strategy meet the requirements of the system.
[0027] A recommendation strategy optimization system based on a user behavior model, comprising:
[0028] Dataset Generation Module: It is used to generate an offline user-recommendation system interaction dataset. It retrieves the interaction data of users within a period of time from the log system of the recommendation system. Each piece of data includes: current timestamp, user ID, list of historical videos watched by the user, recommended video information, and user click feedback information. After sorting the corresponding interaction data of the same user ID according to the timestamp, the interaction trajectory data of the user and the recommendation system is obtained, and the interaction trajectory data constitutes the user-recommendation system interaction dataset.
[0029] User Behavior Model Training Module: It uses the Generative Adversarial Imitation Learning algorithm (GAIL) to train the user behavior model. By simultaneously modeling the explicit feedback of users such as click behavior and implicit feedback such as the interval time of the user's next request, a more accurate prediction result of the user behavior feedback is obtained.
[0030] Recommendation Strategy Training Module: It initializes a recommendation strategy, interacts with the above-trained user behavior model to collect data, and uses the Proximal Policy Optimization (PPO) algorithm of reinforcement learning to optimize the relevant metrics of the recommendation strategy until convergence or reaching a given number of training times, and outputs the optimal recommendation strategy.
[0031] Optimal Recommendation Strategy Deployment Module: It deploys the optimal recommendation strategy trained from the above Recommendation Strategy Training Module to the real recommendation system to replace the original recommendation strategy in the recommendation system and interact with real users;
[0032] Evaluation Module: It uses online data to evaluate whether the interaction metrics of the recommendation strategy meet the requirements of the system. If not, it continues to collect new user-recommendation system interaction data and merge it with the dataset of the Dataset Generation Module, and execute the User Behavior Model Training Module, Recommendation Strategy Training Module, and Optimal Recommendation Strategy Deployment Module until the set metrics of the recommendation strategy meet the requirements of the system.
[0033] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the recommendation strategy optimization method based on the user behavior model as described above.
[0034] A computer-readable storage medium stores a computer program for executing the recommendation strategy optimization method based on the user behavior model as described above.
[0035] Beneficial Effects: Compared with the prior art, the recommendation strategy optimization method and system based on the user behavior model provided by the present invention construct a user behavior model based on the Generative Adversarial Imitation Learning algorithm (GAIL) from the offline interaction data between the user and the recommendation system, and optimize the video recommendation strategy based on the reinforcement learning technology. The advantages are as follows:
[0036] First, during the training process of the present invention, real-time interaction with the recommendation system is not required. An offline dataset is generated by collecting historical log data of user interactions with the recommendation system, and a user behavior model is constructed based on the offline dataset. The recommendation strategy interacts with the user behavior model, and the reinforcement learning technique is used to optimize the recommendation strategy to obtain the optimal recommendation strategy, significantly improving the immediate interaction metrics (such as viewing time, click-through rate, like rate, etc.) and long-term interaction metrics of the recommendation strategy.
[0037] Second, the present invention constructs a user behavior model from offline data. The user behavior model can be used in the evaluation process of the user recommendation strategy, improving the interpretability and optimization sustainability of the recommendation strategy. At the same time, based on the reinforcement learning technique, a recommendation strategy with good generalization and strong robustness is trained on the user behavior model with a low-cost trial-and-error cost, effectively reducing the poor experience and loss risk caused by the user directly interacting with the sub-optimal recommendation strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the schematic diagram of the method of the embodiment of the present invention;
[0039] Figure 2 is the schematic diagram of the user behavior model of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.
[0041] A method for optimizing a video recommendation strategy based on a user behavior model uses the reinforcement learning technique to train a user behavior model from offline interaction data between a user and a video recommendation system, and optimizes the video recommendation strategy based on the reinforcement learning technique. First, a user-recommendation system interaction dataset is collected, then a generative adversarial imitation learning algorithm (GAIL) is used to construct a user behavior model, and an optimal recommendation strategy is trained based on the user behavior model. Then, the optimal recommendation strategy is deployed in a real recommendation system and user-recommendation system interaction data is collected. The training of the user behavior model and the optimal recommendation strategy, as well as the deployment of the optimal recommendation strategy, are repeated until the recommendation strategy meets the platform requirements.
[0042] As Figure 1 shown, first, offline data of user-recommendation system interactions is collected, then a generative adversarial imitation learning algorithm is used to train a user behavior model, and an optimal recommendation strategy is trained on the user behavior model based on the reinforcement learning technique. Then, the recommendation strategy is deployed in a real recommendation system and user-recommendation system interaction data is collected. The training and deployment are repeated until the relevant metrics of the recommendation strategy meet the system requirements.
[0043] As Figure 2 shown, the video tag information recommended by the recommendation system is first converted by the Embedding Model to obtain the Embedding representation of the corresponding video tag information, which is concatenated with the user status information and the predicted click-through rate (CTR) of the video. Then, a sequential feature extraction unit (GRU) is used to extract the corresponding list features, and finally, a multi-layer perceptron neural network (MLP) outputs the corresponding user feedback, including the predicted click-through rate given by the user and the interval time for the user to visit the recommendation system next time.
[0044] The method for optimizing the video recommendation strategy based on the user behavior model includes the following steps in the embodiment:
[0045] Step 101: Collect the historical data of the interaction between real users and the recommendation system, and organize it into a trajectory dataset Data = {τ 1 , τ 2 , …, τ N} according to the time series, where N is the number of trajectories.
[0046] Step 102: Initialize the user feedback network U θ , the recommendation strategy network R Φ and the discriminator D ψ .
[0047] Step 103: Sample a batch of data from the dataset. Each piece of data is the starting point in the trajectory, which contains the timestamp information corresponding to this data, the user's click history list , the user's click feedback at the previous moment . The user's click history list is converted into the corresponding user click history state through the Embedding Model.
[0048] Step 104: Input the timestamp information , the user's historical state , and the user's click feedback information into the recommendation strategy network R Φ to generate the corresponding candidate video information V t :
[0049]
[0050] Step 105: Input the user's historical click state and the candidate video information V t into the user feedback network U θ to obtain the click feedback information at the next moment and the interval time information for the next request :
[0051]
[0052] Step 106: Add the current timestamp information to the interval time information for the next request to obtain the timestamp information for the next request:
[0053]
[0054] Iterate steps 103 - 106 to generate a batch of trajectory data D of user interactions with the recommendation policy network g ;
[0055] Step 107: Train the discriminator D in the manner of generative adversarial imitation learning (GAIL). ψ Sample a batch of real interaction trajectory data D of users and the recommendation system from the offline dataset of real user interactions with the recommendation system r , and input the above - generated trajectory data D g into the discriminator D simultaneously ψ , and optimize according to the following objective:
[0056]
[0057] Step 108: Train the user behavior network in the manner of generative adversarial imitation learning (GAIL). In the embodiment, use the reinforcement learning algorithm PPO to update the user behavior network, and the optimization objective is the discounted cumulative reward on the trajectory generated by the user , where:
[0058]
[0059] Among them, is the discount factor of the reward, usually set as a real number between (0, 1], and the reward at time t r t = logD(τ t ) is set as the natural logarithm transformation value of the score output by the discriminator for the trajectory τ t .
[0060] Step 109: Repeat steps 101 - 108 until the discriminator loss function converges or reaches the given number of training times.
[0061] Step 110: Output the optimal user behavior network as the user behavior model M u .
[0062] Step 111: Train the recommendation policy.
[0063] First, initialize a recommendation strategy π θ , and interact with the above-trained user behavior model M u to collect empirical data of the recommendation strategy , where s t represents the state of the recommendation system at time t, a t represents the action of the recommendation system at time t, r t is the immediate reward given by the user model. Use the reinforcement learning algorithm PPO to optimize the recommendation strategy, and the optimization objective is as follows:
[0064]
[0065]
[0066] where the Clip function restricts the difference in the probability ratio between the new strategy π θ and the old strategy not to exceed , where is a hyperparameter, which is usually set to a decimal close to 0, such as 0.2. It restricts the probability ratio of the new and old strategies to a symmetric interval near 1. Adv is the advantage function of ( s t , a t ), which is used to estimate the advantage of selecting the action s t under the state a t .
[0067] Iteratively optimize until the recommendation strategy converges or reaches the given number of training times, and output the optimal recommendation strategy .
[0068] Step 112: Deploy the above optimal video recommendation strategy to the real video recommendation system, and use the online data to evaluate whether the relevant metrics of the recommendation strategy meet the requirements of the system
[0069] Step 113: If the evaluation of the relevant metrics of the recommendation strategy in Step 112 does not meet the requirements of the system, then use this recommendation strategy to interact with the user behavior model M u to sample a batch of interaction trajectory data Data2 = , and add it to the offline dataset Data in Step 101 to form an expanded offline dataset, that is: Data = Data ∪ Data2
[0070] Step 114: Repeat steps 103 - 113 until the evaluation metrics of the recommendation strategy meet the requirements of the system, at which point the entire process ends and the video recommendation strategy that meets the system requirements is output.
[0071] A video recommendation strategy optimization system based on a user behavior model, including:
[0072] Dataset generation module: Used to generate an offline user - recommendation system interaction dataset. It retrieves the interaction data of users within a period of time from the log system of the recommendation system. Each piece of data includes: current timestamp, user ID, user's watched historical video list, recommended video information, and user click feedback information. After sorting the corresponding interaction data of the same user ID according to the timestamp, the interaction trajectory data of the user and the recommendation system is obtained, and the interaction trajectory data constitutes the user - recommendation system interaction dataset.
[0073] User behavior model training module: Trains the user behavior model using the Generative Adversarial Imitation Learning algorithm (GAIL). By simultaneously modeling the explicit feedback of users such as click behavior and implicit feedback such as the interval time of the user's next request, a more accurate prediction result of user behavior feedback is obtained.
[0074] Recommendation strategy training module: Initializes a recommendation strategy, interacts with the above - trained user behavior model to collect data, and uses the Proximal Policy Optimization (PPO) reinforcement learning algorithm to optimize the relevant metrics of the recommendation strategy until convergence or a given number of training times is reached, and outputs the optimal recommendation strategy.
[0075] Optimal recommendation strategy deployment module: Deploys the optimal recommendation strategy obtained from training the user behavior model to the real - world recommendation system to replace the original recommendation strategy in the recommendation system and interact with real users.
[0076] Evaluation module: Uses online data to evaluate whether the interaction metrics of the recommendation strategy meet the requirements of the system. If not, continue to collect new user - recommendation system interaction data and merge it with the dataset of the dataset generation module, and execute the user behavior model training module, the recommendation strategy training module, and the optimal recommendation strategy deployment module until the relevant metrics of the recommendation strategy meet the system requirements.
[0077] Obviously, those skilled in the art should understand that each step of the video recommendation strategy optimization method based on the user behavior model in the above embodiments of the present invention or each module of the video recommendation strategy optimization system based on the user behavior model can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
Claims
1. A method for optimizing a recommendation strategy based on a user behavior model, characterized in that, It includes the following steps: 1) Generate an offline user-recommendation system interaction dataset; retrieve the interaction data of users within a period of time from the log system of the recommendation system, and after sorting the corresponding interaction data of the same user ID according to the time stamp, obtain the interaction trajectory data of the user and the recommendation system, and the interaction trajectory data constitutes the user-recommendation system interaction dataset; 2) Use the generative adversarial imitation learning algorithm to train the user behavior model; 3) Train the recommendation strategy; Initialize a recommendation strategy, interact with the trained user behavior model to collect data, and use the reinforcement learning algorithm PPO to optimize the metrics of the recommendation strategy until convergence or a given number of training times is reached, and output the optimal recommendation strategy; 4) Deploy and evaluate the optimal recommendation strategy; deploy the optimal recommendation strategy to the recommendation system, and use online data to evaluate whether the interaction metrics of the recommendation strategy meet the requirements of the system; 5) If the result of the online evaluation does not meet the system requirements, continue to collect new user-recommendation system interaction data, and repeat steps 1) to 4) until the relevant metrics of the recommendation strategy meet the system requirements; The steps of using the generative adversarial imitation learning algorithm to train the user behavior model are as follows: Step 1: Initialize the user behavior network, the recommendation strategy network, and the discriminator D; Step 2: Sample a batch of data from the user-recommendation system interaction dataset; each piece of data is the starting point in the trajectory, which contains the time stamp information corresponding to the data, the user click history list, and the user's click feedback information at the previous moment. The user click history list is converted into the corresponding user click history state through the Embedding Model; Step 3: Input the time stamp information, the user click history state, and the user's click feedback at the previous moment into the recommendation strategy network to generate the corresponding candidate video information; Step 4: Input the user click history state and the candidate video into the user behavior network to obtain the user's click feedback information on the candidate video and the interval time information of the next request; Step 5: Add the current timestamp with the interval time information for the next request to obtain the timestamp information for the next request, and iterate Steps 3 to 5 to generate a batch of trajectory data D of user interactions with the recommendation policy network g ; Step 6: Update the discriminator parameters; sample a batch of real user and recommendation system interaction trajectory data D from the historical offline interaction dataset between real users and the video recommendation system r , and input the generated trajectory data D g into the discriminator D simultaneously, and optimize the following objective loss function: Among them, τ represents the trajectory in the dataset, D(τ) indicates that the input data of discriminator D is τ, E represents expectation, log represents taking the natural logarithm with base e. The discriminator D maximizes the discriminator score under the real interaction trajectory data of users and the recommendation system D r and minimizes the discriminator score under the generated interaction trajectory data of users and the recommendation policy network D g to distinguish whether the trajectory comes from the interaction trajectory data D r or from the interaction trajectory data D g ; Step 7: Update the parameters of the user behavior network, iterate Steps 2 to 5 to generate a batch of trajectory datasets {τ 1 , τ 2 , …, τ N} generated by the interaction between the user behavior network and the recommendation policy network. The optimization objective is to generate the discounted cumulative reward on the interaction trajectory data : Among them, is the discount coefficient of the reward, set as a real number between (0, 1], and the reward at time t r t = logD(τ t ) is set as the natural logarithm transformation value of the score output by the discriminator for the trajectory τ t ; Step 8: Repeat steps 2-7 until the loss function of the discriminator D converges or a given number of training times is reached; Step 9: Output the final user behavior network as the user behavior model, and the training process ends.
2. The method for optimizing a recommendation strategy based on a user behavior model according to claim 1, wherein The optimal recommendation strategy is directly deployed to the video recommendation system, and online data is used to evaluate whether the interaction metrics of the recommendation strategy meet the requirements of the system. If not, continue to collect new user-recommendation system interaction data; the new interaction data is further used as offline data for the training of the user behavior model and the training of the optimal recommendation strategy, and the strategy performance is improved iteratively until the relevant metrics of the trained recommendation strategy meet the system requirements.
3. The method for optimizing a recommendation strategy based on a user behavior model according to claim 1, wherein Retrieve the interaction data of users within a period of time from the log system of the recommendation system. Each piece of data includes: the current time stamp, the user ID, the user's viewing history video list, the recommended video information, and the user's click feedback information.
4. A recommendation strategy optimization system based on a user behavior model, characterized in that, It includes: Dataset Generation Module: It is used to generate an offline user-recommendation system interaction dataset. It retrieves the interaction data of users within a period of time from the log system of the recommendation system. After sorting the corresponding interaction data of the same user ID according to the timestamp, it obtains the interaction trajectory data of the user and the recommendation system. The interaction trajectory data constitutes the user-recommendation system interaction dataset; User Behavior Model Training Module: It uses the generative adversarial imitation learning algorithm to train the user behavior model; Recommendation Strategy Training Module: It initializes a recommendation strategy, interacts with the trained user behavior model to collect data, and uses the Proximal Policy Optimization (PPO) algorithm in reinforcement learning to optimize the relevant metrics of the recommendation strategy until convergence or a given number of training times is reached, and then outputs the optimal recommendation strategy; Optimal Recommendation Strategy Deployment Module: It deploys the optimal recommendation strategy output by the recommendation strategy training module into the recommendation system to replace the original recommendation strategy in the recommendation system to interact with users; Evaluation Module: It uses online data to evaluate whether the interaction metrics of the recommendation strategy meet the requirements of the system. If not, it continues to collect new user-recommendation system interaction data and merge it with the dataset in the dataset generation module, and then executes the user behavior model training module, the recommendation strategy training module, and the optimal recommendation strategy deployment module until the set metrics of the recommendation strategy meet the requirements of the system; The steps of using the generative adversarial imitation learning algorithm to train the user behavior model are as follows: Step 1: Initialize the user behavior network, the recommendation strategy network, and the discriminator D; Step 2: Sample a batch of data from the user-recommendation system interaction dataset; each piece of data is the starting point in the trajectory, which contains the timestamp information corresponding to the data, the user's click history list, and the click feedback information of the user at the previous moment. The user's click history list is converted into the corresponding user click history state through the Embedding Model; Step 3: Input the timestamp information, the user click history state, and the click feedback of the user at the previous moment into the recommendation strategy network to generate the corresponding candidate video information; Step 4: Input the user click history state and the candidate video into the user behavior network to obtain the click feedback information of the user for the candidate video and the interval time information for the next request; Step 5: Add the current timestamp to the interval time information for the next request to obtain the timestamp information for the next request, and iterate Steps 3 to 5 to generate a batch of trajectory data D of user interactions with the recommendation policy network g ; Step 6: Update the discriminator parameters; sample a batch of real user and recommendation system interaction trajectory data D from the historical offline interaction dataset of real users and the video recommendation system r , and input the generated trajectory data D g into the discriminator D simultaneously, and optimize the following objective loss function: Among them, τ represents the trajectory in the dataset, D(τ) indicates that the input data of the discriminator D is τ, E represents the expectation, log represents taking the natural logarithm with base e, and the discriminator D maximizes the discriminator score under the real interaction trajectory data of users and the recommendation system D r while minimizing the discriminator score under the generated interaction trajectory data of users and the recommendation policy network D g to distinguish whether the trajectory comes from the interaction trajectory data D r or from the interaction trajectory data D g ; Step 7: Update the user behavior network parameters, iterate from Step 2 to Step 5 to generate a batch of trajectory datasets {τ 1 , τ 2 , …, τ N} generated by the interaction between the user behavior network and the recommendation policy network. The optimization objective is to generate the discounted cumulative reward on the interaction trajectory data : Among them, is the discount factor for the reward, set as a real number between (0, 1], and the reward at time t r t = logD(τ t ) is set as the natural logarithm transformation value of the score output by the discriminator for the trajectory τ t ; Step 8: Repeat Steps 2-7 until the loss function of the discriminator D converges or a given number of training times is reached; Step 9: Output the final user behavior network as the user behavior model, and the training process ends.
5. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the recommendation strategy optimization method based on the user behavior model as described in any one of Claims 1-3.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the recommendation strategy optimization method based on the user behavior model as described in any one of Claims 1-3.
Citation Information
Patent Citations
Application recommending method and system based on user portrait behavior analysis, storage medium and computer device
CN107423442A
Search and recommendation fusion system based on unified user behavior modeling
CN113761383A