Multi-task Reinforcement Learning User Operation Method and System Based on User Model Learning
Through the multi-task reinforcement learning method based on user model learning, the general operation strategy model is trained, and the problems of low operation efficiency and high cost of traditional methods in multi-city market scenarios are solved, and automated and efficient multi-city user operations are achieved.
Patent Information
- Application Number
- CN202210537142.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-18
AI Technical Summary
Traditional user operation methods are difficult to quickly and efficiently complete user operation operations in multi-city market scenarios, and the cost is high and the process is complex, making it difficult to form a general and digital operation process.
A multi-task reinforcement learning method based on user model learning is adopted, task correlation is discovered through clustering, a general operation strategy model is designed, and a general intelligent model is trained using imitation learning and reinforcement learning algorithms to realize automated decision-making and multi-city operations.
It realizes automated and efficient multi-city user operations, reduces labor costs, simplifies operational processes, and utilizes data relevance through multi-task learning, ensuring the universality and robustness of the strategy.
Smart Images

Figure CN114912357B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-task reinforcement learning user operation method and system based on user model learning to implement a general operation system that can meet the user operation requirements of multiple cities, belonging to the field of user operation of mobile platforms. Background Art
[0002] With the continuous development of the mobile Internet in China, all walks of life have started to develop towards the online platform direction. For example, traditional public transportation facilities are difficult to meet the travel needs of some users. Therefore, mobile travel platforms such as Didi have emerged, aiming to create a faster, more convenient and comfortable travel mode. For different platforms in the same industry, in order to attract new users and ensure user stickiness, the competition between different platforms is very fierce, and user operation has become one of the most core tasks of these platforms. In the real scenario, each platform will operate many cities at the same time. Due to the differences in user habits in each city, the optimal operation strategies often vary greatly. How to quickly and efficiently complete the operation of users in multiple cities has become a difficult problem for the platform.
[0003] Traditional methods rely on the manual operation team to summarize experience, and these experiences are obtained by operation personnel through data analysis of the historical data of each city. Over-reliance on the manual operation team will consume a large amount of manpower and generate high costs, and it is difficult to form a general and digital operation process. Some more advanced platforms have also introduced deep learning and reinforcement learning technologies to train neural network models to assist manual operation. However, these methods either still rely on manual work in some processes or still only consider a single city scenario. When there are multiple cities, a large amount of repetitive work will occur in the process. For example, multiple policy models need to be repeatedly launched in the model deployment stage. Moreover, the data of different cities are completely separated, and the partial similarity between city data is not utilized. Once the data volume of a new city is small or the quality is very poor, it will be very difficult to initialize a better-performing operation strategy based on this not-so-good data.
[0004] In recent years, deep reinforcement learning has begun to be widely applied to complex sequential decision-making problems in the real world, such as robot control, playing video games, and recommendation systems. By using a deep neural network policy model trained with reinforcement learning algorithms, it can replace manual decision-making. Traditional reinforcement learning methods require a large number of interactive trial-and-error processes with the environment, which will bring great harm and costs in the real environment. Therefore, in this invention, a method based on user model learning is proposed to restore the user model environment through offline user behavior data and use the user model environment to approximately replace the real environment. In addition, current reinforcement learning methods are difficult to apply to multi-task scenarios, and the trained policies have poor scalability and often can only meet the needs of a specific environment decision. Once the environment changes slightly, the policy will fail. Summary of the Invention
[0005] Object of the Invention: In the user operation tasks of the mobile platform, it is necessary to operate users in multiple cities simultaneously, and the optimal user operation strategies in these cities often vary greatly. Traditional methods generally rely on a large amount of manual data analysis work or use machine learning methods to train a user operation policy model for each city separately. The former has high costs, a simple process, and is difficult to digitize, while the latter will produce a large number of repetitive processes and is difficult to utilize the correlation of data. To solve the problems of the previous methods, this invention proposes a multi-task reinforcement learning method based on user model learning and a general operation system implemented based on this method. The entire system can automatically replace manual decision-making while also designing the user operation policy model to be a general structure that can learn multiple tasks, so as to achieve the user operation task under multi-city conditions with only one trained operation policy model.
[0006] This invention discovers the relevance of tasks based on clustering methods, transfers this prior relevance knowledge to the design of the model structure, and then uses a feasible reinforcement learning algorithm to train the general intelligent agent model of the algorithm with the learned user models of multiple cities as the multi-task environment, and finally obtains a user operation policy model that can adapt to multi-city conditions, so as to build an automatic and efficient general operation system that meets multi-cities.
[0007] Technical Solution: A multi-task reinforcement learning user operation method based on user model learning includes:
[0008] Collect the platform operation and user feedback data of each city in the online environment of the operation platform in the recent period;
[0009] Perform feature engineering to convert the original platform operation and user feedback data into user trajectory data and user behavior data suitable for machine learning respectively;
[0010] Use the user trajectory data of each city to train an encoder network for feature extraction, and output the feature vectors of each user in each city;
[0011] Based on the feature vectors of each user in each city, perform clustering operations, and then construct a general network model structure according to the clustering results;
[0012] Use the imitation learning method to restore the user model of each city from the user behavior data of each city;
[0013] Select a feasible reinforcement learning algorithm, use the general network model structure to initialize the agent model required by the reinforcement learning algorithm, and then train the general agent model of the reinforcement learning algorithm with the user models of multiple cities as a multi-task environment;
[0014] Deploy the general operation policy model in the trained general agent model to the online environment of multiple cities for user operation decision-making, and generate a new round of platform operation and user feedback data.
[0015] Specifically, the present invention includes the following steps:
[0016] (1) Collect the platform operations and self-feedback records of all users in each city on the online platform for a recent period of time;
[0017] (2) Perform feature engineering to convert the historical data of the platform operations and self-feedback records of each user into trajectory data at daily intervals that can be used for reinforcement learning;
[0018] (3) Use this trajectory data to train an encoder network for extracting user features, and the encoder network outputs the respective feature vectors of each user in each city;
[0019] (4) Use the feature vectors of each user in each city to perform clustering operations, and construct a general network model structure according to the results of the clustering operations;
[0020] (5) Use the imitation learning method to imitate the user behavior in the real user behavior data to construct the user model of each city, and these user models serve as the multi-task environment for subsequent reinforcement learning;
[0021] (6) Use the general network model structure to initialize the general agent model required by a feasible reinforcement learning algorithm, and train the general agent model simultaneously with the user models of multiple cities as a multi-task environment, and output the general operation policy model in the agent model;
[0022] (7) Deploy the trained general operation policy model to the actual environment of each city to guide user operation decision-making and generate a new round of platform operation and user feedback data.
[0023] In (1) above, the platform operation and user feedback records of all users in each city in the recent period include: the values of the platform's operation on each user within a specified time range every day, including the number of operation times and the intensity of each operation-related action. The intensity is used to measure the intensity of the platform's operation on the user. For example, for user promotion operation, it corresponds to the size of the discount; the user feedback record refers to the number of times the user gives feedback on the platform after receiving the platform's operation and the platform revenue generated by each feedback.
[0024] In (2) above, feature engineering converts the original platform operation and user feedback data into user trajectory data and user behavior data suitable for machine learning respectively. Let the collected data range from the 1st day to the 2nth day. First, obtain the initialized user portrait: taking the (n + 1)-th day as the benchmark, the portrait of the user on that day is some statistical feature data calculated based on the platform operation and user feedback records obtained from the 1st day to the nth day of the user's past history, and use s 1 to represent the user's initial portrait (corresponding to the (n + 1)-th day). Similarly, when the platform operation actions, user feedback actions, and platform return values are predefined, the platform operation actions, user feedback actions, and platform return value data for each day from the (n + 1)-th day to the 2nth day can be calculated, and are represented by a t , u t , and r t respectively (n + 1 ≤ t ≤ 2n). At the same time, according to the known transition rule: s t+1 = T(s t , a t , u t ), when we know the user portrait, platform operation action, and user feedback action on the current day, the user portrait of the next day can be calculated. In this way, starting from the user's initial portrait, based on the transition rule and the platform operation actions, user feedback actions, and platform return value data for each day from the (n + 1)-th day to the 2nth day, a trajectory data of any user from the (n + 1)-th day to the 2nth day is obtained (in the trajectory, the subscript 1 corresponds to the (n + 1)-th day):
[0025] τ ={(s 1 ,a 1 ,r 1 ,s 2 ), (s 2 ,a 2 ,r 2 ,s 3 ), …,(s n ,a n ,r n ,s n+1 )}
[0026] The trajectory data of all users in a city constitutes the trajectory dataset D of this city. If {1, …, L} represents L different cities, then the total user trajectory training data is D sum ={D 1 , …, D L}. At the same time, in order to learn the user model, it is also necessary to define the behavior data of any user from the (n + 1)-th day to the 2n-th day:
[0027] β={((s 1 ,a 1 ),u 1 ), ((s 2 ,a 2 ),u 2 ), …, ((s n ,a n ),u n )}
[0028] Similarly, the behavior data of all users in a city constitutes the user behavior dataset B of this city. The total user behavior training data is B sum ={B 1 , …, B L}.
[0029] In (3) above, the process of training the encoder network for feature extraction and outputting the feature vector includes:
[0030] (301) Select the neural network model structure for processing time-series data to initialize the encoder network . The encoder network inputs a time-series trajectory data τ of a certain user and outputs the feature vector υ corresponding to this user.
[0031] (302) Train the encoder network based on the contrastive loss. Let for any two users i, j and their trajectory data τ i , τ j . Use y ∈ {1, …, L} to represent which city the user belongs to. Then the contrastive loss generated by this pair of users i, j is:
[0032]
[0033] where m is a constant parameter, 1{ y i = y j} is an expression of the bool function, y i ,y j correspond to the cities to which users i, j belong respectively, v iand v j They respectively correspond to the feature vectors of users i and j, and the ‖·‖ expression calculates the distance of the vectors.
[0034] (303)The total contrast loss is the sum of the contrast losses of all user pairs between any two batches of users (the cities can be the same). Each batch of users is taken from different cities. It is denoted by and we hope the smaller the better. Based on gradient descent, the encoder network parameter σ is updated as follows:
[0035]
[0036] λ 1 is the learning rate, a hyperparameter.
[0037] (304)Train the initialized encoder network until convergence. For any user in the training dataset, use the converged encoder network, input the corresponding user trajectory data, and output its feature vector.
[0038] In the above (4), the process from clustering to constructing the general network structure includes:
[0039] (401)Use the feature vectors of all users in all cities as the training dataset V sum for clustering. Select any feasible clustering method to divide these users into a hierarchical clustering structure. From top to bottom, initialize the clustering process. At the beginning, all users in all cities belong to the same cluster, which is the first layer (the initial current layer 1).
[0040] (402)L is the number of cities. Assume 2 n-1 ≤L≤2 n . Repeat the following process n times: Assume the current layer is i, where 1 ≤ i ≤ n. Traverse each cluster in the current layer in turn. Use the clustering method to divide each cluster in the current layer into two smaller sub-clusters. All the new sub-clusters are used as one of the clusters in the (i + 1)-th layer, and at the same time, update the (i + 1)-th layer as the current layer. Finally, a hierarchical clustering structure in the shape of a binary tree with n + 1 layers is obtained.
[0041] (403)Equivalently map the finally obtained hierarchical clustering structure in the shape of a binary tree to the general network model structure to construct the general network model. Each node of the binary tree corresponds to a module of the neural network, and the edges of the binary tree correspond to the forward propagation connection relationships of the neural network modules.
[0042] In the above (5), when using the method of imitation learning to imitate the user behavior in the real user behavior data, it means:
[0043] For B sumFor the user behavior data of each city, use the method of imitation learning to learn a user model that maps from (user profile, platform operation actions) to user feedback actions. Each city has a user model. Finally, obtain M sum ={M 1 , …, M L}, representing the user models of L different cities.
[0044] In the above (6), select any feasible reinforcement learning algorithm to train the general intelligent agent model of the reinforcement learning algorithm, including:
[0045] In the algorithm initialization process of (601), construct all neural network models related to the intelligent agent with a general network model structure. And initialize the online sampling pool O of each city sum ={O 1 , …, O L}, initialize any sampling pool in the set O sum to an empty set. The subsequent data of the online sampling pool will be sampled from the user model M of the corresponding city sum ={M 1 , …, M L}. O L represents the sampling pool of the L-th city.
[0046] In the algorithm training process of (602), use the general intelligent agent to sample in each user model respectively, and add the sampled data to the corresponding online sampling pool. At each training step, alternately traverse each city, sample a small batch of data from the online sampling pool of the current city, and use this batch of data to optimize the loss function related to the algorithm. The algorithm is trained until the model converges to obtain the trained general operation policy model. The online environment of the operation platform is the real platform environment, and the general intelligent agent is used to interact and sample in the virtual user environment model.
[0047] In the above (7), deploying the trained general operation policy model to the actual environment of each city means:
[0048] Take out the general operation policy model after the algorithm converges, and use the general operation policy model to guide user operation in the online environment of each city: for any user, input their latest user profile and output the operation for the user.
[0049] The multi-task reinforcement learning user operation system based on user model learning includes:
[0050] The data collection module is used to collect the platform operations and self-feedback records of all users in each city in the online environment of the operation platform in the recent period;
[0051] The feature engineering module converts the historical data of each user's platform operations and self-feedback records into trajectory data at daily intervals, which can be used for reinforcement learning.
[0052] The encoder network training module uses the trajectory data to train an encoder network for extracting user features. The encoder network outputs the feature vectors of each user in each city.
[0053] The clustering module uses the feature vectors of each user in each city to perform clustering operations and constructs a general network model structure based on the results of the clustering operations.
[0054] The user model construction module uses the method of imitation learning to imitate the user behaviors in the real user behavior data to construct the user models for each city. These user models serve as the multi-task environment for subsequent reinforcement learning.
[0055] The general operation policy model training module uses the general network model structure to initialize the general agent model required for a feasible reinforcement learning algorithm, and simultaneously trains the general agent model with the user models of multiple cities as the multi-task environment, and outputs the general operation policy model in the agent model.
[0056] The model deployment module deploys the trained general operation policy model to the actual environments of each city to guide user operation decisions and generate a new round of platform operation and user feedback data.
[0057] The implementation methods of the modules in the system are the same as the corresponding steps in the multi-task reinforcement learning user operation method based on user model learning, and will not be elaborated here.
[0058] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the multi-task reinforcement learning user operation method based on user model learning as described above.
[0059] A computer-readable storage medium stores a computer program for executing the multi-task reinforcement learning user operation method based on user model learning as described above.
[0060] Beneficial effects: Compared with the prior art, the multi-task reinforcement learning user operation method and system based on user model learning provided by the present invention have the following advantages:
[0061] 1) Based on a data-driven and end-to-end deep learning framework, except for the definition of data features in the early stage, almost no human intervention is required in the whole process, which saves costs and is more efficient and intelligent.
[0062] 2) The reinforcement learning method based on user model learning avoids frequently deploying poor operation strategies to collect reinforcement learning data in the real environment. Instead, it can approximately replace this process by collecting data from the user model, ensuring low cost and feasibility in the practical sense.
[0063] 3) Based on the idea of multi-task learning, the present invention can effectively utilize the correlation between multi-city data to mine general knowledge. Even if the data quality of a certain city is mediocre, under the constraint of data from other cities, it can ensure a basic performance guarantee for the final strategy in all cities, that is, to ensure the generality and robustness of the learned strategy. At the same time, it should be noted that since only one set of models needs to be trained, compared with the method of training multiple sets of models for each city, it greatly reduces the computational resource overhead and simplifies the deployment process. Only one general policy model needs to be deployed. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a flowchart of the method in the embodiment of the present invention;
[0065] Figure 2 is a schematic diagram of the hierarchical clustering process in the embodiment of the present invention;
[0066] Figure 3 is a schematic diagram of mapping the hierarchical clustering result to the general network structure in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.
[0068] As Figure 1 shown, for the multi-task reinforcement learning user operation method based on user model learning, taking the user coupon distribution operation of a mobile travel platform as an example, the multi-city user coupon distribution operation method corresponding to the multi-task reinforcement learning user operation method cyclically performs six steps:
[0069] Step 1:
[0070] Collect data online on the travel platform. Select a user set for each city and collect an offline data set of this batch of users. The offline data should include the taxi records and the records of obtaining promotional coupons of the selected user set in the past two months (calculated as 60 days) in this city.
[0071] Step 2:
[0072] The general process of the system to perform feature engineering on each piece of offline data to obtain user trajectory data and user behavior data has been described in detail in the technical solution. Next, an example of data feature definition is given. Table 1 gives a simple definition of user portraits. Then consider defining the coupon issuance action, user action, and return value:
[0073] Coupon issuance action: The number of coupons issued on the same day and the average discount (or average denomination) of the coupons issued
[0074] User action: The number of taxi orders on the same day and the average amount of the orders
[0075] Return value: The total amount of taxi rides on the same day minus the amount deducted by the coupons
[0076] Define the initial user portrait based on the historical data from the 1st day to the 30th day. Obtain the coupon issuance action, user action, and return value for each day according to the data for each day from the 31st day to the 60th day.
[0077] Table 1. Simple definition of user portraits
[0078] Feature Name Description total_num Total number of the user's historical orders (with a discount factor of 0.99) average_num Average of the user's historical daily order numbers (excluding days with zero orders) average_fee Average of the user's historical daily average order prices (excluding days with zero prices)
[0079] The transition rules are naturally generated according to the definitions of user portraits, coupon issuance actions, and user actions. Given the user portrait, coupon issuance action, and user action on the same day, the user portrait for the next day can be calculated according to the transition rules. If we use state to represent the user portrait on the same day and the user action on the same day is act. Since act[0] represents the number of orders of the user on the same day and act[1] represents the average amount of the user's orders on the same day, in order to obtain the user portrait next_state for the next day, according to the definition of the user portrait, we can directly calculate each dimension of next_state using act and state. It should be noted that here, since the definition of the user portrait is relatively simple, the influence of the coupon issuance action on the same day can be ignored when calculating the user portrait for the next day.
[0080] Step 3:
[0081] Select the Transformer Net to initialize the feature encoder network. A complete Transformer Net usually consists of n encoder layers and m decoder layers. Since only feature variables need to be extracted from the trajectory data at this step, the Transformer Net here actually only requires n encoder layers, and generally n can be set to 6. Each encoder consists of two components: the self-attention mechanism and the feed-forward neural network. The self-attention mechanism accepts the input encodings from the previous encoder and weighs the correlations between them to generate output encodings. The feed-forward neural network further processes each output encoding separately. The encoder network is trained based on the contrastive loss. Let for any two users i, j and their trajectory data τ i , τ j . Represent which city the user belongs to with y ∈ {1, …, L}, then the contrastive loss generated by this pair of users i, j is:
[0082]
[0083] where m is a constant parameter, and the total contrastive loss is the sum of the contrastive losses of all user pairs between any two batches of users (the cities can be the same). We train the feature encoder network based on minimizing the total contrastive loss. The trained feature encoder network takes the trajectory of each user as input and outputs their feature vectors.
[0084] Step Four:
[0085] Use the feature vectors of all users in all cities as the training dataset V for clustering sum , select the k-means clustering algorithm, and partition these users into a hierarchical clustering structure. From top to bottom, initialize the clustering process. At the beginning, all users in all cities belong to the same cluster, which is the first layer (the initial current layer 1). L is the number of cities, assuming 2 n-1 ≤ L ≤ 2 n , loop through the following process n times: Assume the current layer is i, 1 ≤ i ≤ n, traverse each cluster in the current layer in turn, and use the k-means clustering algorithm to partition each cluster in the current layer into two smaller sub-clusters. All the new sub-clusters are one of the clusters in the i + 1 layer, and at the same time update the i + 1 layer as the current layer. Finally, a hierarchical clustering structure in the shape of a binary tree with n + 1 layers is obtained.
[0086] For the k-means clustering algorithm, given the set of user feature vectors (v 1 , v 2 , …, v n) The k-means algorithm divides these n feature vectors into k sets, minimizing the within-group sum of squares, that is, finding the clustering S that satisfies the following formula i :
[0087]
[0088] where u i is the mean of all points in S i . In this embodiment, as Figure 2 shown, each time the k-means algorithm is called, two sub-clusters are divided from the existing user clustering set of cities. Taking the division of two sub-clusters {A, E, F} and {B, C, D} from the initial cluster {A, B, C, D, E, F} as an example:
[0089] Suppose there are a total of 1000 user feature vectors {v 1 , v 2 , …, v 1000} for all cities in the initial clustering set. Randomly select two objects as the centers of the two sub-clusters. During the assignment process, the 1000 user feature vectors are assigned to the center point closest to them. In this way, two current clusters will be obtained. Take the center points of the two current clusters as the new center points and repeat the assignment process until the center points no longer change or reach the maximum number of iterations of the algorithm. It should be noted that for any city, the cluster with the largest number of users of that city in the clustering result is used as the cluster where the city is located.
[0090] Figure 3 shows how to map to the structure of a neural network through the binary tree structure of hierarchical clustering. It can be seen that this is an equivalent mapping. Each node on the tree corresponds to a module of the neural network. Each module consists of one (or multiple) hidden layers, and each hidden layer consists of multiple neurons. The connection relationship between modules is the same as that between tree nodes, and the forward propagation direction of the network is the same as the direction from the root node of the tree to the bottom.
[0091] Step Five:
[0092] The general process of restoring the user model using imitation learning. In this embodiment, the most commonly used behavior cloning algorithm in imitation learning can be selected to learn the user model of each city. The behavior cloning algorithm uses the maximum likelihood method to learn a user model that maps from (user portrait, coupon issuance action) to user actions. Each city has a corresponding user model.
[0093] Step Six:
[0094] When there is a similar Figure 3The general network structure shown, and user models for multiple cities, select a reinforcement learning algorithm and train it in a multi-task manner. TD3 is a classic reinforcement learning algorithm. In this invention example, the general process of its multi-task training is outlined:
[0095] Input: User models for multiple cities {M 1 , M 2 , …, M L}, initialize the online sampling pools for multiple cities that are initially empty {O 1 , O 2 , …, O L}, and initialize the Q-value network using the general network structure , the policy network , and the target networks corresponding to these networks .
[0096] 1) Copy the parameters of the policy network model and the Q-value network model to the target network: ;
[0097] 2) Use to sample data from the user environment models {M 1 , M 2 , …, M L} for each city respectively, and add the sampled data to the corresponding online sampling pools {O 1 , O 2 , …, O L} respectively;
[0098] 3) Sample a small batch of data from the online sampling pools {O 1 , O 2 , …, O L} for each city respectively , and each β i has N pieces of data;
[0099] 4) Update based on the following objective formula:
[0100]
[0101] where
[0102]
[0103] γ is the discount factor, c is a positive constant, and ε is noise sampled from a normal distribution;
[0104] policy_delay represents a positive integer. If the loop count of this time satisfies j % policy_delay = 0:
[0105] Update based on the following objectives :
[0106]
[0107] , where ρ is a non - negative constant less than 1;
[0108] 5) Loop back to 2) until the policy network model converges and ends.
[0109] Output: General coupon - issuing operation policy network .
[0110] Step Seven:
[0111] When there is a trained general coupon - issuing operation policy network , deploy it to the online coupon - issuing operation system. For any user in any city in the training set, input their latest user portrait s, and the general coupon - issuing operation policy network outputs a coupon - issuing action a for him, and perform user - specific coupon - issuing operations based on this coupon - issuing action.
[0112] The multi - city user coupon - issuing operation system based on user model learning includes:
[0113] A data collection module, used to collect the taxi - taking records and coupon - obtaining records of all users in each city in a recent period of time from the online travel platform;
[0114] A feature engineering module, which converts the historical data of each user's taxi - taking records and coupon - obtaining records into trajectory data at daily intervals that can be used for reinforcement learning;
[0115] An encoder network training module, which uses the trajectory data to train an encoder network for extracting user features, and the encoder network outputs the feature vectors of each user in each city;
[0116] A clustering module, which uses the feature vectors of each user in each city to perform clustering operations, and constructs a general network model structure according to the results of the clustering operations;
[0117] A user model construction module, which uses the method of imitation learning to imitate the user behaviors in the real user behavior data to construct the user models of each city, and these user models are used as the multi - task environment for subsequent reinforcement learning;
[0118] The general operation strategy model training module uses a general network model structure to initialize the general agent model required for a feasible reinforcement learning algorithm, and trains the general agent model with the user models of multiple cities as a multi-task environment, and outputs the general coupon-issuing operation strategy model in the agent model;
[0119] The model deployment module deploys the trained general coupon-issuing operation strategy model to the actual environment of each city to guide user-specific coupon-issuing operation decisions and generate a new round of platform operation and user feedback data.
[0120] Obviously, those skilled in the art should understand that each step of the multi-task reinforcement learning user operation method based on user model learning in the above embodiments of the present invention or each module of the multi-task reinforcement learning user operation system based on user model learning can be implemented by a general computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
Claims
1. A multi-task reinforcement learning user operation method based on user model learning, characterized in that, it includes the following steps: Step (1), collect the platform operations and self-feedback records of all users in each city within a specified time range on the online platform; Step (2), perform feature engineering to convert the historical data of the platform operations and self-feedback records of each user into trajectory data for reinforcement learning; Step (3), use the trajectory data to train an encoder network for extracting user features, and the encoder network outputs the feature vectors of each user in each city; Step (4), use the feature vectors of each user in each city to perform clustering operations, and construct a general network model structure according to the results of the clustering operations; Step (5), use the method of imitation learning to imitate the user behaviors in the real user behavior data to construct the user model of each city; Step (6), use the general network model structure to initialize the general agent model required by the reinforcement learning algorithm, and train the general agent model simultaneously with the user models of multiple cities as a multi-task environment, and output the general operation strategy model in the agent model; Step (7), deploy the trained general operation strategy model to the actual environment of each city to guide the user operation decision-making and generate a new round of platform operation and user feedback data; In step (2), feature engineering converts the historical data of each user's platform operations and their own feedback records into trajectory data for reinforcement learning. Let the collected data range from the 1st day to the 2nth day. First, obtain the initialized user profile: Taking the (n + 1)th day as the benchmark, the user profile on that day is the statistical feature data calculated based on the platform operations and user feedback records obtained from the user's past history from the 1st day to the nth day, denoted by s 1 to represent the user's initial profile. When the platform operation actions, user feedback actions, and platform reward values are predefined, the platform operation action, user feedback action, and platform reward value data for each day from the (n + 1)th day to the 2nth day can be calculated, denoted by a t , u t , and r t respectively, where n + 1 ≤ t ≤ 2n. At the same time, according to the known transition rule: s t+1 = T(s t , a t , u t ), when the user profile, platform operation action, and user feedback action of the current day are known, the user profile of the next day can be calculated; Starting from the initial portrait of the user, based on the transfer rules and the platform operation actions, user feedback actions and platform reward value data every day from the (n + 1)-th day to the 2n-th day, a trajectory data of any user from the (n + 1)-th day to the 2n-th day is obtained: τ = {(s 1 , a 1 , r 1 , s 2 ), (s 2 , a 2 , r 2 , s 3 ), …, (s n , a n , r n , s n+1 )} The trajectory data of all users in a city constitutes the trajectory dataset D of this city; if {1, …, L} represents L different cities, then the total user trajectory training data is D sum ={D 1 , …, D L}; The behavior data of any user from the (n + 1)-th day to the 2n-th day is as follows: β = {((s 1 , a 1 ), u 1 ), ((s 2 , a 2 ), u 2 ), …, ((s n , a n ), u n )} Similarly, all user behavior data of a city constitutes the user behavior data set B of this city; the total user behavior training data is B sum ={B 1 ,…,B L}; In the step (3), the process of training the encoder network for feature extraction and outputting the feature vectors includes: (301) Select the neural network model structure for processing time series data to initialize the encoder network ω σ ; The encoder network inputs a time series trajectory data τ of a certain user and outputs a feature vector υ corresponding to this user; (302) Train the encoder network based on the contrast loss; (303) The total contrastive loss is the sum of the contrastive losses of all user pairs between any two batches of users selected from different cities, denoted by , and the encoder network parameters σ are updated as follows based on gradient descent: (304) Train the initialized encoder network until convergence. For any user in the training dataset, use the converged encoder network, input the corresponding user trajectory data, and output its feature vector.
2. The multi-task reinforcement learning user operation method based on user model learning according to claim 1, characterized in that, in the step (1), the platform operations and user feedback records of all users in each city within a specified time range include: the values of the platform's operation operations on him every day within the specified time range, including the number of operation operations and the intensity of each operation-related action; the user feedback record refers to the number of times the user gives feedback on the platform after receiving the platform's operation operations and the platform revenue generated by each feedback.
3. The multi-task reinforcement learning user operation method based on user model learning according to claim 1, characterized in that, in the step (4), from clustering to constructing the general network structure includes: (401) Use the feature vectors of all users in all cities as the training data set V for clustering sum , and divide users into a hierarchical clustering structure; from top to bottom, initialize the clustering process. At the beginning, all users in all cities belong to the same cluster, which is the first layer; (402) Let 2 n-1 ≤L≤2 n , perform the following process n times in a loop: Let the current layer be i, where 1 ≤ i ≤ n. Traverse each cluster in the current layer in sequence. Using the clustering method, divide each cluster in the current layer into two smaller sub-clusters. All the new sub-clusters are used as one of the clusters in the (i + 1)-th layer, and at the same time, update the (i + 1)-th layer as the current layer; finally, obtain a hierarchical clustering structure in the shape of a binary tree with n + 1 layers; (403) Equivalently map the finally obtained binary tree-shaped hierarchical clustering structure to the general network model structure to construct the general network model; each node of the binary tree corresponds to a module of the neural network, and the edges of the binary tree correspond to the forward propagation connection relationships of the neural network modules.
4. The multi-task reinforcement learning user operation method based on user model learning according to claim 1, characterized in that, in the step (5), the method of imitation learning is used, and the user behavior in the real user behavior data to be imitated refers to: For the user behavior training data B for each city sum in the total user behavior training data B, using the method of imitation learning, a user model that maps from (user profile, platform operation actions) to user feedback actions is learned, and there is a user model for each city; obtaining M sum ={M 1 ,…,M L}, representing the user models of L different cities.
5. The multi-task reinforcement learning user operation method based on user model learning according to claim 1, characterized in that, in the step (6), a reinforcement learning algorithm is selected to train the general operation policy model of the algorithm, including: During the initialization process of the (601) algorithm, neural network models related to all agents are constructed using a general network model structure; and the online sampling pool O of each city is initialized. sum ={O 1 ,…,O L}, and any sampling pool in the set O sum is initialized as an empty set; subsequent data in the online sampling pool will be sampled from the user model M sum ={M 1 ,…,M L}; (602) During the algorithm training process, a general agent is used to sample in each user environment model respectively, and the sampled data is added to the corresponding online sampling pool; at each training step, each city is traversed alternately, and a part of the data is sampled from the online sampling pool of the current city, and this part of the data is used to optimize the loss function related to the algorithm; the algorithm is trained until the model converges to obtain the trained general operation policy model.
6. A multi-task reinforcement learning user operation system based on user model learning, characterized in that, it includes: a data collection module, which is used to collect the platform operations and self-feedback records of all users in each city within a specified time range in the online environment of the operation platform; a feature engineering module, which converts the historical data of the platform operations and self-feedback records of each user into trajectory data that can be used for reinforcement learning at daily intervals; an encoder network training module, which uses the trajectory data to train an encoder network for extracting user features, and the encoder network outputs the feature vectors of each user in each city; a clustering module, which uses the feature vectors of each user in each city to perform clustering operations, and constructs a general network model structure according to the results of the clustering operations; a user model construction module, which uses the method of imitation learning to imitate the user behavior in the real user behavior data to construct the user model of each city, and these user models are used as the multi-task environment for subsequent reinforcement learning; a general operation policy model training module, which uses the general network model structure to initialize the general agent model required by a feasible reinforcement learning algorithm, and simultaneously trains the general agent model with the user models of multiple cities as the multi-task environment, and outputs the general operation policy model in the agent model; a model deployment module, which deploys the trained general operation policy model to the actual environment of each city to guide user operation decisions and generate a new round of platform operation and user feedback data; In the feature engineering module, feature engineering converts the historical data of each user's platform operations and their own feedback records into trajectory data for reinforcement learning. Let the collected data range from the 1st day to the 2nth day. First, obtain the initial user profile: Based on the (n + 1)th day, the user profile on that day is the statistical feature data calculated from the user's past history from the 1st day to the nth day, based on the obtained platform operations and user feedback records, and is represented by s 1 to represent the user's initial profile; When the platform operation actions, user feedback actions, and platform reward values are predefined, the platform operation actions, user feedback actions, and platform reward value data for each day from the (n + 1)th day to the 2nth day can be calculated, and are represented by a t , u t and r t respectively, where n + 1 ≤ t ≤ 2n. At the same time, according to the known transition rule: s t+1 = T(s t , a t , u t ), when the user profile, platform operation action, and user feedback action of the current day are known, the user profile of the next day can be calculated; Starting from the initial portrait of the user, based on the transfer rules and the platform operation actions, user feedback actions and platform return value data from the (n + 1)-th day to the 2n-th day every day, a trajectory data of any user from the (n + 1)-th day to the 2n-th day is obtained: τ = {(s 1 , a 1 , r 1 , s 2 ), (s 2 , a 2 , r 2 , s 3 ), …, (s n , a n , r n , s n+1 )} The trajectory data of all users in a city constitutes the trajectory dataset D of this city; if {1, …, L} represents L different cities, then the total user trajectory training data is D sum ={D 1 , …, D L}; The behavior data of any user from the (n + 1)-th day to the 2n-th day is as follows: β = {((s 1 , a 1 ), u 1 ), ((s 2 , a 2 ), u 2 ), …, ((s n , a n ), u n )} Similarly, the user behavior data of all users in a city constitutes the user behavior data set B of this city; the total user behavior training data is B sum ={B 1 ,…,B L}; In the encoder network training module, the process of training the encoder network for extracting features and outputting feature vectors includes: (301) Select the neural network model structure for processing time series data to initialize the encoder network ω σ ; The encoder network inputs a time series trajectory data τ of a certain user and outputs a feature vector υ corresponding to this user; (302) Training the encoder network based on the contrast loss; (303) The total contrastive loss is the sum of the contrastive losses for all user pairs between any two batches of users from different cities. It is denoted by and the encoder network parameters σ are updated as follows based on gradient descent: (304) Training the initialized encoder network until convergence. For any user in the training dataset, using the converged encoder network, inputting the corresponding user trajectory data, and outputting its feature vector.
7. A computer device, the computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the multi-task reinforcement learning user operation method based on user model learning as described in any one of claims 1-5.
8. A computer-readable storage medium, the computer-readable storage medium stores a computer program for executing the multi-task reinforcement learning user operation method based on user model learning as described in any one of claims 1-5.
Citation Information
Patent Citations
A method and device for predicting user behaviors through deep reinforcement learning
CN109559216A
End-to-end electric energy transaction market user decision-making method based on reinforcement learning
CN112529610A