Content Recommendation Method, Apparatus, Electronic Device, and Storage Medium
By obtaining and adjusting the recommended actions of account status and generating a network with action recommendation information, the problem of difficult content recommendation strategies in the existing technology is solved, efficient content recommendation strategy optimization is achieved, and users can obtain content efficiency.
Patent Information
- Application Number
- CN202210608734.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-31
AI Technical Summary
When optimizing content recommendation strategies, it is difficult for the prior art to pay attention to long-term indicators efficiently, resulting in low user efficiency in obtaining content.
Reconstruct the network by obtaining the account status of the current account and inputting it into the recommended action to reconstruct the network. Determine the action adjustment information based on the account status characteristics, adjust the recommended action, and obtain the adjusted recommended action. Use action recommendation information to generate network prediction interaction degree, determine target recommendation actions, and apply them to target applications.
It realizes the optimization of content recommendation strategies within the scope of online deployment strategies, reduces the difficulty and cost of optimization, and improves the quantity and efficiency of users obtaining content through target applications.
Smart Images

Figure CN114925278B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a content recommendation method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of Internet technologies, users can access media content such as pictures, videos, and music provided by the Internet through mobile phones, laptop computers, etc. For example, users can watch short video content in real time through a short video application installed on a mobile terminal. In the field of content recommendation, in order to increase the amount of content obtained by users through a target application, it is very important to optimize long-term metrics such as the degree of interaction between users and the target application (e.g., user retention).
[0003] However, when optimizing the long-term metrics of content recommendation strategies, related technologies often need to search and optimize in the entire huge recommendation strategy space, and the feedback signals related to user retention are very sparse, and the optimization difficulty cost is very high. This makes the content recommendation strategies in related technologies unable to efficiently focus on the optimization of long-term metrics, so that the content recommendation strategies in related technologies often cannot accurately recommend the content required by users to users, which is not conducive to increasing the amount of content obtained by users through the target application and affects the efficiency of users obtaining content from the target application. Summary of the Invention
[0004] The present disclosure provides a content recommendation method, apparatus, electronic device, and storage medium to at least solve the problem that the content acquisition efficiency of users in related technologies is not high. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a content recommendation method, the method including:
[0006] Obtaining the account status of the current account, and inputting the account status of the current account into a recommendation action reconstruction network to obtain at least one reconstructed recommendation action: the recommendation action reconstruction network is trained by using the account status of a sample account and the corresponding sample recommendation action; the sample recommendation action is a content recommendation action generated by a preset recommendation strategy in response to the account status of the sample account:
[0007] Determining action adjustment information corresponding to each of the reconstructed recommendation actions according to the account status feature corresponding to the account status of the current account, and adjusting the corresponding reconstructed recommendation action by using the action adjustment information to obtain an adjusted recommendation action;
[0008] Determine the action recommendation information corresponding to each of the adjusted recommended actions: The action recommendation information is used to characterize the predicted interaction degree between the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommended action to the current account;
[0009] Determine the adjusted recommended actions whose action recommendation information meets the preset conditions as the target recommended actions; the target recommended actions are used for the target application to recommend content matching the target recommended actions to the current account.
[0010] In a possible implementation manner, the determining the action recommendation information corresponding to each of the adjusted recommended actions includes:
[0011] Input the account status of the current account and the adjusted recommended action into an action recommendation information generation network to obtain the action recommendation information corresponding to the adjusted recommended action;
[0012] Among them, the action recommendation information generation network is trained using the account status of the sample account, the sample recommended action, and the actual recommendation information corresponding to the sample recommended action; the actual recommendation information is determined according to the actual interaction degree between the sample account, the target application, and the content recommended by the target application after the target application recommends content matching the sample recommended action to the sample account.
[0013] In a possible implementation manner, the determining the action adjustment information corresponding to each of the reconstructed recommended actions according to the account status characteristics corresponding to the account status of the current account includes:
[0014] Input the account status of the current account and the reconstructed recommended action into an action residual generation network to obtain the action residual corresponding to the reconstructed recommended action; wherein, the action residual is used to characterize the action adjustment information corresponding to the reconstructed recommended action;
[0015] The adjusting the corresponding reconstructed recommended action using the action adjustment information to obtain the adjusted recommended action includes:
[0016] Superimpose the action residual corresponding to the reconstructed recommended action on the corresponding reconstructed recommended action to obtain the adjusted recommended actions corresponding to each of the reconstructed recommended actions.
[0017] In a possible implementation manner, the action residual generation network includes a status feature extraction module and an action residual generation module, and the inputting the account status of the current account and the reconstructed recommended action into the action residual generation network to obtain the action residual corresponding to the reconstructed recommended action includes:
[0018] Input the account status of the current account into the status feature extraction module to obtain the account status feature corresponding to the account status of the current account;
[0019] Input the account status feature and the reconstructed recommendation action into the action residual generation module to obtain the action residual corresponding to the reconstructed recommendation action.
[0020] In a possible implementation manner, the account status includes an account session status and an account request status. The account session status is used to represent the interaction status with the target application, and the account request status is used to represent the interaction status with the content recommended by the target application. The status feature extraction module includes a session status encoder and a request status encoder. The step of inputting the account status of the current account into the status feature extraction module to obtain the account status feature corresponding to the account status of the current account includes:
[0021] Input the account session status into the session status encoder so that the session status encoder extracts the session status feature corresponding to the account session status, and input the account request status into the request status encoder so that the session status encoder extracts the request status feature corresponding to the account request status;
[0022] Use the session status feature and the request status feature as the account status feature.
[0023] In a possible implementation manner, the method further includes:
[0024] Input the account status of the sample account and the sample recommendation action into the recommendation action reconstruction network to be trained to obtain the predicted recommendation action for the sample account;
[0025] Based on the difference between the predicted recommendation action and the sample recommendation action, train the recommendation action reconstruction network to be trained until the preset training end condition is met to obtain the trained recommendation action reconstruction network.
[0026] In a possible implementation manner, the method further includes:
[0027] Input the account status of the sample account and the predicted recommendation action into the action residual generation network to be trained to obtain the action residual corresponding to the predicted recommendation action;
[0028] Superimpose the action residual corresponding to the predicted recommendation action on the predicted recommendation action to obtain the superimposed recommendation action corresponding to the predicted recommendation action;
[0029] Input the account status of the sample account and the superimposed recommended action into the action recommendation information generation network to be trained, and obtain the action recommendation information corresponding to the superimposed recommended action;
[0030] Based on the action recommendation information corresponding to the superimposed recommended action, train the action residual generation network to be trained and the action recommendation information generation network to be trained using the method of reinforcement learning until the preset training end condition is met, and obtain the trained action residual generation network and the trained action recommendation information generation network; wherein, the action residual generation network is the policy function in the method of reinforcement learning, the action recommendation information generation network is the value function in the method of reinforcement learning, and the action recommendation information corresponding to the superimposed recommended action is the reward data for the policy function in the method of reinforcement learning.
[0031] In a possible implementation manner, the action residual generation network includes a state feature extraction module, and the state feature extraction module is used to output the session state feature corresponding to the account session state in the account status, and the account session state is used to represent the interaction state with the target application; the training of the action residual generation network to be trained and the action recommendation information generation network to be trained using the method of reinforcement learning based on the action recommendation information corresponding to the superimposed recommended action includes:
[0032] In the process of updating the parameters of the state feature extraction module using the method of reinforcement learning, use the first regularization term and the second regularization term to control the parameter update process of the state feature extraction module;
[0033] Wherein, the first regularization term is used to maximize the mutual information between the session state feature and the reward data, and the second regularization term is used to minimize the entropy value of the session state feature distribution.
[0034] In a possible implementation manner, the recommended action reconstruction network includes an action encoder and an action decoder, and the inputting the account status of the sample account and the sample recommended action into the recommended action reconstruction network to be trained to obtain the predicted recommended action for the sample account includes:
[0035] Input the account status of the sample account and the sample recommended action into the action encoder, so that the action encoder projects the vector corresponding to the sample recommended action into a latent variable space in response to the account status of the sample account, and samples at least one predicted latent variable in the latent variable space;
[0036] Input the at least one predicted latent variable and the account status of the sample account into the action decoder, so that the action decoder decodes the at least one predicted latent variable to obtain a predicted recommendation action for the sample account.
[0037] In a possible implementation manner, the inputting the account status of the current account into the recommendation action reconstruction network to obtain at least one reconstructed recommendation action includes:
[0038] Input the account status of the current account into the action encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account status of the current account;
[0039] Input the at least one target latent variable and the account status of the current account into the action decoder, so that the action decoder decodes the at least one target latent variable to obtain the at least one reconstructed recommendation action.
[0040] According to the second aspect of the embodiments of the present disclosure, a content recommendation device is provided, including:
[0041] A reconstruction unit, configured to execute obtaining the account status of the current account, inputting the account status of the current account into the recommendation action reconstruction network to obtain at least one reconstructed recommendation action: the recommendation action reconstruction network is trained by using the account status of the sample account and the corresponding sample recommendation actions; the sample recommendation actions are content recommendation actions generated by a preset recommendation strategy in response to the user status of the sample account:
[0042] An adjustment unit, configured to execute determining action adjustment information corresponding to each of the reconstructed recommendation actions according to the account status feature corresponding to the account status of the current account, and adjusting the corresponding reconstructed recommendation action by using the action adjustment information to obtain an adjusted recommendation action;
[0043] A determination unit, configured to execute determining action recommendation information corresponding to each of the adjusted recommendation actions: the action recommendation information is used to characterize the predicted interaction degree among the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account;
[0044] A recommendation unit, configured to execute determining the adjusted recommendation action whose action recommendation information meets a preset condition as the target recommendation action; the target recommendation action is used for the target application to recommend content matching the target recommendation action to the current account.
[0045] In an exemplary embodiment, the determining unit is specifically configured to input the account status of the current account and the adjusted recommendation action into an action recommendation information generation network to obtain action recommendation information corresponding to the adjusted recommendation action; wherein, the action recommendation information generation network is trained by using the account status of the sample account, the sample recommendation action, and the actual recommendation information corresponding to the sample recommendation action; the actual recommendation information is determined according to the actual interaction degree between the sample account and the target application and the content recommended by the target application after the target application recommends content matching the sample recommendation action to the sample account.
[0046] In an exemplary embodiment, the adjusting unit is specifically configured to input the account status of the current account and the reconstructed recommendation action into an action residual generation network to obtain an action residual corresponding to the reconstructed recommendation action; wherein, the action residual is used to represent the action adjustment information corresponding to the reconstructed recommendation action; the step of adjusting the corresponding reconstructed recommendation action by using the action adjustment information to obtain an adjusted recommendation action includes: superimposing the action residual corresponding to the reconstructed recommendation action on the corresponding reconstructed recommendation action to obtain adjusted recommendation actions corresponding to each of the reconstructed recommendation actions.
[0047] In an exemplary embodiment, the action residual generation network includes a status feature extraction module and an action residual generation module, and the adjusting unit is specifically configured to input the account status of the current account into the status feature extraction module to obtain an account status feature corresponding to the account status of the current account; input the account status feature and the reconstructed recommendation action into the action residual generation module to obtain an action residual corresponding to the reconstructed recommendation action.
[0048] In an exemplary embodiment, the account status includes an account session status and an account request status, the account session status is used to represent the interaction status with the target application, the account request status is used to represent the interaction status with the content recommended by the target application, the status feature extraction module includes a session status encoder and a request status encoder, and the adjusting unit is specifically configured to input the account session status into the session status encoder to enable the session status encoder to extract a session status feature corresponding to the account session status, and input the account request status into the request status encoder to enable the session status encoder to extract a request status feature corresponding to the account request status; use the session status feature and the request status feature as the account status feature.
[0049] In an exemplary embodiment, the device is configured to perform the following operations: input the account status of the sample account and the sample recommendation action into the recommendation action reconstruction network to be trained, and obtain a predicted recommendation action for the sample account; based on the difference between the predicted recommendation action and the sample recommendation action, train the recommendation action reconstruction network to be trained until a preset training end condition is satisfied, and obtain the trained recommendation action reconstruction network.
[0050] In an exemplary embodiment, the device is configured to perform the following operations: input the account status of the sample account and the predicted recommendation action into the action residual generation network to be trained, and obtain an action residual corresponding to the predicted recommendation action; superimpose the action residual corresponding to the predicted recommendation action on the predicted recommendation action to obtain a superimposed recommendation action corresponding to the predicted recommendation action; input the account status of the sample account and the superimposed recommendation action into the action recommendation information generation network to be trained, and obtain action recommendation information corresponding to the superimposed recommendation action; based on the action recommendation information corresponding to the superimposed recommendation action, train the action residual generation network to be trained and the action recommendation information generation network to be trained by using a reinforcement learning method until a preset training end condition is satisfied, and obtain the trained action residual generation network and the trained action recommendation information generation network; wherein, the action residual generation network is a policy function in the reinforcement learning method, the action recommendation information generation network is a value function in the reinforcement learning method, and the action recommendation information corresponding to the superimposed recommendation action is reward data for the policy function in the reinforcement learning method.
[0051] In an exemplary embodiment, the action residual generation network includes a state feature extraction module, and the state feature extraction module is configured to output a session state feature corresponding to the account session state in the account status, and the account session state is used to represent an interaction state with the target application; the device is configured to perform the following operation during the process of updating the parameters of the state feature extraction module by using the reinforcement learning method: use a first regularization term and a second regularization term to control the parameter update process of the state feature extraction module; wherein, the first regularization term is used to maximize the mutual information between the session state feature and the reward data, and the second regularization term is used to minimize the entropy value of the session state feature distribution.
[0052] In an exemplary embodiment, the recommendation action reconstruction network includes an action encoder and an action decoder. The device is configured to perform the operation of inputting the account status of the sample account and the sample recommendation action into the action encoder, so that the action encoder projects the vector corresponding to the sample recommendation action into a latent variable space in response to the account status of the sample account, and samples at least one predicted latent variable in the latent variable space; inputting the at least one predicted latent variable and the account status of the sample account into the action decoder, so that the action decoder decodes the at least one predicted latent variable to obtain a predicted recommendation action for the sample account.
[0053] In an exemplary embodiment, the reconstruction unit is specifically configured to perform the operation of inputting the account status of the current account into the action encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account status of the current account; inputting the at least one target latent variable and the account status of the current account into the action decoder, so that the action decoder decodes the at least one target latent variable to obtain the at least one reconstructed recommendation action.
[0054] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the content recommendation method as described in the first aspect or any possible implementation manner of the first aspect.
[0055] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the content recommendation method as described in the first aspect or any possible implementation manner of the first aspect.
[0056] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product. The program product includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of the device reads and executes the computer program from the readable storage medium, so that the device executes the content recommendation method as described in any possible implementation manner of the first aspect.
[0057] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: By obtaining the account status of the current account and inputting the account status of the current account into the recommendation action reconstruction network, at least one reconstructed recommendation action is obtained. The recommendation action reconstruction network is trained by using the account status of the sample account and the sample recommendation actions generated by the preset recommendation strategy in response to the account status of the sample account. And according to the account status features corresponding to the account status of the current account, the action adjustment information corresponding to each reconstructed recommendation action is determined, and the corresponding reconstructed recommendation action is adjusted by using the action adjustment information to obtain the adjusted recommendation action. In this way, through the recommendation action reconstruction network, the video recommendation actions that the online deployment strategy would output in response to the account status of the current account are reconstructed as much as possible, so as to limit the recommendation strategy space within the range where the online deployment strategy is located.
[0058] At the same time, by using the account status features corresponding to the account status of the current account, the action adjustment information corresponding to each reconstructed recommendation action is determined, and each reconstructed recommendation action is adjusted by using the action adjustment information to obtain the adjusted recommendation action. The action recommendation information corresponding to each adjusted recommendation action is obtained. The action recommendation information is used to represent a long-term indicator of the predicted interaction degree among the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account. Finally, by determining the adjusted recommendation action whose action recommendation information meets the preset condition as the target recommendation operation to instruct the target application to recommend content matching the target recommendation action to the current account, it is possible to optimize the long-term indicator of the video recommendation action output within the range of the online deployment strategy, and further easily control the deviation between the learned strategy and the current deployment strategy, so that when optimizing the long-term indicator of the content recommendation strategy, there is no need to search and optimize in the entire huge recommendation strategy space, reducing the optimization difficulty and cost, and realizing the efficient optimization of the content recommendation strategy focusing on the long-term indicator. As a result, the optimized content recommendation strategy can often accurately recommend the content required by the user to the user, improving the quantity of content obtained by the user through the target application and the efficiency of obtaining content from the target application.
[0059] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0061] Figure 1 It is an application environment diagram of a content recommendation method shown according to an exemplary embodiment.
[0062] Figure 2 is a flowchart of a content recommendation method shown according to an exemplary embodiment.
[0063] Figure 3 is a working flowchart of a content recommendation method shown according to an exemplary embodiment.
[0064] Figure 4 is a flowchart of another content recommendation method shown according to an exemplary embodiment.
[0065] Figure 5 is a model framework diagram of a content recommendation method shown according to an exemplary embodiment.
[0066] Figure 6 is a schematic diagram of the reinforcement learning process of a content recommendation method shown according to an exemplary embodiment.
[0067] Figure 7 is a block diagram of a content recommendation method device shown according to an exemplary embodiment.
[0068] Figure 8 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0069] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0070] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0071] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0072] The content recommendation method provided by the present disclosure can be applied to, for example Figure 1In the application environment shown, the terminal 102 where the target is used communicates with the recommendation server 104 through the network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers.
[0073] In actual applications, the recommendation server 104 obtains the account status of the current account, inputs the account status of the current account into the recommendation action reconstruction network, and obtains at least one reconstructed recommendation action: the recommendation action reconstruction network is obtained by training with the account status of the sample account and the corresponding sample recommendation action; the sample recommendation action is a content recommendation action generated by the preset recommendation strategy in response to the account status of the sample account: according to the account status characteristics corresponding to the account status of the current account, the action adjustment information corresponding to each reconstructed recommendation action is determined, and the corresponding reconstructed recommendation action is adjusted using the action adjustment information to obtain the adjusted recommendation action; the action recommendation information corresponding to each adjusted recommendation action is determined: the action recommendation information is used to characterize the predicted degree of interaction between the current account and the target application, and the content recommended by the target application after the target application recommends the content matching the adjusted recommendation action to the current account; the adjusted recommendation action whose action recommendation information meets the preset conditions is determined as the target recommendation action; the target recommendation action is used for the target application terminal 102 to recommend the content matching the target recommendation action to the current account. In actual applications, the above-mentioned content recommendation method can also be executed on the terminal 102. The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The recommendation server 104 may be implemented using an independent server or a server cluster consisting of multiple servers.
[0074] Figure 2 is a flow chart of a content recommendation method according to an exemplary embodiment. Figure 2 As shown, the content recommendation method can be used for Figure 1 The recommended server includes the following steps.
[0075] In step S210, the account status of the current account is obtained, and the account status of the current account is input into the recommended action reconstruction network to obtain at least one reconstructed recommended action.
[0076] The current account may refer to a user account that currently initiates the content acquisition request.
[0077] Among them, the content can refer to media content such as pictures, videos, and music provided by the target application. Taking the target application as a short video as an example, the content can refer to the short video.
[0078] Among them, the account status is used to depict the current information of the user corresponding to the current account. In practical applications, the account status can include at least one of user inherent attribute information, user historical behavior information, preference information, and current status information in the Context information.
[0079] Among them, the recommendation action can be used to instruct the target application to recommend content matching the recommendation action to the current account. In practical applications, the recommendation action can be represented by an action vector. The action vector can be a representation vector of the items that can be recommended.
[0080] For example, the recommendation server can calculate the similarity between the representation vectors corresponding to each candidate content and the action vector corresponding to the recommendation action respectively, and use the candidate content corresponding to the top N in the similarity ranking as the content recommended by the target application to the current account.
[0081] Among them, the recommendation action reconstruction network is trained by using the account status of the sample account and the sample recommendation action generated in response to the account status of the sample account according to a preset recommendation strategy.
[0082] Among them, the preset recommendation strategy can refer to the content recommendation strategy already used by the recommendation server. In practical applications, the preset recommendation strategy can refer to the online deployment strategy (policy space).
[0083] That is to say, by using the account status of the sample account and the sample recommendation action generated in response to the account status of the sample account according to the preset recommendation strategy to train the recommendation action reconstruction network, the online deployment strategy is reconstructed, so that the recommendation action reconstruction network can output a video recommendation action consistent with the online deployment strategy in response to the account status of the current account.
[0084] It should be noted that the inference process and training process of the recommendation action reconstruction network will be introduced in detail below, and will not be elaborated here.
[0085] In specific implementation, in response to the condition that the current account initiates a content acquisition request (for example, the current user starts the target application and inputs a video recommendation operation to the target application), the recommendation server can obtain the account status of the current account and input the account status of the current account into the recommendation action reconstruction network to obtain at least one reconstructed recommendation action.
[0086] Specifically, the recommendation server inputs the account status of the current account into the recommendation action reconstruction network, and the recommendation action reconstruction network adopts at least one target latent variable in the latent variable space in response to the account status of the current account; the recommendation action reconstruction network decodes the target latent variable and restores it to at least one reconstructed recommendation action.
[0087] Among them, this latent variable space is obtained by the recommendation action reconstruction network projecting the sample recommendation action into a latent variable space during the training process in response to the account status of the sample account.
[0088] In step S220, according to the account status features corresponding to the account status of the current account, determine the action adjustment information corresponding to each of the reconstructed recommendation actions, and use the action adjustment information to adjust the corresponding reconstructed recommendation actions to obtain adjusted recommendation actions.
[0089] In a specific implementation, after the recommendation server obtains at least one reconstructed recommendation action for the current account, the recommendation server can extract the account status features corresponding to the account status of the current account, and determine the action adjustment information corresponding to each reconstructed recommendation action according to the account status features corresponding to the account status of the current account; then, the recommendation server uses the action adjustment information to adjust each reconstructed recommendation action to obtain adjusted recommendation actions.
[0090] Specifically, the recommendation server can input the account status of the current account and at least one reconstructed recommendation action for the current account into the action residual generation network, extract the account status features corresponding to the account status of the current account through the action residual generation network, and predict the action residuals corresponding to each reconstructed recommendation action according to the account status features corresponding to the account status of the current account. Then, the recommendation server applies each action residual to the corresponding reconstructed recommendation action to obtain adjusted recommendation actions. In practical applications, when using vectors to represent the reconstructed recommendation actions, the action adjustment information may include amplitude adjustment information and direction adjustment information for the vector representation corresponding to the reconstructed recommendation action, that is, the action adjustment information includes at least one of the action adjustment amplitude and the action adjustment direction.
[0091] It should be noted that the inference process and training process of this action residual generation network will be introduced in detail below, and will not be elaborated here for the time being.
[0092] In step S230, determine the action recommendation information corresponding to each of the adjusted recommendation actions.
[0093] Among them, the action recommendation information is used to characterize the predicted interaction degree among the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account. In practical applications, the action recommendation information may include an action recommendation value, which may also be named action value.
[0094] In practical applications, the action recommendation value can be calculated based on the predicted duration until the current account next starts the target application and the predicted number of clicks on the content recommended by the target application, so as to characterize the predicted interaction frequency among the current account, the target application, and the content recommended by the target application.
[0095] In a specific implementation, after the recommendation server obtains the adjusted recommendation action, it can obtain the action recommendation value corresponding to each adjusted recommendation action. Specifically, the recommendation server can input the account status of the current account and the adjusted recommendation action into the action value generation network to obtain the action recommendation value corresponding to the adjusted recommendation action. It should be noted that the inference process and training process of this action value generation network will be introduced in detail below, and will not be elaborated here.
[0096] In step S240, the adjusted recommendation action for which the action recommendation information meets the preset condition is determined as the target recommendation action; the target recommendation action is used to instruct the target application to recommend content matching the target recommendation action to the current account.
[0097] In a specific implementation, after the recommendation server determines the action recommendation value corresponding to each adjusted recommendation action, the recommendation server can select the target recommendation operation with the highest action recommendation value from among the adjusted recommendation actions, so as to instruct the target application to recommend content matching the target recommendation action to the current account.
[0098] Specifically, the recommendation server can calculate the similarity between the feature vector corresponding to each candidate content and the action vector corresponding to the target recommendation action respectively, and use the candidate content corresponding to the top N similarity rankings as the content recommended by the target application to the current account.
[0099] For the convenience of understanding by those skilled in the art, Figure 3An example provides a workflow diagram of a content recommendation method. Among them, the workflow can be divided into three steps. The first step is the Reconstruction process: the recommendation server can obtain the account status of the current account and input the account status of the current account into the recommendation action reconstruction network to obtain at least one reconstructed recommendation action, so as to reconstruct the recommendation actions that the online deployment model may make. The second step is the Prediction process: the recommendation server can input the account status of the current account and at least one reconstructed recommendation action for the current account into the action residual generation network. Through the action residual generation network, the account status features corresponding to the account status of the current account are extracted, and according to the account status features corresponding to the account status of the current account, the action residuals corresponding to each reconstructed recommendation action are predicted. Then, the recommendation server applies each action residual to the corresponding reconstructed recommendation action to obtain the adjusted recommendation action. The third step is the Selection process: the recommendation server can input the account status of the current account and the adjusted recommendation action into the action value generation network to obtain the action recommendation value corresponding to the adjusted recommendation action. The recommendation server can select the target recommendation operation with the highest action recommendation value among the adjusted recommendation actions based on the action recommendation values corresponding to each adjusted recommendation action, so as to instruct the target application to recommend content matching the target recommendation action to the current account.
[0100] In the above content recommendation method, by obtaining the account status of the current account and inputting the account status of the current account into the recommendation action reconstruction network, at least one reconstructed recommendation action is obtained. The recommendation action reconstruction network is trained by using the account status of the sample account and the sample recommendation actions generated by the preset recommendation strategy in response to the account status of the sample account. And according to the account status features corresponding to the account status of the current account, the action adjustment information corresponding to each reconstructed recommendation action is determined, and the corresponding reconstructed recommendation action is adjusted by using the action adjustment information to obtain the adjusted recommendation action. In this way, through the recommendation action reconstruction network, the video recommendation actions that the online deployment strategy will output in response to the account status of the current account are reconstructed as much as possible, so as to limit the recommendation strategy space within the range of the online deployment strategy.
[0101] Meanwhile, by utilizing the account status features corresponding to the account status of the current account, the action adjustment information corresponding to each reconstructed recommendation action is determined, and each reconstructed recommendation action is adjusted by using the action adjustment information to obtain the adjusted recommendation actions; the action recommendation information corresponding to each adjusted recommendation action is obtained: the action recommendation information is used to represent a long-term metric of the predicted interaction degree among the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account; finally, the adjusted recommendation actions that satisfy the preset conditions are determined as the target recommendation operations to instruct the target application to recommend content matching the target recommendation action to the current account, thereby realizing the optimization of the long-term metric for the video recommendation actions output within the scope of the online deployment strategy, and further realizing the easy control of the deviation between the learned strategy and the current deployment strategy, so that when optimizing the long-term metric of the content recommendation strategy, there is no need to search and optimize in the entire huge recommendation strategy space, reducing the optimization difficulty and cost, and realizing the efficient optimization of the content recommendation strategy focusing on the long-term metric, so that the optimized content recommendation strategy can often accurately recommend the content required by the user to the user, improving the quantity of content obtained by the user through the target application and the efficiency of obtaining content from the target application.
[0102] In one exemplary embodiment, determining the action recommendation information corresponding to each adjusted recommendation action includes: inputting the account status of the current account and the adjusted recommendation action into an action recommendation information generation network to obtain the action recommendation information corresponding to the adjusted recommendation action.
[0103] Among them, the action recommendation information generation network is trained by using the account status of the sample account, the sample recommendation action, and the actual recommendation information corresponding to the sample recommendation action. In practical applications, the action recommendation information generation network can be named the action value generation network.
[0104] Among them, the actual recommendation information is determined according to the actual interaction degree among the sample account, the target application, and the content recommended by the target application after the target application recommends content matching the sample recommendation action to the sample account.
[0105] In practical applications, the actual recommendation information can be an actual recommendation value. The actual recommendation value can be calculated according to the actual duration until the sample account next starts the target application and the actual number of clicks of the sample account on the content recommended by the target application, so as to represent the actual interaction frequency among the sample account, the target application, and the content recommended by the target application.
[0106] In a specific implementation, when the recommendation server obtains the action recommendation values corresponding to the adjusted recommendation actions, the recommendation server can input the account status of the current account and the adjusted recommendation actions into the action value generation network to obtain the action recommendation values corresponding to the adjusted recommendation actions, so as to predict the predicted duration until the target application is next launched by the current account and the predicted number of clicks on the content recommended by the target application by the current account, that is, the predicted interaction frequency between the current account and the target application and the content recommended by the target application.
[0107] In the technical solution of this embodiment, by inputting the account status of the current account and the adjusted recommendation actions into the action value generation network, the action recommendation values corresponding to the adjusted recommendation actions are obtained. Among them, the action value generation network is trained using the account status of the sample account, the sample recommendation actions, and the actual recommendation values corresponding to the sample recommendation actions. In this way, the action value generation network can predict the action recommendation values that can effectively represent the predicted interaction frequency between the current account and the target application and the content recommended by the target application according to the account status of the current account and the adjusted recommendation actions.
[0108] In an exemplary embodiment, according to the account status features corresponding to the account status of the current account, determining the action adjustment information corresponding to each reconstructed recommendation action includes: inputting the account status of the current account and the reconstructed recommendation actions into the action residual generation network to obtain the action residuals corresponding to the reconstructed recommendation actions; where the action residuals are used to represent the action adjustment information corresponding to the reconstructed recommendation actions; adjusting the corresponding reconstructed recommendation actions using the action adjustment information to obtain the adjusted recommendation actions, including: superimposing the action residuals corresponding to the reconstructed recommendation actions on the corresponding reconstructed recommendation actions to obtain the adjusted recommendation actions corresponding to each reconstructed recommendation action.
[0109] Among them, the action residuals are used to represent the action adjustment methods corresponding to the reconstructed recommendation actions.
[0110] In a specific implementation, when the recommendation server determines the action adjustment methods corresponding to each reconstructed recommendation action according to the account status features corresponding to the account status of the current account, it specifically includes: the recommendation server can input the account status of the current account and the reconstructed recommendation actions into the action residual generation network, so that the action residual generation network extracts the account status features corresponding to the account status of the current account, and uses the account status features and the reconstructed recommendation actions to generate the action residuals corresponding to each reconstructed recommendation action, so as to adaptively determine the action adjustment methods corresponding to each reconstructed recommendation action based on the account status of the current account.
[0111] Then, after the recommendation server determines the action residuals corresponding to each reconstructed recommendation action, the recommendation server superimposes the action residuals corresponding to the reconstructed recommendation action onto the reconstructed recommendation action to adjust each reconstructed recommendation action in an action adjustment manner, obtaining the adjusted recommendation actions.
[0112] In the technical solution of this embodiment, by inputting the account status of the current account and the reconstructed recommendation action into the action residual generation network, obtaining the action residuals corresponding to the reconstructed recommendation action, and superimposing the action residuals corresponding to the reconstructed recommendation action onto the reconstructed recommendation action, the adjusted recommendation actions corresponding to the reconstructed recommendation action can be obtained, which can optimize the video recommendation actions output within the scope of the online deployment strategy, and further easily control the deviation between the learned strategy and the current deployment strategy, so that when optimizing the long-term metrics of the content recommendation strategy, there is no need to search and optimize in the huge entire recommendation strategy space.
[0113] In an exemplary embodiment, the action residual generation network includes a state feature extraction module and an action residual generation module. The state feature extraction module includes a session state encoder and a request state encoder. The account status includes a user session state and a user request state. The user session state is used to represent the interaction state with the target application, and the user request state is used to represent the interaction state with the content recommended by the target application.
[0114] Inputting the account status of the current account and the reconstructed recommendation action into the action residual generation network to obtain the action residuals corresponding to the reconstructed recommendation action includes: inputting the account status of the current account into the state feature extraction module to obtain the account status features corresponding to the account status of the current account; wherein, inputting the user session state into the session state encoder to enable the session state encoder to extract the session state features corresponding to the user session state, and inputting the user request state into the request state encoder to enable the session state encoder to extract the request state features corresponding to the user request state; using the session state features and the request state features as the account status features. Inputting the account status features and the reconstructed recommendation action into the action residual generation module to obtain the action residuals corresponding to the reconstructed recommendation action.
[0115] In practical applications, the event of a user starting the target application can be defined as a session; each event of a user clicking on the content recommended by the target application can be defined as a request in this session.
[0116] Based on the above definitions, the interaction state between the current account and the target application is named the user session state; the interaction state between the current account and the content recommended by the target application is named the user request state. The account status includes the user session state and the user request state.
[0117] In a specific implementation, when the action residual generation network is a residual actor, the residual actor can include three components, namely, a high-level state encoder, a low-level state encoder, and a residual sub-actor.
[0118] Among them, the high-level state encoder is used to extract the state features of the user session state.
[0119] Among them, the low-level state encoder is used to extract the state features of the user request state.
[0120] In the process that the recommendation server inputs the account state and the reconstructed recommendation action of the current account into the action residual generation network to obtain the action residual corresponding to the reconstructed recommendation action, the recommendation server can input the session state features in the account state into the high-level state encoder so that the high-level state encoder extracts the session state features corresponding to the user session state, and input the user request state in the account state into the low-level state encoder so that the low-level state encoder extracts the request state features corresponding to the user request state.
[0121] Then, the recommendation server inputs the session state features, the request state features, and the reconstructed recommendation action into the residual sub-actor to obtain the action residual corresponding to the reconstructed recommendation action.
[0122] For the technical solution of this embodiment, considering the unique session-request double-layer structure in sequential recommendation, we decompose the state into a high-level state and a low-level state to facilitate the representation of the state, and then can predict effective residuals so that the action residuals added to the reconstructed action can improve the reconstructed action.
[0123] In an exemplary embodiment, the method further includes: inputting the account state and the sample recommendation action of the sample account into the recommendation action reconstruction network to be trained to obtain the predicted recommendation action for the sample account; training the recommendation action reconstruction network to be trained based on the difference between the predicted recommendation action and the sample recommendation action until the trained recommendation action reconstruction network is obtained.
[0124] In a specific implementation, before using the recommended action reconstruction network to obtain at least one reconstructed recommended action corresponding to the account status of the current account, the recommended action reconstruction network needs to be trained. The recommendation server can use the account status of the sample account and the sample recommended action, that is, the [status, action] binary tuple, as the training sample data for training the recommended action reconstruction network. Among them, the recommendation server can input the account status of the sample account and the sample recommended action into the recommended action reconstruction network to be trained, and obtain the predicted recommended action for the sample account.
[0125] Then, the recommendation server updates the parameters of the recommended action reconstruction network to be trained by comparing the difference between the predicted recommended action and the sample recommended action, and uses the gradient descent and backpropagation algorithms to train the recommended action reconstruction network to be trained until the trained recommended action reconstruction network is obtained.
[0126] In an exemplary embodiment, the recommended action reconstruction network includes an action encoder and an action decoder. Inputting the account status of the sample account and the sample recommended action into the recommended action reconstruction network to be trained to obtain the predicted recommended action for the sample account includes: inputting the account status of the sample account and the sample recommended action into the action encoder, so that the action encoder projects the vector corresponding to the sample recommended action into a latent variable space in response to the account status of the sample account, and samples at least one predicted latent variable in the latent variable space; inputting at least one predicted latent variable and the account status of the sample account into the action decoder, so that the action decoder decodes at least one predicted latent variable to obtain the predicted recommended action for the sample account.
[0127] In a specific implementation, when the recommended action reconstruction network is a conditional variational auto-encoder, the conditional variational auto-encoder includes an action encoder (CVAE-Encoder) and an action decoder (CVAE-Decoder). In the process of the recommendation server inputting the account status of the sample account and the sample recommended action into the recommended action reconstruction network to be trained and obtaining the predicted recommended action for the sample account, the recommendation server can input the account status of the sample account and the sample recommended action into the action encoder, so that the action encoder projects the vector corresponding to the sample recommended action into a latent variable space in response to the account status of the sample account, and samples at least one predicted latent variable in the latent variable space; then, the recommendation server inputs at least one predicted latent variable and the account status of the sample account into the action decoder, so that the action decoder decodes at least one predicted latent variable to obtain the predicted recommended action for the sample account.
[0128] In the technical solution of this embodiment, by inputting the account status of the sample account and the sample recommendation action into the action encoder, the action encoder projects the vector corresponding to the sample recommendation action into a latent variable space in response to the account status of the sample account, and samples at least one predicted latent variable in the latent variable space; and inputs at least one predicted latent variable and the account status of the sample account into the action decoder, so that the action decoder decodes at least one predicted latent variable to obtain a predicted recommendation action for the sample account, so that the trained recommendation action reconstruction network can subsequently reconstruct as much as possible the video recommendation action that the online deployment policy may output in response to the account status of the current account.
[0129] In an exemplary embodiment, inputting the account status of the current account into the recommendation action reconstruction network to obtain at least one reconstructed recommendation action includes: inputting the account status of the current account into the action encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account status of the current account; and inputting at least one target latent variable and the account status of the current account into the action decoder, so that the action decoder decodes at least one target latent variable to obtain at least one reconstructed recommendation action.
[0130] In a specific implementation, when the recommendation server inputs the account status of the current account into the recommendation action reconstruction network to obtain at least one reconstructed recommendation action, the recommendation server inputs the account status of the current account into the action encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account status of the current account; then, the recommendation server inputs at least one target latent variable and the account status of the current account into the action decoder, so that the action decoder decodes at least one target latent variable to obtain at least one reconstructed recommendation action.
[0131] In the technical solution of this embodiment, by inputting the account status of the current account into the action encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account status of the current account; and inputting at least one target latent variable and the account status of the current account into the action decoder, so that the action decoder decodes at least one target latent variable to obtain at least one reconstructed recommendation action, thereby enabling the recommendation action reconstruction network to reconstruct as much as possible the video recommendation action that the online deployment policy will output in response to the account status of the current account, thereby restricting the recommendation policy space within the range where the online deployment policy is located.
[0132] In an exemplary embodiment, the method further includes: inputting the account status of the sample account and the predicted recommendation action into the action residual generation network to be trained to obtain the action residual corresponding to the predicted recommendation action; superimposing the action residual corresponding to the predicted recommendation action on the predicted recommendation action to obtain the superimposed recommendation action corresponding to the predicted recommendation action; inputting the account status of the sample account and the superimposed recommendation action into the action value generation network to be trained to obtain the action recommendation value corresponding to the superimposed recommendation action; and training the action residual generation network to be trained and the action value generation network to be trained by using the method of reinforcement learning based on the action recommendation value corresponding to the superimposed recommendation action until the trained action residual generation network and the trained action value generation network are obtained.
[0133] In a specific implementation, the recommendation server uses the actor-critic framework in deep reinforcement learning to train the action residual generation network to be trained and the action value generation network to be trained. That is, the action residual generation network is used as the policy function in the method of reinforcement learning, the action value generation network is used as the value function in the method of reinforcement learning, and the action recommendation information corresponding to the superimposed recommendation action is the reward data for the policy function in the method of reinforcement learning.
[0134] Specifically, the recommendation server may input the predicted recommendation action output by the recommendation action reconstruction network for the account status of the sample account together with the account status of the sample account into the action residual generation network to be trained to obtain the action residual corresponding to the predicted recommendation action; then, the recommendation server superimposes the action residual corresponding to the predicted recommendation action on the predicted recommendation action to obtain the superimposed recommendation action corresponding to the predicted recommendation action. Then, the recommendation server inputs the account status of the sample account and the superimposed recommendation action into the action value generation network to be trained to obtain the action recommendation value corresponding to the superimposed recommendation action; then, the recommendation server uses the action recommendation value corresponding to the superimposed recommendation action as the reward for the action residual generation network to be trained, and uses the method of policy gradient to train the action residual generation network to be trained and the action value generation network to be trained by using the method of reinforcement learning until the trained action residual generation network and the trained action value generation network are obtained.
[0135] In an exemplary embodiment, the action residual generation network includes a state feature extraction module, which is configured to output session state features corresponding to the account session state in the account state, and the account session state is used to represent the interaction state with the target application; based on the action recommendation information corresponding to the superimposed recommended actions, the method of reinforcement learning is used to train the to-be-trained action residual generation network and the to-be-trained action recommendation information generation network, including: in the process of updating the parameters of the state feature extraction module by using the method of reinforcement learning, a first regularization term and a second regularization term are used to control the parameter update process of the state feature extraction module; wherein, the first regularization term is used to maximize the mutual information between the session state features and the reward data, and the second regularization term is used to minimize the entropy value of the session state feature distribution.
[0136] Specifically, in the process of training the to-be-trained action residual generation network, in the process of updating the parameters of the high-level state encoder in the to-be-trained action residual generation network, the recommendation server can use two information-theory-based regularization terms to control the parameter update of the high-level state encoder. Among them, one regularization term of the two regularization terms is used to maximize the mutual information between the session state features and the reward, and the other regularization term in the information-theory-based regularization term is used to minimize the entropy of the state feature distribution, ensuring that the learned features are both concise and clear and contain as much information related to user retention as possible.
[0137] Figure 4 is a flowchart of another content recommendation method shown according to an exemplary embodiment, as Figure 3 shown, this method is used for Figure 1 the recommendation server in, and includes the following steps.
[0138] In step S410, the account state of the current account is obtained, and the account state of the current account is input into the recommended action reconstruction network to obtain at least one reconstructed recommended action: the recommended action reconstruction network is trained by using the account state of the sample account and the corresponding sample recommended action; the sample recommended action is a content recommendation action generated by a preset recommendation strategy in response to the account state of the sample account.
[0139] In step S420, the account state of the current account and the reconstructed recommended action are input into the action residual generation network to obtain the action residual corresponding to the reconstructed recommended action.
[0140] In step S430, the action residual corresponding to the reconstructed recommended action is superimposed on the reconstructed recommended action to obtain the adjusted recommended action corresponding to the reconstructed recommended action.
[0141] In step S440, the account status of the current account and the adjusted recommended action are input into an action recommendation information generation network to obtain the action recommendation information corresponding to the adjusted recommended action. Among them, the action recommendation information generation network is trained using the account status of sample accounts, sample recommended actions, and the actual recommendation information corresponding to the sample recommended actions. The actual recommendation information is determined based on the actual interaction degree between the sample account and the target application and the content recommended by the target application after the target application recommends content matching the sample recommended action to the sample account.
[0142] In step S450, the adjusted recommended action for which the action recommendation information meets the preset condition is determined as the target recommendation operation. The target recommendation operation is used for the target application to recommend content matching the target recommendation action to the current account.
[0143] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a content recommendation method described above, and will not be elaborated here.
[0144] It should be understood that although Figure 2 and Figure 4 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 and Figure 4 at least a part of the steps in
[0145] For the convenience of those skilled in the art, Figure 5 an exemplary model framework diagram of a content recommendation method is shown, as Figure 5As shown in the figure, it includes: a conditional variational auto-encoder 510, a residual actor 520, and a state-action value network 530; among them, the conditional variational auto-encoder 510 includes an action encoder 511 (CVAE-Encoder) and an action decoder 512 (CVAE-Decoder); the residual actor 520 includes a high-level state encoder 521, a low-level state encoder 522, and a residual sub-actor 523.
[0146] Reconstruction stage: The recommendation server inputs the account state of the current account into the action encoder in the conditional variational auto-encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account state of the current account; then, the recommendation server inputs at least one target latent variable and the account state of the current account into the action decoder in the conditional variational auto-encoder, so that the action decoder decodes at least one target latent variable to obtain at least one reconstructed recommendation action.
[0147] Prediction stage: The recommendation server can input the session state feature in the account state into the high-level state encoder, so that the high-level state encoder extracts the session state feature corresponding to the user session state, and input the user request state in the account state into the low-level state encoder, so that the low-level state encoder extracts the request state feature corresponding to the user request state. Then, the recommendation server inputs the session state feature, the request state feature, and the reconstructed recommendation action into the residual sub-actor to obtain the action residual corresponding to the reconstructed recommendation action. The action residual corresponding to the reconstructed recommendation action is superimposed on the reconstructed recommendation action to obtain the adjusted recommendation action corresponding to the reconstructed recommendation action.
[0148] Selection stage: The recommendation server can input the account state of the current account and the adjusted recommendation action into the state-action value network to obtain the action recommendation value corresponding to the adjusted recommendation action; the recommendation server can select the target recommendation operation with the highest action recommendation value among the adjusted recommendation actions based on the action recommendation values corresponding to the adjusted recommendation actions, so as to instruct the target application to recommend content matching the target recommendation action to the current account.
[0149] The content recommendation method adopting the above model framework obtains the account status of the current account and inputs the account status of the current account into the recommendation action reconstruction network to obtain at least one reconstructed recommendation action. The recommendation action reconstruction network is trained with the account status of the sample account and the sample recommendation actions generated by the preset recommendation strategy in response to the account status of the sample account. And according to the account status features corresponding to the account status of the current account, the action adjustment methods corresponding to each reconstructed recommendation action are determined, and each reconstructed recommendation action is adjusted by using the action adjustment method. In this way, through the recommendation action reconstruction network, the video recommendation actions that the online deployment strategy would output in response to the account status of the current account are reconstructed as much as possible, so as to limit the recommendation strategy space within the range where the online deployment strategy is located.
[0150] Meanwhile, by using the account status features corresponding to the account status of the current account, the action adjustment methods corresponding to each reconstructed recommendation action are determined, and each reconstructed recommendation action is adjusted by using the action adjustment method to obtain the adjusted recommendation actions. The action recommendation values corresponding to each adjusted recommendation action are obtained. The action recommendation value is used to represent a long-term indicator of the predicted interaction frequency between the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account. Finally, by determining the adjusted recommendation action with the highest action recommendation value as the target recommendation operation to instruct the target application to recommend content matching the target recommendation action to the current account, the long-term indicator optimization of the video recommendation actions output within the range of the online deployment strategy can be realized, and further the deviation between the learned strategy and the current deployment strategy can be easily controlled, so that when optimizing the long-term indicator of the content recommendation strategy, there is no need to search and optimize in the huge entire recommendation strategy space, reducing the optimization difficulty and cost, and realizing the efficient optimization of the content recommendation strategy focusing on the long-term indicator. As a result, the optimized content recommendation strategy can often accurately recommend the content required by the user to the user, improving the quantity of content obtained by the user through the target application and the efficiency of obtaining content from the target application.
[0151] Therefore, the present disclosure verifies the content recommendation method adopting the above model framework on a large-scale real dataset containing millions of sessions and tens of millions of requests. The statistical data of the dataset is shown in Table 1. The present disclosure uses the interval duration (return time) between sessions, the number of requests contained in a session (session length), and their combination to measure the algorithm effect.
[0152]
[0153] Table 1: Dataset statistical information
[0154] As shown in Table 2 and Figure 6As shown, compared with existing deep reinforcement learning algorithms (DDPG, TD3), offline deep learning algorithms (TD3_BC, IQL), and common imitation learning algorithms, our algorithm can achieve significant performance improvement in all tasks and realize at least a 10% gain.
[0155]
[0156] Table 2: Performance Comparison
[0157] It can be understood that the same / similar parts among the various embodiments of the above methods in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments. For the relevant parts, refer to the descriptions of other method embodiments.
[0158] Figure 7 It is a block diagram of a content recommendation device shown according to an exemplary embodiment. Referring to Figure 7 , the device includes:
[0159] A reconstruction unit 710, configured to execute obtaining the account status of the current account, inputting the account status of the current account into a recommendation action reconstruction network to obtain at least one reconstructed recommendation action: the recommendation action reconstruction network is trained using the account status of a sample account and the corresponding sample recommendation actions; the sample recommendation actions are content recommendation actions generated by a preset recommendation strategy in response to the user status of the sample account:
[0160] An adjustment unit 720, configured to execute determining action adjustment information corresponding to each of the reconstructed recommendation actions according to the account status features corresponding to the account status of the current account, and using the action adjustment information to adjust the corresponding reconstructed recommendation actions to obtain adjusted recommendation actions;
[0161] A determination unit 730, configured to execute determining action recommendation information corresponding to each of the adjusted recommendation actions: the action recommendation information is used to characterize the predicted interaction degree between the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account;
[0162] A recommendation unit 740, configured to execute determining the adjusted recommendation actions that satisfy a preset condition as target recommendation actions; the target recommendation actions are used for the target application to recommend content matching the target recommendation actions to the current account.
[0163] In an exemplary embodiment, the determining unit 730 is specifically configured to input the account status of the current account and the adjusted recommended action into an action recommendation information generation network to obtain action recommendation information corresponding to the adjusted recommended action; wherein, the action recommendation information generation network is trained by using the account status of the sample account, the sample recommended action, and the actual recommendation information corresponding to the sample recommended action; the actual recommendation information is determined according to the actual interaction degree between the sample account and the target application and the content recommended by the target application after the target application recommends content matching the sample recommended action to the sample account.
[0164] In an exemplary embodiment, the adjusting unit 720 is specifically configured to input the account status of the current account and the reconstructed recommended action into an action residual generation network to obtain an action residual corresponding to the reconstructed recommended action; wherein, the action residual is used to represent the action adjustment information corresponding to the reconstructed recommended action; the step of adjusting the corresponding reconstructed recommended action by using the action adjustment information to obtain an adjusted recommended action includes: superimposing the action residual corresponding to the reconstructed recommended action on the corresponding reconstructed recommended action to obtain adjusted recommended actions corresponding to the respective reconstructed recommended actions.
[0165] In an exemplary embodiment, the action residual generation network includes a state feature extraction module and an action residual generation module. The adjusting unit 720 is specifically configured to input the account status of the current account into the state feature extraction module to obtain an account status feature corresponding to the account status of the current account; and input the account status feature and the reconstructed recommended action into the action residual generation module to obtain an action residual corresponding to the reconstructed recommended action.
[0166] In an exemplary embodiment, the account status includes an account session status and an account request status. The account session status is used to represent the interaction status with the target application, and the account request status is used to represent the interaction status with the content recommended by the target application. The state feature extraction module includes a session status encoder and a request status encoder. The adjusting unit 720 is specifically configured to input the account session status into the session status encoder to enable the session status encoder to extract a session status feature corresponding to the account session status, and input the account request status into the request status encoder to enable the session status encoder to extract a request status feature corresponding to the account request status; and use the session status feature and the request status feature as the account status feature.
[0167] In an exemplary embodiment, the device is configured to perform the following operations: input the account status of the sample account and the sample recommendation action into the recommendation action reconstruction network to be trained, and obtain a predicted recommendation action for the sample account; based on the difference between the predicted recommendation action and the sample recommendation action, train the recommendation action reconstruction network to be trained until a preset training end condition is met, and obtain the trained recommendation action reconstruction network.
[0168] In an exemplary embodiment, the device is configured to perform the following operations: input the account status of the sample account and the predicted recommendation action into the action residual generation network to be trained, and obtain an action residual corresponding to the predicted recommendation action; superimpose the action residual corresponding to the predicted recommendation action on the predicted recommendation action to obtain a superimposed recommendation action corresponding to the predicted recommendation action; input the account status of the sample account and the superimposed recommendation action into the action recommendation information generation network to be trained, and obtain action recommendation information corresponding to the superimposed recommendation action; based on the action recommendation information corresponding to the superimposed recommendation action, train the action residual generation network to be trained and the action recommendation information generation network to be trained by using a reinforcement learning method until a preset training end condition is met, and obtain the trained action residual generation network and the trained action recommendation information generation network; wherein, the action residual generation network is a policy function in the reinforcement learning method, the action recommendation information generation network is a value function in the reinforcement learning method, and the action recommendation information corresponding to the superimposed recommendation action is reward data for the policy function in the reinforcement learning method.
[0169] In an exemplary embodiment, the action residual generation network includes a state feature extraction module, and the state feature extraction module is configured to output a session state feature corresponding to the account session state in the account status, and the account session state is used to characterize the interaction state with the target application; the device is configured to perform the following operation during the process of updating the parameters of the state feature extraction module by using the reinforcement learning method: use a first regularization term and a second regularization term to control the parameter update process of the state feature extraction module; wherein, the first regularization term is used to maximize the mutual information between the session state feature and the reward data, and the second regularization term is used to minimize the entropy value of the session state feature distribution.
[0170] In an exemplary embodiment, the recommended action reconstruction network includes an action encoder and an action decoder. The device is configured to input the account status of the sample account and the sample recommended action into the action encoder, so that the action encoder projects the vector corresponding to the sample recommended action into a latent variable space in response to the account status of the sample account, and samples at least one predicted latent variable in the latent variable space; input the at least one predicted latent variable and the account status of the sample account into the action decoder, so that the action decoder decodes the at least one predicted latent variable to obtain a predicted recommended action for the sample account.
[0171] In an exemplary embodiment, the reconstruction unit 710 is specifically configured to input the account status of the current account into the action encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account status of the current account; input the at least one target latent variable and the account status of the current account into the action decoder, so that the action decoder decodes the at least one target latent variable to obtain the at least one reconstructed recommended action.
[0172] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0173] Figure 8 is a block diagram of an electronic device 800 for performing a content recommendation method shown according to an exemplary embodiment. For example, the electronic device 800 may be a server. Referring to Figure 8 , the electronic device 800 includes a processing component 820, which further includes one or more processors, and memory resources represented by a memory 822 for storing instructions executable by the processing component 820, such as application programs. The application programs stored in the memory 822 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 820 is configured to execute instructions to perform the above method.
[0174] The electronic device 800 may further include: a power component 824 configured to perform power management of the electronic device 800, a wired or wireless network interface 826 configured to connect the electronic device 800 to a network, and an input / output (I / O) interface 828. The electronic device 800 may operate based on an operating system stored in the memory 822, such as Windows Server, Mac OSX, Unix, Linux, FreeBSD, or the like.
[0175] In an exemplary embodiment, a computer-readable storage medium including instructions is further provided, such as a memory 822 including instructions, and the above instructions can be executed by a processor of the electronic device 800 to complete the above method. The storage medium may be a computer-readable storage medium. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0176] In an exemplary embodiment, a computer program product is further provided, and the computer program product includes instructions, and the above instructions can be executed by a processor of the electronic device 800 to complete the above method.
[0177] It should be noted that the above-mentioned device, electronic device, computer-readable storage medium, computer program product, etc. may also include other implementation manners according to the description of the method embodiment. The specific implementation manners may refer to the description of the relevant method embodiment, and will not be elaborated herein one by one.
[0178] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0179] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A content recommendation method, characterized in that, it includes: Obtain the account status of the current account, and input the account status of the current account into a recommendation action reconstruction network to obtain at least one reconstructed recommendation action: the recommendation action reconstruction network is trained using the account status of sample accounts and the corresponding sample recommendation actions; the sample recommendation actions are content recommendation actions generated by a preset recommendation strategy in response to the account status of the sample accounts; the reconstructed recommendation actions are characterized by action vectors for recommendation: Input the account status of the current account and the reconstructed recommendation action into an action residual generation network to obtain the action residual corresponding to the reconstructed recommendation action, and superimpose the action residual corresponding to the reconstructed recommendation action on the corresponding reconstructed recommendation action to obtain the adjusted recommendation actions corresponding to each of the reconstructed recommendation actions; the action residual is used to characterize the action adjustment information corresponding to the reconstructed recommendation action; Determine the action recommendation information corresponding to each of the adjusted recommendation actions: the action recommendation information is used to characterize the predicted interaction degree between the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account; Determine the adjusted recommendation actions whose action recommendation information meets the preset conditions as the target recommendation actions; the target recommendation actions are used for the target application to recommend content matching the target recommendation actions to the current account.
2. The content recommendation method according to claim 1, characterized in that, the determination of the action recommendation information corresponding to each of the adjusted recommendation actions includes: Input the account status of the current account and the adjusted recommendation action into an action recommendation information generation network to obtain the action recommendation information corresponding to the adjusted recommendation action; wherein, the action recommendation information generation network is trained using the account status of the sample accounts, the sample recommendation actions, and the actual recommendation information corresponding to the sample recommendation actions; the actual recommendation information is determined according to the actual interaction degree between the sample account, the target application, and the content recommended by the target application after the target application recommends content matching the sample recommendation action to the sample account.
3. The content recommendation method according to claim 1, characterized in that, the action residual generation network includes a state feature extraction module and an action residual generation module, and the input of the account status of the current account and the reconstructed recommendation action into the action residual generation network to obtain the action residual corresponding to the reconstructed recommendation action includes: Input the account status of the current account into the state feature extraction module to obtain the account status feature corresponding to the account status of the current account; Input the account status feature and the reconstructed recommendation action into the action residual generation module to obtain the action residual corresponding to the reconstructed recommendation action.
4. The content recommendation method according to claim 3, characterized in that, The account status includes an account session status and an account request status. The account session status is used to represent the interaction status with the target application, and the account request status is used to represent the interaction status with the content recommended by the target application. The status feature extraction module includes a session status encoder and a request status encoder. Inputting the account status of the current account into the status feature extraction module to obtain the account status feature corresponding to the account status of the current account includes: Inputting the account session status into the session status encoder so that the session status encoder extracts the session status feature corresponding to the account session status, and inputting the account request status into the request status encoder so that the session status encoder extracts the request status feature corresponding to the account request status; Using the session status feature and the request status feature as the account status feature.
5. The content recommendation method according to claim 1, wherein, the method further includes: Inputting the account status of the sample account and the sample recommendation action into the to-be-trained recommendation action reconstruction network to obtain a predicted recommendation action for the sample account; Training the to-be-trained recommendation action reconstruction network based on the difference between the predicted recommendation action and the sample recommendation action until a preset training end condition is satisfied to obtain the trained recommendation action reconstruction network.
6. The content recommendation method according to claim 5, wherein, the method further includes: Inputting the account status of the sample account and the predicted recommendation action into the to-be-trained action residual generation network to obtain an action residual corresponding to the predicted recommendation action; Superimposing the action residual corresponding to the predicted recommendation action on the predicted recommendation action to obtain a superimposed recommendation action corresponding to the predicted recommendation action; Inputting the account status of the sample account and the superimposed recommendation action into the to-be-trained action recommendation information generation network to obtain action recommendation information corresponding to the superimposed recommendation action; Training the to-be-trained action residual generation network and the to-be-trained action recommendation information generation network by using a reinforcement learning method based on the action recommendation information corresponding to the superimposed recommendation action until a preset training end condition is satisfied to obtain the trained action residual generation network and the trained action recommendation information generation network; wherein, the action residual generation network is a policy function in the reinforcement learning method, the action recommendation information generation network is a value function in the reinforcement learning method, and the action recommendation information corresponding to the superimposed recommendation action is reward data for the policy function in the reinforcement learning method.
7. The content recommendation method according to claim 6, wherein, The action residual generation network includes a state feature extraction module, which is used to output the session state features corresponding to the account session state in the account state, and the account session state is used to characterize the interaction state with the target application; the method for training the action residual generation network to be trained and the action recommendation information generation network to be trained by using the reinforcement learning method based on the action recommendation information corresponding to the superimposed recommendation actions includes: In the process of updating the parameters of the state feature extraction module by using the reinforcement learning method, a first regularization term and a second regularization term are used to control the parameter update process of the state feature extraction module; Wherein, the first regularization term is used to maximize the mutual information between the session state features and the reward data, and the second regularization term is used to minimize the entropy value of the session state feature distribution.
8. The content recommendation method according to claim 5, characterized in that The recommended action reconstruction network includes an action encoder and an action decoder. The method for inputting the account state of the sample account and the sample recommended action into the recommended action reconstruction network to be trained to obtain the predicted recommended action for the sample account includes: Inputting the account state of the sample account and the sample recommended action into the action encoder, so that the action encoder projects the vector corresponding to the sample recommended action into a latent variable space in response to the account state of the sample account, and samples at least one predicted latent variable in the latent variable space; Inputting the at least one predicted latent variable and the account state of the sample account into the action decoder, so that the action decoder decodes the at least one predicted latent variable to obtain the predicted recommended action for the sample account.
9. The content recommendation method according to claim 8, characterized in that The method for inputting the account state of the current account into the recommended action reconstruction network to obtain at least one reconstructed recommended action includes: Inputting the account state of the current account into the action encoder, so that the action encoder samples at least one target latent variable in the latent variable space in response to the account state of the current account; Inputting the at least one target latent variable and the account state of the current account into the action decoder, so that the action decoder decodes the at least one target latent variable to obtain the at least one reconstructed recommended action.
10. A content recommendation device, characterized in that it includes: A reconstruction unit, configured to execute obtaining the account state of the current account, inputting the account state of the current account into the recommended action reconstruction network to obtain at least one reconstructed recommended action: the recommended action reconstruction network is trained by using the account state of the sample account and the corresponding sample recommended action; the sample recommended action is a content recommendation action generated by a preset recommendation strategy in response to the user state of the sample account; the reconstructed recommended action is characterized by an action vector for recommendation: An adjustment unit, configured to execute inputting the account status of the current account and the reconstructed recommendation action into an action residual generation network to obtain an action residual corresponding to the reconstructed recommendation action, and superimposing the action residual corresponding to the reconstructed recommendation action on the corresponding reconstructed recommendation action to obtain adjusted recommendation actions corresponding to the respective reconstructed recommendation actions; the action residual is used to represent action adjustment information corresponding to the reconstructed recommendation action. A determination unit, configured to execute determining action recommendation information corresponding to the respective adjusted recommendation actions: the action recommendation information is used to represent a predicted interaction degree among the current account, the target application, and the content recommended by the target application after the target application recommends content matching the adjusted recommendation action to the current account. A recommendation unit, configured to execute determining an adjusted recommendation action for which the action recommendation information meets a preset condition as a target recommendation action. The target recommendation action is used for the target application to recommend content matching the target recommendation action to the current account.
11. An electronic device Characterized in that it includes: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the content recommendation method according to any one of claims 1 to 9.
12. A computer-readable storage medium Characterized in that when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the content recommendation method according to any one of claims 1 to 9.
13. A computer program product, the computer program product includes instructions Characterized in that when the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the content recommendation method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Recommendation method, feature generation network training method and device, and electronic equipment
CN114154050A
Content recommendation model training method, content recommendation method and device
CN114386507A