Recommendation model training method and related device

Through inverse reinforcement learning, the satisfaction prediction model is trained, and the training objectives of the recommendation model are optimized, which solves the problem that users are not interested in the recommended content in the existing recommendation system, and improves the recommendation accuracy and targetedness.

CN119939238APending Publication Date: 2025-05-06HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311464415.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

After the existing recommendation system continues to recommend content related to user interests, it is easy for users to be uninterested in the recommended content, resulting in low recommendation accuracy.

Method used

Through inverse reinforcement learning, the satisfaction prediction model is trained, and the user's satisfaction changes in the interaction process are captured, and the training objectives of the recommendation model are optimized to improve the degree of matching between the feedback actions output by the recommended model and user satisfaction.

Benefits of technology

It improves the recommendation accuracy of the recommendation model, reduces the user's disinterest in the recommended content, and enhances the targetedness of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939238A_ABST
    Figure CN119939238A_ABST
Patent Text Reader

Abstract

A recommendation model training method is applied to content recommendation in the technical field of artificial intelligence. According to the method, based on the interaction process of the user and the recommendation system, a satisfaction prediction model is trained through inverse reinforcement learning, that is, the satisfaction change condition of the user in the interaction process of the user and the recommendation system is learned through the inverse reinforcement learning, and then the satisfaction of the user in various interaction states can be captured through the satisfaction prediction model. Therefore, in the process of training the recommendation model, the user satisfaction corresponding to each piece of training data is predicted through the satisfaction prediction model, so that the recommendation model can learn and recommend contents with high user satisfaction as much as possible when the recommendation model is trained, and the recommendation accuracy of the recommendation model obtained through training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a training method for a recommendation model and related devices. Background Art

[0002] With the rapid development of the Internet, the Internet has become an important way to deliver content, and more and more users are using the Internet to obtain content such as videos, news or product information. At present, recommendation systems are used on the Internet to recommend personalized content to users.

[0003] The current recommendation system infers the user's interests based on the user's historical behavior in accessing content, and recommends content based on the user's interests, thereby improving the relevance of the recommended content to the user, and further achieving targeted content recommendations.

[0004] Specifically, current recommendation systems usually predict users' interests based on the characteristics of the content they have visited and the characteristics of the users themselves, and then filter out the content that is most relevant to the users' interests from a large amount of content, and recommend the filtered content to the users. However, existing recommendation systems will continuously recommend content related to users' interests, which may lead to users gradually becoming uninterested in the recommended content, resulting in a low recommendation accuracy rate for the recommendation system. Summary of the invention

[0005] The present invention provides a method for training a recommendation model, which can improve the recommendation accuracy of the trained recommendation model.

[0006] The first aspect of the present application provides a training method for a recommendation model, which is applied to content recommendation in the field of AI. The method includes: first obtaining a training set, which includes multiple training data. The multiple training data all include the interaction history between the user and the recommendation system, the recommended content, and the feedback action for the recommended content. Among them, the interaction history between the user and the recommendation system includes the content recommended to the user by the recommendation system in the previous recommendation process and the feedback action taken by the user for recommending the previously recommended content.

[0007] Then, based on the training set, the satisfaction prediction model is obtained through inverse reinforcement learning training. The input of the satisfaction prediction model is the interaction history, recommended content and feedback action, and the satisfaction prediction model is used to predict user satisfaction. That is, the interaction history in the training set can actually be regarded as the excellent behavior adopted by the observed users (that is, how to make the decision strategy with the greatest satisfaction), so the unknown reward function can be inferred through inverse reinforcement learning based on these interaction histories, thereby capturing the user's satisfaction in various interaction states.

[0008] Secondly, each training data in the training set is input into the satisfaction prediction model to predict the user satisfaction corresponding to each training data in the training set.

[0009] Finally, the recommendation model is trained based on the training set and the user satisfaction corresponding to the training data in the training set to obtain the trained recommendation model. The training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data. Specifically, when the user satisfaction corresponding to the training data is a high value, if the feedback action output by the recommendation model is an action belonging to positive feedback (such as a click action), then the greater the probability of the feedback action output by the recommendation model, the higher the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

[0010] In this way, by setting the training goals of the recommendation model during the training process, including improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the recommendation model can learn during the training process to recommend content with high user satisfaction as much as possible, and not recommend content with low user satisfaction.

[0011] In this solution, based on the interaction process between the user and the recommendation system, the satisfaction prediction model is trained through inverse reinforcement learning, that is, the changes in the user's satisfaction during the interaction with the recommendation system are learned through inverse reinforcement learning, and then the satisfaction of the user in various interaction states can be captured through the satisfaction prediction model. In this way, in the process of training the recommendation model, the user satisfaction corresponding to each training data is first predicted through the satisfaction prediction model, so that when training the recommendation model, the recommendation model can learn to recommend content with high user satisfaction as much as possible, thereby improving the recommendation accuracy of the trained recommendation model.

[0012] In a possible implementation, a training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

[0013] That is to say, the recommendation model actually has two training goals during the training process. One goal is to learn as much as possible the behavioral strategy of how to adopt feedback actions represented in the training data, so that the feedback actions output by the recommendation model are as close as possible to the feedback actions decided by the user. The other goal is to learn to recommend content with high user satisfaction as much as possible, that is, for content with high user satisfaction, the probability of outputting positive feedback actions is as high as possible, and for content with low user satisfaction, the probability of outputting positive feedback actions is as low as possible.

[0014] In one possible implementation, a recommendation model is trained based on a training set and user satisfaction corresponding to the training data in the training set, specifically including: firstly, inputting the interaction history and recommendation content in the target training data into the recommendation model, and constructing a first loss function based on the feedback action output by the recommendation model and the feedback action in the target training data, wherein the target training data is any training data in the training set, and the first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data. In simple terms, the feedback action output by the recommendation model can be represented as a vector represented by multiple probability values, and the feedback action in the target training data can also be represented as a vector represented by multiple probability values, so the first loss function can actually be the difference between the vector corresponding to the feedback action output by the recommendation model and the vector corresponding to the feedback action in the target training data.

[0015] Then, based on the feedback actions output by the recommendation model and the user satisfaction corresponding to the target training data, a second loss function is constructed, and the second loss function is obtained based on the product sum between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions. The value of the second loss function may have a negative correlation with the product sum between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions, that is, the larger the product sum between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions, the smaller the value of the second loss function.

[0016] Finally, based on the first loss function and the second loss function, the recommendation model is updated. For example, a weighted sum is performed on the first loss function and the second loss function to obtain a total loss function; and based on the total loss function, the recommendation model is updated by a gradient descent method.

[0017] In this scheme, by setting two loss functions to correspond to the two optimization objectives of the recommendation model, the recommendation model can learn to reduce the difference between the feedback action output by the recommendation model and the feedback action in the training data as much as possible during the training process, and improve the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, ensuring that the recommendation model can recommend content with high user click-through rate and high user satisfaction as much as possible, thereby improving the recommendation accuracy of the recommendation model.

[0018] In a possible implementation, the probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

[0019] In one possible implementation, the training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. Adopting different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

[0020] That is to say, during the training process of the satisfaction prediction model, for input data with the same state and different feedback actions, the satisfaction prediction model needs to learn to determine as different user satisfaction as possible based on the input data, so as to ensure that the user satisfaction corresponding to different feedback actions under the same state has sufficient differentiation.

[0021] In one possible implementation, the satisfaction prediction model includes a trunk network and multiple branch networks, and the multiple branch networks are connected to the trunk network. The input of the trunk network is the interaction history and recommended content, and the multiple branch networks are used to predict user satisfaction under different feedback actions.

[0022] In this solution, by setting the satisfaction prediction model to be a network architecture connecting a backbone network and multiple branch networks, the satisfaction prediction model can output different user satisfaction under different feedback actions for the same state, thereby improving the efficiency of the satisfaction prediction model in outputting user satisfaction.

[0023] In a possible implementation, the trained recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

[0024] In a possible implementation, the interaction history includes one or more contents recommended by the recommendation system and the user's feedback actions on the one or more contents.

[0025] The second aspect of the present application provides a content recommendation method, including: obtaining first content and a first interaction history, the first interaction history being used to indicate the historical interaction behavior between the recommendation system and the target user; inputting the first content and the interaction history into a recommendation model to obtain the target user's feedback action on the first content predicted by the recommendation model; wherein the recommendation model is trained based on a training set and user satisfaction corresponding to the training data in the training set, and the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the user satisfaction corresponding to the training data in the training set is predicted based on a satisfaction prediction model, and the satisfaction prediction model is trained based on the training set through inverse reinforcement learning.

[0026] In a possible implementation, a training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

[0027] In one possible implementation, the recommendation model is updated based on a first loss function and a second loss function. The first loss function is constructed based on the feedback action output by the recommendation model and the feedback action in the target training data after the interaction history and recommended content in the target training data are input into the recommendation model. The target training data is any training data in the training set. The first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data. The second loss function is constructed based on the feedback action output by the recommendation model and the user satisfaction corresponding to the target training data. The second loss function is obtained by multiplying the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions based on the product.

[0028] In a possible implementation, the probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

[0029] In one possible implementation, the training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. Adopting different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

[0030] In one possible implementation, the satisfaction prediction model includes a trunk network and multiple branch networks, and the multiple branch networks are connected to the trunk network. The input of the trunk network is the interaction history and recommended content, and the multiple branch networks are used to predict user satisfaction under different feedback actions.

[0031] In a possible implementation, the recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

[0032] In a possible implementation, the interaction history includes one or more contents recommended by the recommendation system and the user's feedback actions on the one or more contents.

[0033] The third aspect of the present application provides a training device for a recommendation model, including: an acquisition module, used to acquire a training set, the training set includes multiple training data, and the multiple training data all include the interaction history between the user and the recommendation system, the recommended content, and the feedback action for the recommended content; a processing module, used to obtain a satisfaction prediction model through inverse reinforcement learning training based on the training set, the input of the satisfaction prediction model is the interaction history, the recommended content, and the feedback action, and the satisfaction prediction model is used to predict user satisfaction; the processing module is also used to input the training set into the satisfaction prediction model to predict the user satisfaction corresponding to the training data in the training set; the processing module is also used to train the recommendation model based on the training set and the user satisfaction corresponding to the training data in the training set to obtain a trained recommendation model; wherein the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

[0034] In a possible implementation, a training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

[0035] In one possible implementation, the processing module is specifically used to: input the interaction history and recommended content in the target training data into the recommendation model, and construct a first loss function based on the feedback action output by the recommendation model and the feedback action in the target training data, the target training data is any training data in the training set, and the first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data; construct a second loss function based on the feedback action output by the recommendation model and the user satisfaction corresponding to the target training data, the second loss function is obtained by the product sum of the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions; based on the first loss function and the second loss function, update the recommendation model.

[0036] In a possible implementation, the probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

[0037] In a possible implementation, the processing module is specifically used to: perform weighted summation on the first loss function and the second loss function to obtain a total loss function; and based on the total loss function, update the recommendation model by a gradient descent method.

[0038] In one possible implementation, the training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. Adopting different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

[0039] In one possible implementation, the satisfaction prediction model includes a trunk network and multiple branch networks, and the multiple branch networks are connected to the trunk network. The input of the trunk network is the interaction history and recommended content, and the multiple branch networks are used to predict user satisfaction under different feedback actions.

[0040] In a possible implementation, the trained recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

[0041] In a possible implementation, the interaction history includes one or more contents recommended by the recommendation system and the user's feedback actions on the one or more contents.

[0042] In a fourth aspect, the present application provides a content recommendation device, comprising: an acquisition module, used to acquire first content and a first interaction history, the first interaction history being used to indicate the historical interaction behavior between the recommendation system and the target user; a recommendation module, used to input the first content and the interaction history into a recommendation model, and obtain the feedback action of the target user with respect to the first content predicted by the recommendation model; wherein the recommendation model is trained based on a training set and user satisfaction corresponding to the training data in the training set, and a training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the user satisfaction corresponding to the training data in the training set is predicted based on a satisfaction prediction model, and the satisfaction prediction model is trained based on the training set through inverse reinforcement learning.

[0043] In a possible implementation, a training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

[0044] In one possible implementation, the recommendation model is updated based on a first loss function and a second loss function. The first loss function is constructed based on the feedback action output by the recommendation model and the feedback action in the target training data after the interaction history and recommended content in the target training data are input into the recommendation model. The target training data is any training data in the training set. The first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data. The second loss function is constructed based on the feedback action output by the recommendation model and the user satisfaction corresponding to the target training data. The second loss function is obtained by multiplying the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions based on the product.

[0045] In a possible implementation, the probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

[0046] In one possible implementation, the training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. Adopting different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

[0047] In one possible implementation, the satisfaction prediction model includes a trunk network and multiple branch networks, and the multiple branch networks are connected to the trunk network. The input of the trunk network is the interaction history and recommended content, and the multiple branch networks are used to predict user satisfaction under different feedback actions.

[0048] In a possible implementation, the recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

[0049] In a possible implementation, the interaction history includes one or more contents recommended by the recommendation system and the user's feedback actions on the one or more contents.

[0050] In a fifth aspect, the present application provides a training device for a recommendation model, which may include a processor, the processor and a memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation of the first aspect is implemented. For the steps in each possible implementation of the first aspect executed by the processor, the details can be referred to the first aspect, and no further description is given here.

[0051] In a sixth aspect, the present application provides a content recommendation device, which may include a processor, the processor and a memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method of the second aspect or any implementation of the second aspect is implemented. For the steps in each possible implementation of the second aspect executed by the processor, the details can be referred to the second aspect, which will not be repeated here.

[0052] The seventh aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes a method implemented in any one of the first or second aspects.

[0053] An eighth aspect of the present application provides a circuit system, the circuit system includes a processing circuit, and the processing circuit is configured to execute a method implemented in any one of the first or second aspects above.

[0054] The ninth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute a method implemented in any one of the first or second aspects.

[0055] The tenth aspect of the present application provides a chip system, which includes a processor for supporting an electronic device to implement the functions involved in any implementation of the first aspect or the second aspect, for example, processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the electronic device. The chip system can be composed of a chip, or it can include a chip and other discrete devices.

[0056] The beneficial effects of the second to tenth aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A schematic diagram of content recommendation for an application store provided in an embodiment of the present application;

[0058] Figure 2 A schematic diagram of a recommendation system provided in an embodiment of the present application;

[0059] Figure 3 A schematic diagram of a system architecture 300 provided in an embodiment of the present application;

[0060] Figure 4 A flowchart of a training method for a recommendation model provided in an embodiment of the present application;

[0061] Figure 5 A schematic diagram of the structure of a satisfaction prediction model provided in an embodiment of the present application;

[0062] Figure 6 A schematic diagram of a model training framework provided in an embodiment of the present application;

[0063] Figure 7 A schematic diagram of the structure of a training device for a recommendation model provided in an embodiment of the present application;

[0064] Figure 8 A schematic diagram of the structure of a content recommendation device provided in an embodiment of the present application;

[0065] Fig. 9 A schematic diagram of the structure of an execution device provided in an embodiment of the present application;

[0066] Fig.10 A schematic diagram of the structure of a chip provided in an embodiment of the present application;

[0067] Fig.11 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those of ordinary skill in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0069] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchanged where appropriate, so that the embodiments can be implemented in a sequence other than that illustrated or described in the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The naming or numbering of the steps that appear in the present application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The process steps that have been named or numbered can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. There may be other division methods when it is implemented in actual applications. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. In addition, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed in multiple circuit units, and some or all of the units may be selected according to actual needs to achieve the purpose of the present application.

[0070] To facilitate understanding, some technical terms involved in the embodiments of the present application are first introduced below.

[0071] (1) Content Recommendation Model

[0072] The content recommendation model is a neural network that uses machine learning algorithms to analyze and learn based on the user's historical access content and the user's own characteristics, determine the probability of the user accessing specific content (i.e., click-through rate), and then recommend content that the user is likely to be interested in based on the probability of the user accessing various content.

[0073] (2) Click-through rate

[0074] Click-through rate refers to the probability that a user clicks on a certain displayed content in a specific environment.

[0075] (3) Neural Network

[0076] A neural network may be composed of neural units, and a neural unit may refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input, and the output of the operation unit may be:

[0077]

[0078] Where s=1, 2, ...n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolution layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the characteristics of the local receptive field. The local receptive field can be an area composed of several neural units.

[0079] (4) Deep Neural Network (DNN)

[0080] Deep neural network, also known as multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. From the position of different layers of DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since DNN has many layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the kth neuron in the L-1th layer to the jth neuron in the Lth layer are defined as It should be noted that the input layer does not have a W parameter. In a deep neural network, more hidden layers allow the network to better describe complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix formed by many layers of vector W).

[0081] (5) Loss Function

[0082] In the process of training a neural network, because we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the neural network). For example, if the predicted value of the network is high, adjust the weight vector to make it predict a lower value, and continue to adjust until the neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function, which are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the neural network becomes a process of minimizing this loss as much as possible.

[0083] (6) Back propagation algorithm

[0084] The neural network can use the error back propagation (BP) algorithm to correct the size of the parameters in the initial prediction model during the training process, so that the error loss of the prediction model becomes smaller and smaller. Specifically, the forward transmission of the input signal to the output will generate error loss, and the error loss information is back-propagated to update the parameters in the initial prediction model, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, aiming to obtain the optimal parameters of the prediction model, such as the weight matrix.

[0085] Specifically, during the model training process, the back propagation algorithm is usually used to calculate the gradient of each node in the model, so as to adjust the weight parameters of the node based on the gradient of each node, thereby reducing the loss function value of the model as much as possible. Among them, the gradient represents the rate of change of a function at a certain point. In addition, the gradient of each node in the model can be determined by taking partial derivatives.

[0086] (7) Gradient descent

[0087] Gradient descent is a first-order optimization algorithm that is often used in machine learning to recursively approximate the minimum deviation prediction model. To use gradient descent to find the local minimum of a function, it is necessary to iteratively search for points at a specified step distance in the opposite direction of the gradient (or approximate gradient) corresponding to the current point on the function. Gradient descent is one of the most commonly used methods for solving the prediction model parameters of machine learning algorithms, that is, unconstrained optimization problems.

[0088] Specifically, when solving the minimum value of the loss function, we can use the gradient descent method to iterate step by step to obtain the minimized loss function and prediction model parameter values. Conversely, if we need to solve the maximum value of the loss function, we need to use the gradient ascent method to iterate.

[0089] (8) Inverse reinforcement learning

[0090] Inverse reinforcement learning is a method to infer an optimized reward function from observed behavior. Inverse reinforcement learning is a type of reinforcement learning, and the difference between inverse reinforcement learning and traditional reinforcement learning is that reinforcement learning attempts to find an excellent strategy under a given reward function, while inverse reinforcement learning attempts to infer an unknown reward function from observed excellent behavior.

[0091] (9) Markov process

[0092] A Markov process is a random process that satisfies the Markov property (i.e., memorylessness). In simpler terms, a Markov process is a process whose future outcomes can be predicted based solely on its current state. In other words, conditional on the current state of the system, its future and past states are independent.

[0093] (10) Regularization

[0094] Regularization is a technique that improves the performance of the model on the test set by modifying the learning algorithm, such as adding additional constraints and penalties to the original constraint function, in order to reduce the generalization error and improve the generalization ability of the model.

[0095] The applicant has found that current recommendation systems usually predict user interests based on the characteristics of the content that the user has visited and the characteristics of the user himself, and then filter out the content that is most relevant to the user's interests from a large amount of content, and recommend the filtered content to the user. In short, after predicting the user's interests based on interaction with the user, the existing recommendation system will continue to recommend content related to the user's interests to the user, which is prone to the phenomenon that the user gradually loses interest in the recommended content.

[0096] For example, if a user clicks on some news related to emerging topics, the recommendation system will think that the user is interested in news related to emerging topics, and thus recommend a large number of such related news to the user. Understandably, once the user gives positive feedback, it seems right to recommend more related news. But on the other hand, as the user continues to browse news on the same topic, the user may gradually get tired of such news because the user has obtained enough information about this topic and does not need to read related news frequently.

[0097] In general, as the recommendation system interacts with users, users' satisfaction with the content recommended by the recommendation system will change. In the early stage when users are eager to learn about a certain topic, when the recommendation system recommends relevant content under the topic to users, users will get higher satisfaction; when users gradually read more content under the current topic, when the recommendation system recommends relevant content under the topic to users again, users may get lower satisfaction or even feel bored.

[0098] Based on this, the embodiment of the present application provides a training method for a recommendation model, which trains a satisfaction prediction model through inverse reinforcement learning based on the interaction process between the user and the recommendation system, that is, through inverse reinforcement learning, the changes in the user's satisfaction during the interaction with the recommendation system are learned, and then the satisfaction of the user in various interaction states can be captured through the satisfaction prediction model. In this way, in the process of training the recommendation model, the user satisfaction corresponding to each training data is first predicted through the satisfaction prediction model, so that when training the recommendation model, the recommendation model can learn to recommend content with high user satisfaction as much as possible, thereby improving the recommendation accuracy of the trained recommendation model.

[0099] The method provided in the embodiment of the present application can be applied to various content recommendation scenarios, such as the recommendation of goods, applications, videos, news or songs, etc. In addition, in the embodiment of the present application, advertisements can be the carriers of these contents, that is, the information of the content can be displayed through the advertisement page, so the method provided in the embodiment of the present application can also be applied to the recommendation scenario of advertisements.

[0100] In a possible scenario, the method provided in the embodiment of the present application may be applied to a scenario where application software or game software is used as recommended content on the interface of an application store. Figure 1 , Figure 1 A schematic diagram of content recommendation for an application store provided in an embodiment of the present application. Figure 1 As shown, for a certain application software on the user's mobile phone, the application store is used by the user to download various application software or game software. In the interface of the application store, various software recommended by the recommendation system are displayed, that is, various application software displayed under "high-quality applications" on the interface.

[0101] In another possible scenario, the method provided in the embodiment of the present application can also be applied to a scenario where a product or application is recommended as content on application software such as social software and video software. For example, when a user is watching a video through a video software, during the period when the user pauses watching the video, the video software can play a corresponding advertisement to recommend a certain product or application.

[0102] In another possible scenario, the method provided by the embodiment of the present application can also be used in a scenario where products are recommended on an online shopping software or web page. For example, when a user searches or browses various products on an online shopping software, the online shopping software can pin some products that the user is likely to be interested in to the top, so as to give priority to recommending these top products to the user.

[0103] In general, the embodiments of the present application do not limit the specific type of recommended content.

[0104] The system scenario of the embodiment of the present application is an application scenario based on machine learning. The following will take the click-through rate prediction scenario in the recommendation system as an example to introduce, where the click-through rate is one of the advertising conversion rates. The click-through rate prediction scenario is a typical scenario in machine learning applications, and its main structure is as follows Figure 2 As shown. Among them, Figure 2 A schematic diagram of a recommendation system provided in an embodiment of the present application. Figure 2 As shown in the figure, the recommendation system includes logs, offline training modules, prediction models, online prediction modules and display lists.

[0105] The basic operating logic of the recommendation system is as follows: users perform a series of actions in the display list on the front end, such as browsing, clicking, commenting, downloading, etc., to generate behavioral data, which is stored in the log. The offline training module in the recommendation system uses data including user behavior logs to perform offline model training, thereby training a prediction model. Then, the prediction model is deployed in the online service environment to form an online prediction module, and the recommendation results are given based on the user's request access, item features and contextual information, and then the recommendation results are displayed in the display list. Finally, the user generates feedback on the recommendation results in the display list to form the user's behavior data, and this behavior data is stored in the log.

[0106] During the entire process of the recommendation system running, thousands of user features and item features can be recorded in the logs. It is impractical to apply all of these features to content recommendations. This is because more features mean more computing resources are needed, and the online latency will increase accordingly, leading to increased costs. In addition, noisy features and redundant features will also cause the results learned by the content recommendation model to deteriorate. Therefore, it is necessary to perform feature screening on all features, and how to more accurately evaluate the importance of features in this process is crucial to whether the accuracy of the content recommendation model can be improved after feature screening.

[0107] See also Figure 3 , Figure 3 A schematic diagram of a system architecture 300 provided in an embodiment of the present application. Figure 3 As shown, in the system architecture 300, the execution device 310 can be implemented by one or more servers. Optionally, the execution device 310 cooperates with other computing devices, such as data storage, routers, load balancers and other devices; the execution device 310 can be arranged on a physical site, or distributed on multiple physical sites. The execution device 310 can use the data in the data storage system 320, or call the program code in the data storage system 320 to implement the feature screening method provided in the embodiment of the present application, and then obtain the features for performing the content recommendation task.

[0108] Users can operate their respective user devices (such as local device 301 and local device 302) to interact with execution device 310. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.

[0109] Each user's local device can interact with the execution device 310 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0110] In one implementation, the execution device 310 is used to implement the feature screening method provided in the embodiment of the present application to obtain features for performing content recommendation tasks. Furthermore, during the process of the local device 301 and the local device 302 accessing content, the execution device 310 predicts the content that the user is likely to be interested in based on the screened features, and then returns the corresponding recommended content to the local device 301 and the local device 302.

[0111] In another implementation, one or more aspects of the execution device 310 may be implemented by each local device. For example, the local device 301 may provide local data or feedback calculation results to the execution device 310, or execute the feature screening method provided in the embodiment of the present application.

[0112] It should be noted that all functions of the execution device 310 can also be implemented by the local device. For example, the local device 301 implements the functions of the execution device 310 and provides services to its own user, or provides services to the user of the local device 302.

[0113] In general, the feature screening method provided in the embodiments of the present application can be applied to electronic devices, such as the above-mentioned execution device 310, local device 301 or local device 302.

[0114] See also Figure 4 , Figure 4 A flowchart of a training method for a recommendation model provided in an embodiment of the present application is shown below. Figure 4 As shown, the training method of the recommendation model includes the following steps 401-404.

[0115] Step 401 : obtaining a training set, wherein the training set includes a plurality of training data, each of which includes the interaction history between the user and the recommendation system, the recommended content, and the feedback action for the recommended content.

[0116] In this embodiment, the training set can be obtained based on the interaction between the deployed recommendation system and the user. Specifically, for each training data included in the training set, each training data includes the interaction history between the user and the recommendation system, the recommended content currently displayed to the user by the recommendation system, and the feedback action taken by the user for the recommended content. Among them, the interaction history between the user and the recommendation system includes the content recommended to the user by the recommendation system in the previous recommendation process and the feedback action taken by the user for the content recommended in the past.

[0117] In addition, the interaction history between the user and the recommendation system may include one or more rounds of interaction between the user and the recommendation system. That is, the interaction history actually includes one or more contents recommended by the recommendation system and the user's feedback actions on the one or more contents, and the one or more contents recommended by the recommendation system included in the interaction history may be recommended in one round of interaction or recommended separately in multiple rounds of interaction. For example, the interaction history between the user and the recommendation system records multiple rounds of interaction between the user and the recommendation system, and each round of interaction includes one or more recommended contents recommended by the recommendation system and the user's feedback actions on the recommended contents.

[0118] Optionally, the type of feedback action that the user can take may be related to the recommendation scenario. For example, the feedback action taken by the user may include click, download or purchase, etc., which is not specifically limited here.

[0119] Step 402: Based on the training set, a satisfaction prediction model is obtained through inverse reinforcement learning training. The input of the satisfaction prediction model is the interaction history, recommended content and feedback action. The satisfaction prediction model is used to predict user satisfaction.

[0120] In interactive recommendation, the user's satisfaction with the recommendation system will change continuously due to the interaction history, but the user often does not directly express his or her satisfaction. Therefore, in this embodiment, inverse reinforcement learning is used to train a satisfaction prediction model for predicting user satisfaction.

[0121] Specifically, for users who interact with the recommendation system, the user can be assumed to be a sequential decision maker, and the user's decision conforms to the Markov process. That is, in each round of interaction with the recommendation system, the user will make decisions based on historical interaction records and the currently recommended content (such as whether to click on the content). After executing each round of decision-making, a satisfaction level will be generated in the user's brain, which will be given to the user as a reward, and the user's goal is to maximize his or her satisfaction through this sequential decision-making.

[0122] In other words, the interaction histories in the training set can be understood as how users make decisions in each interaction state to maximize their satisfaction, that is, the content clicked by users in each round of interaction is actually the content that they are most satisfied with at the moment. In this way, the interaction histories in the training set can actually be regarded as the observed excellent behaviors adopted by users (that is, the decision-making strategies that maximize satisfaction), so the unknown reward function can be inferred based on these interaction histories, thereby capturing the user's satisfaction in various interaction states.

[0123] In the process of training the satisfaction prediction model, the input of the satisfaction prediction model is each data in the training set, namely the interaction history, recommended content and feedback action, and the satisfaction prediction model is used to predict the user satisfaction corresponding to each input data. In simple terms, the satisfaction prediction model predicts the user satisfaction when recommending a certain content to the user under a certain interaction state.

[0124] Specifically, in order to adapt to the inverse reinforcement learning used to train the satisfaction prediction model, the training data can be constructed in the format of <state, action>. Among them, the state includes the interaction history between the user and the recommendation system in the training data (the interaction history includes the recommended content and the user's feedback action) and the recommended content, and the action is the user's feedback action for the recommended content. In this way, after the training data is constructed in the format of <state, action>, the satisfaction prediction model can be trained based on the inverse reinforcement learning method. The specific training process can refer to the existing relevant inverse reinforcement learning implementation method, which will not be repeated here.

[0125] In the recommendation scenario, different feedback actions taken by users for a certain recommended content often represent user satisfaction with great differences. For example, in a certain interactive state, for a certain recommended content, if the user clicks on the recommended content, it means that the user has a high user satisfaction with the recommended content, and if the user does not click on the recommended content, it means that the user has a low user satisfaction with the recommended content. In other words, different feedback actions often represent different user satisfaction, and the difference between different user satisfaction may be large.

[0126] Optionally, in order to ensure that the user satisfaction determined by the satisfaction prediction model when the same state and different feedback actions are input can have sufficient discrimination, during the training process of the satisfaction prediction model, the training goal of the satisfaction prediction model can be set to include increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are used under the same state. Among them, using different feedback actions under the same state means that the interaction history and recommended content input to the satisfaction prediction model are the same, and the feedback actions input to the satisfaction prediction model are different. In other words, during the training process of the satisfaction prediction model, for input data with the same state and different feedback actions, the satisfaction prediction model needs to learn to determine as different user satisfaction as possible based on the input data, so as to ensure that the user satisfaction corresponding to different feedback actions under the same state has sufficient discrimination.

[0127] Optionally, in order to facilitate the output of different user satisfaction under different feedback actions for the same state, the satisfaction prediction model may include a trunk network and multiple branch networks. Multiple branch networks are connected to the trunk network, that is, each branch network is used to receive the output of the trunk network as the input of this network. The input of the trunk network is the interaction history and recommended content, and the multiple branch networks are used to predict the user satisfaction under different feedback actions. In other words, by inputting the interaction history and recommended content into the satisfaction prediction model, the user satisfaction under different feedback actions can be determined based on the output of multiple branch networks. In the case where the input of the satisfaction prediction model also includes a feedback action, the output of the branch network used can be selected based on the feedback action, and then the user satisfaction under the feedback action can be determined.

[0128] For example, see Figure 5 , Figure 5 This is a schematic diagram of a satisfaction prediction model provided in an embodiment of the present application. Figure 5 As shown, the satisfaction prediction model includes a trunk network, branch network 1 and branch network 2, wherein the inputs of branch network 1 and branch network 2 are both the outputs of the trunk network. In the scenario where the feedback action only has a click action and a no-click action, the output of branch network 1 is output data 1, which can be used to determine the user satisfaction when the feedback action is a click; the output of branch network 2 is output data 2, which can be used to determine the user satisfaction when the feedback action is a no-click.

[0129] In this embodiment, the structure of the satisfaction prediction model can be implemented in a variety of ways, as long as it can support the input data as serialized data (i.e., the data in the <state, action> format described above), and the output data can be used to determine user satisfaction, and this embodiment does not specifically limit this. For example, the structure of the satisfaction prediction model is, for example, a deep neural network, a product-based neural network (Product-based neural networks, PNN), a deep interest network (Deep Interest Network, DIN) or a deep interest evolution network (Deep Interest Evolution Network, DIEN).

[0130] Step 403: input the training set into the satisfaction prediction model to predict the user satisfaction corresponding to the training data in the training set.

[0131] After training the satisfaction prediction model, each training data in the training set can be input into the satisfaction prediction model, so as to predict the user satisfaction corresponding to each training data based on the output data of the satisfaction prediction model. In this way, the training data represented in the <state, action> format can be expanded to the <state, action, user satisfaction> format, that is, each training data has a corresponding user satisfaction.

[0132] Step 404: training a recommendation model based on the training set and the user satisfaction corresponding to the training data in the training set to obtain a trained recommendation model.

[0133] Compared with the traditional recommendation model training method, since the user satisfaction corresponding to the training data is newly added to train the recommendation model in this embodiment, the training goal of the recommendation model in the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data. Specifically, the higher the user satisfaction corresponding to the training data, the higher the probability of the feedback action output by the recommendation model being positive feedback, which means the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data is higher; similarly, the lower the user satisfaction corresponding to the training data, the higher the probability of the feedback action output by the recommendation model being positive feedback, which means the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data is higher.

[0134] Generally speaking, the output of the recommendation model is actually usually the probability of adopting various feedback actions. The feedback action with the highest probability is often the most likely feedback action to be adopted. Therefore, the feedback action with the highest probability can also be called the feedback action output by the recommendation model. In addition, for feedback actions, different feedback actions often represent different user satisfaction. For example, feedback actions that are positive feedback (such as click actions) often represent higher user satisfaction, while feedback actions that are negative feedback (such as no click actions) represent lower user satisfaction.

[0135] Therefore, when the user satisfaction corresponding to the training data has been determined, the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data can be determined. Specifically, when the user satisfaction corresponding to the training data is a high value, if the feedback action output by the recommendation model is an action belonging to positive feedback (such as a click action), then the greater the probability of the feedback action output by the recommendation model, the higher the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data; if the feedback action output by the recommendation model is an action belonging to negative feedback (such as a no-click action), then the greater the probability of the feedback action output by the recommendation model, the lower the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

[0136] Similarly, when the user satisfaction corresponding to the training data is a lower value, if the feedback action output by the recommendation model is a negative feedback action, the greater the probability of the feedback action output by the recommendation model, the higher the degree of match between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data; if the feedback action output by the recommendation model is a positive feedback action, the greater the probability of the feedback action output by the recommendation model, the lower the degree of match between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

[0137] In this way, by setting the training goals of the recommendation model during the training process, including improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the recommendation model can learn during the training process to recommend content with high user satisfaction as much as possible, and not recommend content with low user satisfaction.

[0138] In this embodiment, based on the interaction process between the user and the recommendation system, the satisfaction prediction model is trained through inverse reinforcement learning, that is, the changes in the user's satisfaction during the interaction with the recommendation system are learned through inverse reinforcement learning, so that the satisfaction of the user in various interaction states can be captured through the satisfaction prediction model. In this way, in the process of training the recommendation model, the user satisfaction corresponding to each training data is first predicted through the satisfaction prediction model, so that when training the recommendation model, the recommendation model can learn to recommend content with high user satisfaction as much as possible, thereby improving the recommendation accuracy of the trained recommendation model.

[0139] In addition, in addition to improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data. In other words, there are actually two training goals for the recommendation model during the training process. One goal is to learn as much as possible the behavioral strategy of how to adopt feedback actions represented in the training data, so that the feedback action output by the recommendation model is as close as possible to the feedback action decided by the user; the other goal is to learn to recommend content with high user satisfaction as much as possible, that is, for content with high user satisfaction, the probability of outputting positive feedback actions is as high as possible, and for content with low user satisfaction, the probability of outputting positive feedback actions is as low as possible.

[0140] In the above embodiment, the trained recommendation model can be used to perform any one of the following tasks: advertisement recommendation task, product recommendation task, video recommendation task, application recommendation task and news recommendation task. In general, the trained recommendation model can be applied to recommend any type of content, and is not specifically limited here.

[0141] Exemplarily, the process of training the recommendation model based on the training set and the user satisfaction corresponding to the training data in the training set in the above step 404 may specifically include the following steps 4041-4043.

[0142] Step 4041, input the interaction history and recommended content in the target training data into the recommendation model, and construct a first loss function based on the feedback action output by the recommendation model and the feedback action in the target training data. The first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data.

[0143] In this embodiment, the target training data is any training data in the training set. During the training process of the recommendation model, the interaction history between the user and the recommendation system and the recommended content included in the target training data can be input into the recommendation model to obtain the feedback action output by the recommendation model. Specifically, the data output by the recommendation model is actually the probabilities corresponding to various feedback actions, such as the probability of a click operation and the probability of a non-click operation. For example, the data output by the recommendation model may be (0.8, 0.2), where 0.8 represents the probability corresponding to a click operation, and 0.2 represents the probability corresponding to a non-click operation; and the feedback action in the target training data can also be represented in the form of probability. For example, in the case where the feedback action in the target training data is a click operation, the feedback action in the target training data can be represented by (1, 0), where 1 represents the probability corresponding to a click operation, and 0 represents the probability corresponding to a non-click operation.

[0144] In this way, based on the feedback action output by the recommendation model and the feedback action in the target training data, a first loss function can be constructed to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data. In simple terms, the feedback action output by the recommendation model can be represented as a vector represented by multiple probability values, and the feedback action in the target training data can also be represented as a vector represented by multiple probability values. Therefore, the first loss function can actually be the difference between the vector corresponding to the feedback action output by the recommendation model and the vector corresponding to the feedback action in the target training data.

[0145] It should be noted that the first loss function of this embodiment is actually used to achieve the above-mentioned training goal: reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data. Specifically, since the first loss function represents the difference between the vector corresponding to the feedback action output by the recommendation model and the vector corresponding to the feedback action in the target training data, it is possible to reduce the difference between the feedback action output by the recommendation model and the feedback action in the training data by minimizing the first loss function.

[0146] Step 4042, based on the feedback actions output by the recommendation model and the user satisfaction corresponding to the target training data, construct a second loss function, the second loss function being obtained by multiplying the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions.

[0147] Specifically, the value of the second loss function may have a negative correlation with the product sum between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions, that is, the greater the product sum between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions, the smaller the value of the second loss function. In this way, in the training process of the recommendation model, the process of minimizing the value of the second loss function is actually to maximize the product sum between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions. Since the user satisfaction corresponding to various feedback actions is known, and the user satisfaction of positive feedback actions (such as click operations) is high, while the user satisfaction of negative feedback actions (such as non-click operations) is low, if you want to maximize the product sum between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions, it is necessary to maximize the matching degree between the feedback actions output by the recommendation model and the user satisfaction. That is, in the case where the user satisfaction corresponding to the positive feedback action is higher, the probability of the positive feedback action output by the recommendation model needs to be as high as possible.

[0148] In general, the second loss function of this embodiment is actually used to achieve the above-mentioned training goal: to improve the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

[0149] Exemplarily, in some possible content recommendation scenarios (such as advertisement recommendation scenarios), the probabilities of various feedback actions output by the above recommendation model include, for example, the probability of clicking and the probability of not clicking.

[0150] Step 4043: Update the recommendation model based on the first loss function and the second loss function.

[0151] When the first loss function and the second loss function are constructed, a weighted summation can be performed on the first loss function and the second loss function to obtain a total loss function. The weights of the first loss function and the second loss function can be determined or adjusted based on actual conditions, and are not specifically limited here. In this way, based on the total loss function, the recommendation model can be updated by the gradient descent method. That is, based on the total loss function, the gradient of each neural unit in the recommendation model is determined, and then the parameters of each neural unit are updated by gradient descent to achieve the update of the recommendation model, so as to achieve the goal of minimizing the total loss function.

[0152] In this scheme, by setting two loss functions to correspond to the two optimization objectives of the recommendation model, the recommendation model can learn to reduce the difference between the feedback action output by the recommendation model and the feedback action in the training data as much as possible during the training process, and improve the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, ensuring that the recommendation model can recommend content with high user click-through rate and high user satisfaction as much as possible, thereby improving the recommendation accuracy of the recommendation model.

[0153] The above introduces the process of the training method of the recommendation model provided in the embodiment of the present application. For ease of understanding, the following will introduce in detail how to perform the training of the recommendation model with reference to specific examples.

[0154] For example, see Figure 6 , Figure 6 A schematic diagram of a model training framework provided in an embodiment of the present application. Figure 6 As shown, based on the training set, the satisfaction prediction model is trained by constructing an inverse reinforcement learning loss function, so that the satisfaction prediction model can give corresponding user satisfaction for the user behavior strategy. In addition, based on the training set and the user satisfaction output by the satisfaction prediction model, the recommendation model is trained by constructing a cross entropy loss and an alignment loss, respectively, so that the recommendation model can output feedback actions under various interaction states with reference to the user behavior strategy indicated by the training set.

[0155] Specifically, the training of the recommendation model includes the following four steps: 1. Construction of training set; 2. Construction of satisfaction prediction model based on the training set through inverse reinforcement learning; 3. Labeling the training set based on the constructed satisfaction prediction model to obtain the user satisfaction corresponding to each training data; 4. Training the recommendation model based on the labeled training set.

[0156] 1. Construction of training set.

[0157] In order to adapt the inverse reinforcement learning algorithm, the interaction data between the recommendation system and the user can be constructed as a training set. The training set includes multiple training data, and the format of each training data is <state, action>. The state includes the interaction history between the user and the recommendation system in the training data (the interaction history includes the recommended content and the user's feedback action) and the recommended content, and the action is the user's feedback action on the recommended content.

[0158] 2. Based on the training set, a satisfaction prediction model is constructed through inverse reinforcement learning.

[0159] Taking the satisfaction prediction model as a deep neural network as an example, the input of the satisfaction prediction model is the interaction history between the user and the recommendation system and the recommended content in the training data, and the output is the user satisfaction under various feedback actions.

[0160] Specifically, a user state can be represented as s t =(h t-1 ,i t ), where h t-1 is the user's interaction history, i t is the current recommended content. The user's action set can be simplified to a t ={click,not click}, that is, click or not click. After defining the state and action, the Soft-Q value can be further defined by the following formula.

[0161]

[0162] Among them, Q(s,a) represents the expected value of the reward in the long term of inverse reinforcement learning; r(s,a) represents the user satisfaction under different feedback actions; γ represents the manually set hyperparameter; Indicates that in the next state s ′ Expectations under ′ is the next state after executing action a in state s; V(s′) represents the value function of the next state, and V(s ′ )=log∑ a′ expQ(s ′ ,a′); a′ represents the feedback action performed in the state.

[0163] In this embodiment, the satisfaction prediction model is used to represent Q, thereby obtaining a parameterized representation of the Q value Q θ (s,a), that is, Q θ (s,a) represents the output of the satisfaction prediction model.

[0164] In this way, for the satisfaction prediction model, the loss function constructed based on traditional inverse reinforcement learning can be shown as the following formula.

[0165]

[0166] in, represents the inverse reinforcement learning loss function; represents the expectation, which is actually calculated by sampling from the distribution of state s and action a in the training set; ρ E represents the distribution of state s and action a in the training set; α represents a hyperparameter.

[0167] In addition, the action set can be restricted to two actions {click, not click}, in order to make Q θ (s, a) The Q values ​​corresponding to the two actions in the same state s are sufficiently distinguishable, and a regularized loss function called Reward Distinction Enlargement (RDE) is proposed, as shown in the following formula.

[0168]

[0169] in, represents the regularized loss function of RDE; represents the distribution of state s; a P Indicates that the feedback action is click; a N Indicates that the feedback action is not click.

[0170] Therefore, based on the above-mentioned inverse reinforcement learning loss function and the regularization loss function of RDE, the total loss function finally used to train the satisfaction prediction model can be constructed, as shown in the following formula.

[0171]

[0172] in, represents the total loss function used to train the satisfaction prediction model; β represents the weight.

[0173] 3. Label the training set based on the constructed satisfaction prediction model to obtain the user satisfaction corresponding to each training data.

[0174] Since the satisfaction prediction model Q obtained in step 2 above θ (s,a) actually outputs the Q value. In order to obtain the user satisfaction after the user takes feedback action a in state s, the following formula can be used to calculate the user satisfaction.

[0175]

[0176] In this way, for each training data in the training set, after inputting the training data into the satisfaction prediction model, the user satisfaction corresponding to each training data can be obtained based on the above formula. In this way, each training data can actually be expanded from the (s, a) format to the (s, a, r) ​​format, where r represents user satisfaction.

[0177] 4. Train the recommendation model based on the labeled training set.

[0178] In the process of training the recommendation model, a cross entropy loss function (ie, the first loss function of the above embodiment) can be constructed based on the training set, as shown in the following formula.

[0179]

[0180] in, represents the cross entropy loss function; ψ represents the parameter of the recommendation model p; p ψ (a P |s) represents the probability that the feedback action output by the recommendation model p in state s is a click operation; p ψ (a N |s) represents the probability that the feedback action output by the recommendation model p in state s is a no-click operation.

[0181] In addition, in order to enable the recommendation model to learn to recommend content with high user satisfaction as much as possible, an alignment loss function (i.e., the second loss function mentioned above) can be constructed based on the labels in the training set (i.e., the user satisfaction corresponding to the training data), as shown in the following formula.

[0182]

[0183] in, indicates; r θ (s|a) represents the user satisfaction obtained by taking action a in state s; p ψ (a|s) represents the probability of feedback action a output by the recommendation model in state s; Represents a collection of actions.

[0184] Finally, the total loss function for training the recommendation model can be shown as the following formula.

[0185]

[0186] in, represents the total loss function used to train the recommendation model; represents the cross entropy loss function; represents the alignment loss function; Represents weight.

[0187] From the above formula, we can see that the alignment loss maximizes the expectation of user satisfaction r(s,a), which makes the recommendation model not only tend to recommend content with high click-through rate, but also tend to recommend content with higher user satisfaction, thereby considering content recommendation from multiple angles and ensuring the accuracy of content recommendation.

[0188] The above introduces the training method of the recommendation model provided in the embodiment of the present application. The following will introduce the process of implementing content recommendation based on the recommendation model obtained by the training method of the recommendation model mentioned above.

[0189] An embodiment of the present application also provides a content recommendation method, including: obtaining first content and a first interaction history, the first interaction history being used to indicate the historical interaction behavior between the recommendation system and the target user; inputting the first content and the interaction history into a recommendation model to obtain a feedback action of the target user with respect to the first content predicted by the recommendation model; wherein the recommendation model is trained based on a training set and user satisfaction corresponding to the training data in the training set, and a training goal of the recommendation model during the training process includes improving the degree of match between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the user satisfaction corresponding to the training data in the training set is predicted based on a satisfaction prediction model, and the satisfaction prediction model is trained based on the training set through inverse reinforcement learning.

[0190] Specifically, the process of training the recommendation model and the satisfaction prediction model can refer to the above-mentioned embodiment and will not be repeated here.

[0191] The method provided in the embodiment of the present application is described in detail above. Next, the device provided in the embodiment of the present application for executing the above method will be introduced.

[0192] See also Figure 7 , Figure 7 A schematic diagram of a training device for a recommendation model provided in an embodiment of the present application. Figure 7As shown, a training device for a recommendation model includes: an acquisition module 701, which is used to acquire a training set, wherein the training set includes multiple training data, and the multiple training data all include the interaction history between the user and the recommendation system, the recommended content, and the feedback action for the recommended content; a processing module 702, which is used to obtain a satisfaction prediction model through inverse reinforcement learning training based on the training set, wherein the input of the satisfaction prediction model is the interaction history, the recommended content, and the feedback action, and the satisfaction prediction model is used to predict user satisfaction; the processing module 702 is also used to input the training set into the satisfaction prediction model to predict the user satisfaction corresponding to the training data in the training set; the processing module 702 is also used to train the recommendation model based on the training set and the user satisfaction corresponding to the training data in the training set to obtain a trained recommendation model; wherein the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

[0193] In a possible implementation, a training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

[0194] In one possible implementation, the processing module 702 is specifically used to: input the interaction history and recommended content in the target training data into the recommendation model, and construct a first loss function based on the feedback action output by the recommendation model and the feedback action in the target training data, the target training data is any training data in the training set, and the first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data; construct a second loss function based on the feedback action output by the recommendation model and the user satisfaction corresponding to the target training data, the second loss function is obtained by multiplying the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions; based on the first loss function and the second loss function, update the recommendation model.

[0195] In a possible implementation, the probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

[0196] In a possible implementation, the processing module 702 is specifically configured to: perform a weighted summation on the first loss function and the second loss function to obtain a total loss function; and update the recommendation model based on the total loss function by a gradient descent method.

[0197] In one possible implementation, the training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. Adopting different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

[0198] In one possible implementation, the satisfaction prediction model includes a trunk network and multiple branch networks, and the multiple branch networks are connected to the trunk network. The input of the trunk network is the interaction history and recommended content, and the multiple branch networks are used to predict user satisfaction under different feedback actions.

[0199] In a possible implementation, the trained recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

[0200] In a possible implementation, the interaction history includes one or more contents recommended by the recommendation system and the user's feedback actions on the one or more contents.

[0201] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of a content recommendation device provided in an embodiment of the present application. Figure 8 As shown, the content recommendation device includes: an acquisition module 801, which is used to acquire the first content and the first interaction history, and the first interaction history is used to indicate the historical interaction behavior between the recommendation system and the target user; a recommendation module 802, which is used to input the first content and the interaction history into a recommendation model to obtain the feedback action of the target user for the first content predicted by the recommendation model; wherein the recommendation model is trained based on a training set and user satisfaction corresponding to the training data in the training set, and the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, and the user satisfaction corresponding to the training data in the training set is predicted based on a satisfaction prediction model, and the satisfaction prediction model is trained based on the training set through inverse reinforcement learning.

[0202] In a possible implementation, a training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

[0203] In one possible implementation, the recommendation model is updated based on a first loss function and a second loss function. The first loss function is constructed based on the feedback action output by the recommendation model and the feedback action in the target training data after the interaction history and recommended content in the target training data are input into the recommendation model. The target training data is any training data in the training set. The first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data. The second loss function is constructed based on the feedback action output by the recommendation model and the user satisfaction corresponding to the target training data. The second loss function is obtained by multiplying the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions based on the product.

[0204] In a possible implementation, the probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

[0205] In one possible implementation, the training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. Adopting different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

[0206] In one possible implementation, the satisfaction prediction model includes a trunk network and multiple branch networks, and the multiple branch networks are connected to the trunk network. The input of the trunk network is the interaction history and recommended content, and the multiple branch networks are used to predict user satisfaction under different feedback actions.

[0207] In a possible implementation, the recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

[0208] In a possible implementation, the interaction history includes one or more contents recommended by the recommendation system and the user's feedback actions on the one or more contents.

[0209] See also Fig. 9 , Fig. 9 This is a schematic diagram of the structure of an execution device provided in an embodiment of the present application. The execution device 900 can be specifically a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Specifically, the execution device 900 includes: a receiver 901, a transmitter 902, a processor 903 and a memory 904 (wherein the number of processors 903 in the execution device 900 can be one or more, Fig. 9In the example of FIG. 9 , the processor 903 may include an application processor 9031 and a communication processor 9032. In some embodiments of the present application, the receiver 901, the transmitter 902, the processor 903 and the memory 904 may be connected via a bus or other means.

[0210] The memory 904 may include a read-only memory and a random access memory, and provides instructions and data to the processor 903. A portion of the memory 904 may also include a non-volatile random access memory (NVRAM). The memory 904 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0211] The processor 903 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.

[0212] The method disclosed in the above embodiment of the present application can be applied to the processor 903, or implemented by the processor 903. The processor 903 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 903 or an instruction in the form of software. The above processor 903 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0213] The processor 903 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 904, and the processor 903 reads the information in the memory 904 and completes the steps of the above method in combination with its hardware.

[0214] The receiver 901 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the execution device. The transmitter 902 can be used to output digital or character information through the first interface; the transmitter 902 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 902 can also include a display device such as a display screen.

[0215] The electronic device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the method for determining the model structure described in the above embodiment, or so that the chip in the training device executes the method for determining the model structure described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0216] For details, please refer to Fig.10 , Fig.10 A schematic diagram of the structure of a chip provided in an embodiment of the present application, the chip can be expressed as a neural network processor NPU 1000, NPU 1000 is mounted on the host CPU (Host CPU) as a coprocessor, and the Host CPU assigns tasks. The core part of the NPU is the operation circuit 1003, which is controlled by the controller 1004 to extract matrix data from the memory and perform multiplication operations.

[0217] In some implementations, the operation circuit 1003 includes multiple processing units (Process Engine, PE) inside. In some implementations, the operation circuit 1003 is a two-dimensional systolic array. The operation circuit 1003 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1003 is a general-purpose matrix processor.

[0218] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory 1002 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1001 and performs matrix operation with matrix B, and the partial result or final result of the matrix is ​​stored in the accumulator 1008.

[0219] The unified memory 1006 is used to store input data and output data. The weight data is directly transferred to the weight memory 1002 through the direct memory access controller (DMAC) 1005. The input data is also transferred to the unified memory 1006 through the DMAC.

[0220] BIU stands for Bus Interface Unit, i.e., bus interface unit 1010 , which is used for interaction between AXI bus, DMAC and instruction fetch buffer (IFB) 1009 .

[0221] The bus interface unit 1010 (BIU) is used for the instruction fetch memory 1009 to obtain instructions from the external memory, and is also used for the storage unit access controller 1005 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0222] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1006 or to transfer weight data to the weight memory 1002 or to transfer input data to the input memory 1001.

[0223] The vector calculation unit 1007 includes multiple operation processing units, and further processes the output of the operation circuit 1003 when necessary, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, upsampling of feature planes, etc.

[0224] In some implementations, the vector calculation unit 1007 can store the processed output vector to the unified memory 1006. For example, the vector calculation unit 1007 can apply a linear function; or a nonlinear function to the output of the operation circuit 1003, such as linear interpolation of the feature plane extracted by the convolution layer, and then, for example, a vector of accumulated values ​​to generate an activation value. In some implementations, the vector calculation unit 1007 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1003, for example, for use in a subsequent layer in a neural network.

[0225] An instruction fetch buffer 1009 connected to the controller 1004 is used to store instructions used by the controller 1004;

[0226] Unified memory 1006, input memory 1001, weight memory 1002 and instruction fetch memory 1009 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0227] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.

[0228] See also Fig.11 , Fig.11 This is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the above Figure 4 The disclosed methods may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of manufacture.

[0229] Fig.11 Schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0230] In one embodiment, the computer readable storage medium 1100 is provided using a signal bearing medium 1101. The signal bearing medium 1101 may include one or more program instructions 1102, which when executed by one or more processors may provide the above-mentioned Figure 4 Describes the functionality or part of the functionality.

[0231] In some examples, signal bearing medium 1101 may include computer readable medium 1103 such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a ROM or RAM, and the like.

[0232] In some embodiments, the signal bearing medium 1101 may include a computer recordable medium 1104, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, etc. In some embodiments, the signal bearing medium 1101 may include a communication medium 1105, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, etc.). Thus, for example, the signal bearing medium 1101 may be communicated by a wireless form of the communication medium 1105 (e.g., a wireless communication medium that complies with the IEEE 802.X standard or other transmission protocol).

[0233] The one or more program instructions 1102 may be, for example, computer executable instructions or logic implementation instructions. In some examples, the computing device of the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1102 communicated to the computing device via one or more of the computer readable medium 1103, the computer recordable medium 1104, and / or the communication medium 1105.

[0234] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0235] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the method of each embodiment of the present application.

[0236] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0237] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that contains one or more available media integration. Available media can be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state hard disk (SSD)), etc.

Claims

1. A training method for a recommendation model, characterized in that: include: Acquire a training set, wherein the training set includes a plurality of training data, each of which includes an interaction history between a user and a recommendation system, recommended content, and a feedback action for the recommended content; Based on the training set, a satisfaction prediction model is obtained through inverse reinforcement learning training, wherein the input of the satisfaction prediction model is the interaction history, the recommended content, and the feedback action, and the satisfaction prediction model is used to predict user satisfaction; Inputting the training set into the satisfaction prediction model to predict the user satisfaction corresponding to the training data in the training set; Training a recommendation model based on the training set and user satisfaction corresponding to the training data in the training set to obtain a trained recommendation model; Among them, the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

2. The method according to claim 1, characterized in that The training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

3. The method according to claim 1 or 2, characterized in that: The user satisfaction training recommendation model based on the training set and the training data in the training set includes: Inputting the interaction history and the recommended content in the target training data into the recommendation model, and constructing a first loss function based on the feedback action output by the recommendation model and the feedback action in the target training data, wherein the target training data is any one of the training data in the training set, and the first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data; Based on the feedback action output by the recommendation model and the user satisfaction corresponding to the target training data, constructing a second loss function, wherein the second loss function is obtained by summing the products between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions; Based on the first loss function and the second loss function, the recommendation model is updated.

4. The method according to claim 3, characterized in that The probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

5. The method according to claim 3 or 4, characterized in that: The updating of the recommendation model based on the first loss function and the second loss function includes: Performing a weighted summation on the first loss function and the second loss function to obtain a total loss function; Based on the total loss function, the recommendation model is updated by a gradient descent method.

6. The method according to any one of claims 1 to 5, characterized in that: The training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. The adoption of different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

7. The method according to any one of claims 1 to 6, characterized in that: The satisfaction prediction model includes a trunk network and multiple branch networks, the multiple branch networks are all connected to the trunk network, the input of the trunk network is the interaction history and the recommended content, and the multiple branch networks are respectively used to predict user satisfaction under different feedback actions.

8. The method according to any one of claims 1 to 7, characterized in that: The trained recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

9. The method according to any one of claims 1 to 8, characterized in that: The interaction history includes one or more contents recommended by the recommendation system and feedback actions of the user with respect to the one or more contents.

10. A content recommendation method, characterized in that: include: Acquire first content and a first interaction history, where the first interaction history is used to indicate a historical interaction behavior between the recommendation system and the target user; Inputting the first content and the interaction history into a recommendation model to obtain a feedback action of the target user with respect to the first content predicted by the recommendation model; Among them, the recommendation model is trained based on the training set and the user satisfaction corresponding to the training data in the training set, and the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the user satisfaction corresponding to the training data in the training set is predicted based on the satisfaction prediction model, and the satisfaction prediction model is trained based on the training set through inverse reinforcement learning.

11. A training device for a recommendation model, characterized in that: include: An acquisition module, used to acquire a training set, wherein the training set includes a plurality of training data, each of which includes an interaction history between a user and a recommendation system, recommended content, and a feedback action for the recommended content; A processing module, configured to obtain a satisfaction prediction model through inverse reinforcement learning training based on the training set, wherein the input of the satisfaction prediction model is the interaction history, the recommended content, and the feedback action, and the satisfaction prediction model is used to predict user satisfaction; The processing module is further used to input the training set into the satisfaction prediction model to predict the user satisfaction corresponding to the training data in the training set; The processing module is further used to train a recommendation model based on the training set and user satisfaction corresponding to the training data in the training set to obtain a trained recommendation model; Among them, the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data.

12. The device according to claim 11, characterized in that The training goal of the recommendation model during the training process also includes reducing the difference between the feedback action output by the recommendation model and the feedback action in the training data.

13. The device according to claim 11 or 12, characterized in that The processing module is specifically used for: Inputting the interaction history and the recommended content in the target training data into the recommendation model, and constructing a first loss function based on the feedback action output by the recommendation model and the feedback action in the target training data, wherein the target training data is any one of the training data in the training set, and the first loss function is used to indicate the difference between the feedback action output by the recommendation model and the feedback action in the training data; Based on the feedback action output by the recommendation model and the user satisfaction corresponding to the target training data, constructing a second loss function, wherein the second loss function is obtained by summing the products between the probabilities of various feedback actions output by the recommendation model and the user satisfaction corresponding to the various feedback actions; Based on the first loss function and the second loss function, the recommendation model is updated.

14. The device according to claim 13, characterized in that The probabilities of various feedback actions output by the recommendation model include the probability of clicking and the probability of not clicking.

15. The device according to claim 13 or 14, characterized in that The processing module is specifically used for: Performing a weighted summation on the first loss function and the second loss function to obtain a total loss function; Based on the total loss function, the recommendation model is updated by a gradient descent method.

16. The device according to any one of claims 11 to 15, characterized in that: The training goal of the satisfaction prediction model during the training process includes increasing the difference between the user satisfaction output by the satisfaction prediction model when different feedback actions are adopted under the same state. The adoption of different feedback actions under the same state means that the interaction history and recommendation content input into the satisfaction prediction model are the same, but the feedback actions input into the satisfaction prediction model are different.

17. The device according to any one of claims 11 to 16, characterized in that: The satisfaction prediction model includes a trunk network and multiple branch networks, the multiple branch networks are all connected to the trunk network, the input of the trunk network is the interaction history and the recommended content, and the multiple branch networks are respectively used to predict user satisfaction under different feedback actions.

18. The device according to any one of claims 11 to 17, characterized in that: The trained recommendation model is used to perform any one of the following tasks: an advertisement recommendation task, a product recommendation task, a video recommendation task, an application recommendation task, and a news recommendation task.

19. The device according to any one of claims 11 to 18, characterized in that: The interaction history includes one or more contents recommended by the recommendation system and feedback actions of the user with respect to the one or more contents.

20. A content recommendation device, characterized in that: include: An acquisition module, used to acquire first content and a first interaction history, where the first interaction history is used to indicate a historical interaction behavior between the recommendation system and the target user; A recommendation module, configured to input the first content and the interaction history into a recommendation model to obtain a feedback action of the target user with respect to the first content predicted by the recommendation model; Among them, the recommendation model is trained based on the training set and the user satisfaction corresponding to the training data in the training set, and the training goal of the recommendation model during the training process includes improving the matching degree between the feedback action output by the recommendation model and the user satisfaction corresponding to the training data, the user satisfaction corresponding to the training data in the training set is predicted based on the satisfaction prediction model, and the satisfaction prediction model is trained based on the training set through inverse reinforcement learning.

21. A training device for a recommendation model, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 9.

22. A content recommendation device, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method as claimed in claim 10.

23. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 10.

Citation Information

Cited By

  • Distributed artificial intelligence recommendation model training and application system oriented to multi-platform data collaboration

    CN122198043A