Content recommendation method, model training method, and related apparatus

By using two feature extraction networks in the content recommendation model to extract the features of interaction history and content of user interest, and combining them with the features of the content to be recommended, the problem of low prediction accuracy of the existing model is solved, and more accurate user feedback action prediction and more efficient model updates are achieved.

WO2025190196A1PCT designated stage Publication Date: 2025-09-18HUAWEI TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/081524
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2025-03-10
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing content recommendation models have low accuracy in predicting user feedback actions and are unable to effectively utilize massive user and content features, resulting in poor recommendation results.

Method used

Two feature extraction networks are used to extract the interaction history between the user and the recommendation system and the features of the content that the user is interested in. Combined with the features of the content to be recommended, the user's feedback action is predicted, and the content of user interest is used as prior knowledge to improve the prediction accuracy of the model.

Benefits of technology

It improves the accuracy of the content recommendation model in predicting user feedback actions, reduces the demand for training data, improves the update and iteration efficiency of the model, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081524_18092025_PF_FP_ABST
    Figure CN2025081524_18092025_PF_FP_ABST
Patent Text Reader

Abstract

A content recommendation method, which can effectively improve the accuracy of a model for predicting a feedback action of a user, thereby improving the accuracy of content recommendation. In the method, interaction histories between a user and a recommendation system and a feature of content to be recommended are extracted by means of one feature extraction network, a feature of content in which the user is interested is extracted from the interaction histories by means of another feature extraction network, and a feedback action of the user for the content to be recommended is then predicted on the basis of the two extracted features. In the present solution, a feature of content in which a user is interested is additionally extracted from interaction histories, i.e., priori knowledge is additionally provided during content recommendation, and the content in which the user is interested is highlighted in the interaction histories, such that when predicting a feedback action of the user, a model can focus on the content in which the user is interested, thereby effectively improving the accuracy of the model for predicting the feedback action of the user.
Need to check novelty before this filing date? Find Prior Art

Description

A content recommendation method, model training method and related devices

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 13, 2024, with application number 202410289990.9 and invention name “A content recommendation method, model training method and related devices”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a content recommendation method, a model training method, and related devices. Background Art

[0003] With the rapid development of the internet, it has become a crucial means of content delivery. More and more users are accessing content such as videos, news, and product information online. Currently, recommendation systems are used to recommend personalized content to users. These systems infer users' interests and usage behaviors based on information such as the characteristics of content they have already accessed and their own characteristics. They then recommend content based on these interests and behaviors, thereby improving the relevance of delivered content to users and achieving targeted content delivery.

[0004] At present, the recommendation system in related technologies is usually a content recommendation model based on deep learning. By inputting the characteristics of the content that the user has visited and the characteristics of the user himself into the content recommendation model, the content recommendation model predicts the probability of the user accessing specific content based on the multiple input features.

[0005] However, existing content recommendation models often need to learn the characteristics of user-preferred content from a large amount of input features to predict the probability of users accessing specific content. This prediction is difficult, resulting in low prediction accuracy of content recommendation models. Summary of the Invention

[0006] This application provides a content recommendation method that can effectively improve the accuracy of the model in predicting user feedback actions, thereby improving the accuracy of content recommendations.

[0007] In a first aspect, the present application provides a content recommendation method for recommending content to users online. The content recommendation method comprises: first, obtaining interaction history and content to be recommended. The interaction history is obtained based on the interaction between a deployed recommendation system and the user, and includes the recommendation system's historical recommendations and the user's feedback actions in response to the historical recommendations. The content to be recommended refers to content waiting to be recommended to the user.

[0008] Then, the interaction history and the content to be recommended are input into a first feature extraction network to obtain a first feature. The first feature extraction network is a neural network that can extract features from input data.

[0009] Next, the user's interesting content from the historical recommendations is fed into a second feature extraction network to obtain a second feature. The second feature extraction network is also a neural network capable of extracting features from the input data. Furthermore, the second feature extraction network can have a different network structure than the first feature extraction network.

[0010] Finally, based on the first and second features, the user's feedback action regarding the recommended content is predicted; and based on the predicted feedback action, content is recommended to the user. For example, based on the first and second features, a prediction network can be used to predict the user's feedback action regarding the recommended content, thereby outputting a predicted feedback action. Specifically, the prediction network can output the probability that the user will take a certain feedback action regarding the recommended content, such as the probability that the user will click on or download the recommended content. The prediction network can also output the degree to which the user will take a feedback action regarding the recommended content, such as the amount of money the user will spend on the recommended content.

[0011] In this solution, while a feature extraction network extracts features of the user's interaction history with the recommendation system and the content to be recommended, another feature extraction network also extracts features of the user's content of interest in the interaction history (i.e., the input of the other feature extraction network is a subset of the input of the previous feature extraction network). Based on the two extracted features, the user's feedback action on the content to be recommended is predicted. Because this solution additionally extracts features of the user's content of interest in the interaction history, it is equivalent to adding prior knowledge to the content recommendation process. It emphasizes the user's content of interest in the interaction history, allowing the model to focus on the user's content of interest when predicting the user's feedback action. This helps the model more easily learn the patterns of user preferences, thereby effectively improving the model's accuracy in predicting user feedback actions.

[0012] Furthermore, this approach uses user content of interest as the input to extract relevant features, effectively using user content of interest to characterize each user's uniqueness, eliminating the need for separate identifiers for each user. This effectively protects user privacy. Furthermore, in online recommendation scenarios, since content recommendation models require continuous iterative training and updates based on new interactions, characterizing users based on user content of interest allows for a large number of users to be represented using limited content data, thereby improving the update efficiency of the content recommendation model. For example, in a gaming scenario, a library of recommended games may contain tens of thousands or even hundreds of thousands of games, while the number of users may be in the millions or even tens of millions. Therefore, using a limited number of game identifiers can represent a large number of user features, thereby reducing the amount of data required. Since models often require all types of input to achieve good prediction results during training, characterizing users based on user content of interest effectively reduces the amount of data required to train the content recommendation model, thereby reducing the number of training cycles and improving the iterative efficiency of the content recommendation model.

[0013] In one possible implementation, the content of interest to the user is the content in the historical recommended content that corresponds to a positive feedback action, and the positive feedback action includes at least one of click, download, or consumption. In some cases, the feedback action can also be represented by some quantitative numerical values. For example, in the case where the content is a video or a song, after the user clicks on the video or song, it will trigger the state of watching the video or listening to the song, so the feedback action can specifically be the length of time the user watches the video after clicking on the video, or the length of time the user listens to the song after clicking on the song. For another example, in the case where the content is a game, after the user consumes in the game, there will be a corresponding consumption amount, so the feedback action can specifically be the amount of money the user spends in the game, or the number of times the user consumes the game.

[0014] In one possible implementation, predicting a user's feedback action for a proposed content based on the first and second features may include first obtaining features of the proposed content; then, based on the first and second features, predicting the user's feedback action for the proposed content. Specifically, features are extracted for the proposed content as content requiring emphasis, and these features are used together with the first and second features to predict the user's feedback action, thereby improving the accuracy of the model's prediction of user feedback actions.

[0015] Since the content to be recommended itself is a direct factor affecting whether the user will take feedback action, combining the content to be recommended with the content that the user is interested in as prior knowledge that the model needs to focus on can enable the model to pay more attention to the content to be recommended and the content that the user is interested in, thereby improving the accuracy of the model's prediction of the user's feedback action on the content to be recommended.

[0016] In one possible implementation, predicting a user's feedback action regarding the content to be recommended based on the first feature, the second feature, and the features of the content to be recommended may specifically include: fusing the second feature with the features of the content to be recommended to obtain a fused feature; and inputting the first feature and the fused feature into a prediction network to predict the user's feedback action regarding the content to be recommended. Fusion of the second feature with the features of the content to be recommended may, for example, be performed by performing a dot product operation between the second feature and the features of the content to be recommended.

[0017] In this solution, by fusing the second feature with the feature of the content to be recommended, the second feature and the feature of the content to be recommended can together constitute a collaborative feature, which is input into the prediction network as prior knowledge to predict the feedback action, which can effectively improve the prediction accuracy of the prediction network.

[0018] In one possible implementation, the first feature extraction network, the second feature extraction network, and the prediction network are uniformly trained based on the same training objective, and the prediction network is used to predict the feedback action based on the first feature and the second feature.

[0019] That is, the first feature extraction network, the second feature extraction network and the prediction network are trained as a unified whole with the same training objective, thereby ensuring that the final prediction network can predict accurate feedback actions based on the outputs of the first feature extraction network and the second feature extraction network.

[0020] In one possible implementation, a target model including a first feature extraction network, a second feature extraction network, and a prediction network is determined from multiple content recommendation models, and the multiple content recommendation models are trained to output a user's interest level in input content.

[0021] That is, when selecting multiple content recommendation models, the models are not trained to output user feedback actions, but rather to output the user's interest level in the input content. The user's interest level in the input content can be a value measured using a unified standard, representing the user's interest in the input content. By using a unified standard of interest level as the output of the content recommendation model, the content recommendation model can learn the user's willingness to provide feedback on the content, thereby ensuring that the content recommendation model's evaluation metrics are more closely aligned with the user's actual thoughts and ultimately ensuring that the optimal content recommendation model is selected.

[0022] In one possible implementation, when the prediction network is used to predict the amount of money a user will spend on content to be recommended, during the training process of multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user in the content already consumed.

[0023] That is, the higher the user's value measurement of the input content, the more worthy the user believes the input content is to be consumed, and the stronger the user's willingness to consume the input content is; the higher the proportion of the input content in the user's consumed content, the more money the user has spent on the input content, and the stronger the user's willingness to consume the input content is.

[0024] In a possible implementation, the user's value measurement of the input content is obtained based on the user's consumption amount of the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

[0025] For example, the user's value measurement for the input content is calculated by dividing the difference between the user's spending amount on the input content and the mean spending amount for the input content by the standard deviation of the spending amount for the input content. In this way, by incorporating the mean and standard deviation of the spending amount for the input content, combined with the user's spending amount for the input content, a unified standard can be used to assess the value of the input content in the user's mind.

[0026] In a possible implementation, the content to be recommended is any one of the following: advertisement, application, product, video, song or news.

[0027] A second aspect of the present application provides a model training method, including: obtaining training data, the training data including interaction history, recommended content, and target feedback actions of users with respect to the recommended content, the interaction history including historical recommended content of the recommendation system and feedback actions of users with respect to the historical recommended content; inputting the interaction history and recommended content into a first feature extraction network to obtain a first feature; inputting user-interested content in the historical recommended content into a second feature extraction network to obtain a second feature; based on the first feature and the second feature, predicting the user's feedback action with respect to the recommended content through a prediction network to obtain a predicted feedback action; training a target model through a loss function constructed based on the predicted feedback action and the target feedback action to obtain an updated target model, the target model including the first feature extraction network, the second feature extraction network, and the prediction network.

[0028] In a possible implementation, the user's interested content is content in historical recommended content that corresponds to a positive feedback action, where the positive feedback action includes one or more of the following feedback actions: click, download, or consume.

[0029] In one possible implementation, predicting a user's feedback action on recommended content based on a first feature and a second feature includes: obtaining features of the recommended content; and predicting the user's feedback action on the recommended content based on the first feature, the second feature, and the features of the recommended content.

[0030] In one possible implementation, based on the first feature, the second feature, and the feature of the recommended content, predicting the user's feedback action on the recommended content includes: fusing the second feature with the feature of the recommended content to obtain a fused feature; and inputting the first feature and the fused feature into a prediction network to predict the user's feedback action on the recommended content.

[0031] In one possible implementation, the target model is determined from multiple content recommendation models, and the multiple content recommendation models are trained to output the user's interest level in input content.

[0032] In one possible implementation, when the prediction network is used to predict the amount of money a user will spend on recommended content, during the training of multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user in the content already consumed.

[0033] In a possible implementation, the user's value measurement of the input content is obtained based on the user's consumption amount of the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

[0034] In a possible implementation, the recommended content is any one of the following: advertisement, application, product, video, song, or news.

[0035] In a third aspect, the present application provides a content recommendation device, comprising: an acquisition module, configured to acquire interaction history and content to be recommended, wherein the interaction history includes historical recommended content of the recommendation system and user feedback actions regarding the historical recommended content; a processing module, configured to input the interaction history and content to be recommended into a first feature extraction network to obtain a first feature; the processing module is further configured to input user-interested content in the historical recommended content into a second feature extraction network to obtain a second feature; the processing module is further configured to predict the user's feedback action regarding the content to be recommended based on the first feature and the second feature; and the processing module is further configured to recommend content to the user based on the predicted feedback action.

[0036] In a possible implementation, the user's interested content is content in historical recommended content that corresponds to a positive feedback action, where the positive feedback action includes at least one of click, download, or consumption.

[0037] In a possible implementation, the acquisition module is further configured to acquire features of the content to be recommended;

[0038] The processing module is further configured to predict the user's feedback action on the content to be recommended based on the first feature, the second feature, and the feature of the content to be recommended.

[0039] In a possible implementation, the processing module is further configured to:

[0040] Fusing the second feature with the feature of the content to be recommended to obtain a fused feature;

[0041] The first feature and the fused feature are input into the prediction network to predict the user's feedback action on the content to be recommended.

[0042] In one possible implementation, the first feature extraction network, the second feature extraction network, and the prediction network are uniformly trained based on the same training objective, and the prediction network is used to predict the feedback action based on the first feature and the second feature.

[0043] In one possible implementation, a target model including a first feature extraction network, a second feature extraction network, and a prediction network is determined from multiple content recommendation models, and the multiple content recommendation models are trained to output a user's interest level in input content.

[0044] In one possible implementation, when the prediction network is used to predict the amount of money a user will spend on content to be recommended, during the training process of multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user in the content already consumed.

[0045] In a possible implementation, the user's value measurement of the input content is obtained based on the user's consumption amount of the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

[0046] In a possible implementation, the content to be recommended is any one of the following: advertisement, application, product, video, song or news.

[0047] In a fourth aspect, the present application provides a model training device, comprising: an acquisition module for acquiring training data, the training data including interaction history, recommended content, and target feedback actions of users with respect to the recommended content, the interaction history including historical recommended content of the recommendation system and feedback actions of users with respect to the historical recommended content; a recommendation module for inputting the interaction history and recommended content into a first feature extraction network to obtain a first feature; the recommendation module is further configured to input user-interested content in historical recommended content into a second feature extraction network to obtain a second feature; the recommendation module is further configured to predict user feedback actions with respect to the recommended content through a prediction network based on the first feature and the second feature to obtain a predicted feedback action; the recommendation module is further configured to train a target model using a loss function constructed based on the predicted feedback action and the target feedback action to obtain an updated target model, the target model comprising a first feature extraction network, a second feature extraction network, and a prediction network.

[0048] In a possible implementation, the user's interested content is content in historical recommended content that corresponds to a positive feedback action, where the positive feedback action includes at least one of click, download, or consumption.

[0049] In a possible implementation, the acquisition module is further configured to acquire features of the recommended content; and the recommendation module is further configured to predict a user's feedback action on the recommended content based on the first feature, the second feature, and the features of the recommended content.

[0050] In a possible implementation, the recommendation module is further configured to: fuse the second feature with the feature of the recommended content to obtain a fused feature; and input the first feature and the fused feature into a prediction network to predict a user's feedback action on the recommended content.

[0051] In one possible implementation, the target model is determined from multiple content recommendation models, and the multiple content recommendation models are trained to output the user's interest level in input content.

[0052] In one possible implementation, when the prediction network is used to predict the amount of money a user will spend on content to be recommended, during the training process of multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user in the content already consumed.

[0053] In a possible implementation, the user's value measurement of the input content is obtained based on the user's consumption amount of the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

[0054] In a possible implementation, the recommended content is any one of the following: advertisement, application, product, video, song, or news.

[0055] In a fifth aspect, the present application provides a content recommendation device, which may include a processor coupled to a memory, wherein the memory stores program instructions. When the processor executes the program instructions stored in the memory, the method of the first aspect or any implementation of the first aspect is implemented. For details regarding the steps in each possible implementation of the first aspect performed by the processor, please refer to the first aspect and will not be repeated here.

[0056] In a sixth aspect, the present application provides a model training device, which may include a processor coupled to a memory, wherein the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method of the second aspect or any implementation of the second aspect is implemented. For the steps in each possible implementation of the second aspect executed by the processor, please refer to the second aspect for details, and no further description is given here.

[0057] In a seventh aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes a method implemented in any one of the first and second aspects.

[0058] In an eighth aspect, the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute a method implemented in any one of the first or second aspects.

[0059] In a ninth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method implemented in any one of the first or second aspects.

[0060] In a tenth aspect, the present application provides a chip system, which includes a processor for supporting an electronic device to implement the functions involved in any implementation of the first or second aspect above, for example, processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory for storing program instructions and data necessary for the electronic device. The chip system can be composed of a chip or can include a chip and other discrete devices.

[0061] The beneficial effects of the second to tenth aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] FIG1 is a schematic diagram of content recommendation in an application store according to an embodiment of the present application;

[0063] FIG2 is a schematic diagram of a recommendation system provided in an embodiment of the present application;

[0064] FIG3 is a schematic diagram of a system architecture 300 provided in an embodiment of the present application;

[0065] FIG4 is a flow chart of a content recommendation method provided in an embodiment of the present application;

[0066] FIG5 is a schematic diagram illustrating an implementation of a content recommendation method provided in an embodiment of the present application;

[0067] FIG6 is another schematic diagram illustrating another implementation of the content recommendation method provided in an embodiment of the present application;

[0068] FIG7 is a flow chart of a model training method provided in an embodiment of the present application;

[0069] FIG8 is a schematic diagram of a selection process of a content recommendation model provided in an embodiment of the present application;

[0070] FIG9 is a schematic diagram of a process for implementing game recommendations based on a retrained target model according to an embodiment of the present application;

[0071] FIG10 is a schematic diagram of inputting different data on different networks of a target model to implement game recommendations according to an embodiment of the present application;

[0072] FIG11 is a schematic structural diagram of a content recommendation device provided in an embodiment of the present application;

[0073] FIG12 is a schematic structural diagram of a model training device provided in an embodiment of the present application;

[0074] FIG13 is a schematic diagram of a structure of an execution device provided in an embodiment of the present application;

[0075] FIG14 is a schematic structural diagram of a chip provided in an embodiment of the present application;

[0076] FIG15 is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0077] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0078] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.

[0079] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.

[0080] (1) Content recommendation model

[0081] The content recommendation model is a neural network that uses machine learning algorithms to analyze and learn based on the user's historical access content and the user's own characteristics, determine the probability of the user accessing specific content (i.e., click-through rate), and then recommend content that the user is likely to be interested in based on the probability of the user accessing various content.

[0082] (2) Click-through rate

[0083] Click-through rate refers to the probability that a user clicks on a certain displayed content in a specific environment.

[0084] (3) Neural Network

[0085] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:

[0086] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0087] (4) Deep Neural Network (DNN)

[0088] Deep neural networks, also known as multi-layer neural networks, can be understood as neural networks with many hidden layers. There is no special metric for "many" here. Based on the position of different layers in DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscript corresponds to the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).

[0089] (5) Multilayer Perceptron (MLP)

[0090] The multilayer perceptron is a classic neural network model composed of multiple layers of neurons. Generally, a multilayer perceptron consists of an input layer, a hidden layer, and an output layer. The input layer receives input data, the hidden layer learns feature representations, and the output layer produces the final output. Furthermore, each neuron in the hidden and output layers typically has an activation function to introduce nonlinear mapping.

[0091] (6) Loss function

[0092] During neural network training, because we want the output of the neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vectors of each layer of the neural network based on the difference between the two. (Of course, before the first update, there is usually an initialization process, which pre-configures the parameters of each layer in the neural network.) For example, if the network's prediction is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so neural network training becomes a process of minimizing this loss.

[0093] (7) Backpropagation algorithm

[0094] Neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial prediction model during training, reducing the error loss of the prediction model. Specifically, the forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the parameters in the initial prediction model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by error loss, aiming to obtain the optimal prediction model parameters, such as the weight matrix.

[0095] Specifically, during model training, the backpropagation algorithm is typically used to calculate the gradient of each node in the model. This allows the node weight parameters to be adjusted based on the gradient of each node, thereby minimizing the model's loss function. The gradient represents the rate of change of a function at a specific point. Furthermore, the gradient of each node in the model can be determined by taking partial derivatives.

[0096] (8) Gradient descent

[0097] Gradient descent is a first-order optimization algorithm commonly used in machine learning to recursively approximate minimum-deviation prediction models. To find a local minimum of a function using gradient descent, an iterative search must be performed toward a point on the function at a specified step distance in the opposite direction of the gradient (or approximate gradient) corresponding to the current point. Gradient descent is one of the most commonly used methods for solving unconstrained optimization problems, such as predictive model parameters in machine learning algorithms.

[0098] Specifically, when solving for the minimum value of the loss function, we can use the gradient descent method to iterate step by step to obtain the minimized loss function and prediction model parameter values. Conversely, if we need to solve for the maximum value of the loss function, we need to use the gradient ascent method to iterate.

[0099] At present, the recommendation system in related technologies is usually a content recommendation model based on deep learning. By inputting the characteristics of the content that the user has visited and the characteristics of the user himself into the content recommendation model, the content recommendation model predicts the probability of the user accessing specific content based on the multiple input features.

[0100] However, throughout the entire operation of a recommendation system, logs can record thousands of user characteristics (such as age, gender, hobbies, and city) and content characteristics (such as content clicked, not clicked, downloaded, and consumed). Therefore, existing content recommendation models often need to learn the characteristics of user-favorite content from a massive amount of input features to predict the probability of a user accessing a specific piece of content. This makes prediction difficult and results in low prediction accuracy.

[0101] Based on this, while one feature extraction network extracts features of the user's interaction history with the recommendation system and the content to be recommended, another feature extraction network also extracts features of the user's content of interest in the interaction history. Based on these two extracted features, the user's feedback action for the content to be recommended is predicted. Because this solution additionally extracts features of the user's content of interest from the interaction history, it adds prior knowledge to the content recommendation process, highlighting the user's content of interest in the interaction history. This allows the model to focus on the user's content of interest when predicting the user's feedback action, effectively improving the model's accuracy in predicting user feedback actions.

[0102] The method provided in the embodiments of the present application can be applied to various content recommendation scenarios, such as the recommendation of products, applications, videos, news, or songs. In addition, in the embodiments of the present application, advertisements can be the carrier of these contents, that is, the content information can be displayed on the advertisement page. Therefore, the method provided in the embodiments of the present application can also be applied to the recommendation scenario of advertisements.

[0103] In one possible scenario, the method provided by the embodiment of the present application can be applied to a scenario where application software or game software is recommended as content on the interface of an application store. For example, please refer to Figure 1, which is a schematic diagram of content recommendation of an application store provided by an embodiment of the present application. As shown in Figure 1, for a certain application software - application store on the user's mobile phone, the application store is used by the user to download various application software or game software. In the interface of the application store, various software recommended by the recommendation system are displayed, that is, the various application software displayed under the "Quality Applications" on the interface.

[0104] In another possible scenario, the method provided by the embodiments of the present application can also be applied to a scenario where products or applications are recommended as content on social software, video software, and other application software. For example, when a user is watching a video on a video software, during the period when the user pauses watching the video, the video software can play a corresponding advertisement to recommend a certain product or application.

[0105] In another possible scenario, the method provided by the embodiments of the present application can also be used to recommend products on online shopping apps or web pages. For example, when a user searches or browses various products on an online shopping app, the app can pin some products that the user is likely to be interested in to the top of the list, giving priority to recommending these top products to the user.

[0106] In general, the embodiments of the present application do not limit the specific type of recommended content.

[0107] The system scenario applied in the embodiment of the present application is an application scenario based on machine learning. The following will be introduced using the click-through rate prediction scenario in the recommendation system as an example, where the click-through rate is a type of advertising conversion rate. The click-through rate prediction scenario is a typical scenario in machine learning applications, and its main structure is shown in Figure 2. Among them, Figure 2 is a schematic diagram of a recommendation system provided by the embodiment of the present application. As shown in Figure 2, the recommendation system includes a log, an offline training module, a prediction model, an online prediction module, and a display list.

[0108] The basic operating logic of a recommendation system is as follows: users perform a series of actions within the front-end display list, such as browsing, clicking, commenting, and downloading. This generates behavioral data, which is stored in logs. The offline training module in the recommendation system utilizes data, including user behavior logs, and combines user characteristics (such as age, city, purchase history, and download history), product characteristics (such as product category, product description, and product attribute tags), and specific contextual information (such as whether it is a weekend or a holiday) to construct training data. Each sample in the training data includes user characteristics, content characteristics, contextual characteristics, and the user's feedback action on the content (such as the amount a user spent on a game app). Given the training data, the offline training module uses a predefined machine learning algorithm to iteratively update the parameters of the constructed prediction model until the predefined requirements are met. After training is complete, the trained content recommendation model is output to the online prediction module for use. The online prediction module is responsible for receiving the content recommendation model generated by the offline training module. When a user initiates a request, the model's interface is called to predict the user's feedback action on a given content (such as the user's possible consumption value for a game application), and the predicted feedback action is sorted, and the top N content is displayed on the user interface.

[0109] For example, when a user opens a mobile app store, it triggers a request to the recommendation system. The recommendation system then obtains input information related to the request, including user characteristics (such as the user's city and their download history), application characteristics (such as application category and developer), and environmental characteristics (such as time and network conditions). Using this information as input, the online prediction module's interface is called to predict the probability of the user downloading each application. Based on the prediction results, the recommendation system displays the top N applications that the user is most likely to download. Simultaneously, user interaction records in the app store, such as browsing, clicking, and downloading, are stored in log files. The offline training module then retrains and updates the prediction model to further improve its prediction accuracy. This results in a trained prediction model. The prediction model is then deployed in an online service environment to form an online prediction module. This module generates recommendations based on the user's requested access, item characteristics, and contextual information, and then displays the recommendations in a display list. Finally, the user's feedback on the recommendations in the display list forms user behavior data, which is also stored in the log.

[0110] Please refer to Figure 3, which is a schematic diagram of a system architecture 300 provided in an embodiment of the present application. As shown in Figure 3, in the system architecture 300, the execution device 310 can be implemented by one or more servers. Optionally, the execution device 310 cooperates with other computing devices, such as data storage devices, routers, load balancers and other devices; the execution device 310 can be arranged at one physical site, or distributed across multiple physical sites. The execution device 310 can use the data in the data storage system 320, or call the program code in the data storage system 320 to implement the content recommendation method provided in the embodiment of the present application, and then obtain the content that needs to be recommended to the user.

[0111] Users can operate their respective user devices (such as local device 301 and local device 302) to interact with execution device 310. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.

[0112] Each user's local device can interact with the execution device 310 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0113] In one implementation, the execution device 310 is used to implement the model training method provided in the embodiments of the present application to obtain a model for performing content recommendation. Furthermore, during the process of local devices 301 and 302 accessing content, the execution device 310 predicts the feedback actions that users may take for various contents based on the trained model and the content recommendation method of this embodiment, and then returns the corresponding recommended content (i.e., content that users will take positive feedback actions) to the local devices 301 and 302.

[0114] In another implementation, one or more aspects of the execution device 310 can be implemented by each local device. For example, the local device 301 can provide local data or feedback calculation results to the execution device 310, or execute the content recommendation method provided in the embodiment of the present application.

[0115] It should be noted that all functions of the execution device 310 can also be implemented by the local device. For example, the local device 301 implements the functions of the execution device 310 and provides services to its own user, or provides services to the user of the local device 302.

[0116] In general, the content recommendation method provided in the embodiments of the present application can be applied to electronic devices, such as the aforementioned execution device 310 , local device 301 , or local device 302 .

[0117] Please refer to Figure 4, which is a flowchart of a content recommendation method provided by an embodiment of the present application. As shown in Figure 4, the content recommendation method includes the following steps 401-405.

[0118] Step 401: Obtain interaction history and content to be recommended. The interaction history includes historical recommended content of the recommendation system and user feedback actions on the historical recommended content.

[0119] In this embodiment, the interaction history can be obtained based on the interaction behavior between the deployed recommendation system and the user. Specifically, the interaction history between the user and the recommendation system includes the content recommended to the user by the recommendation system in the previous recommendation process (i.e., the historical recommended content of the recommendation system) and the feedback actions taken by the user in response to the historical recommended content of the recommendation system. Exemplarily, the interaction history can be represented by the following content: content that the user did not click, content that the user marked as not wanting to be recommended again, content that the user clicked, content that the user downloaded, and content that the user consumed. In this way, by using multiple types of content to represent the interaction history in batches, it is possible to effectively represent the various content that appears in the interaction history and the feedback actions corresponding to the various content.

[0120] In addition, the interaction history between the user and the recommendation system may also include the user's characteristic information and the environment information when the user interacts with the recommendation system. The user's characteristic information includes, for example, the user's age, gender, interests, and the city where the user is located. The environment information when the user interacts with the recommendation system includes, for example, the time when the user interacts with the recommendation system (whether it is evening, holiday, etc.) and the weather when the user interacts with the recommendation system (whether it is raining or sunny, etc.).

[0121] The interaction history between a user and a recommendation system may include one or more rounds of interaction between the user and the recommendation system. That is, the interaction history actually includes one or more content recommended by the recommendation system and the user's feedback actions in response to the one or more content. Moreover, the one or more content recommended by the recommendation system included in the interaction history may be recommended in one round of interaction or recommended in multiple rounds of interaction. For example, the interaction history between a user and a recommendation system records multiple rounds of interaction between the user and the recommendation system, and each round of interaction includes one or more recommended content recommended by the recommendation system and the user's feedback actions in response to the recommended content.

[0122] In addition, the content to be recommended refers to content waiting to be recommended to the user. Optionally, the content to be recommended is any one of the following: advertisements, applications, products, videos, songs, or news. This embodiment does not limit the type of content to be recommended. It should be noted that the content to be recommended is not necessarily recommended to the user. Instead, it is necessary to predict the user's feedback action for each content to be recommended based on the content recommendation method provided by this embodiment, and then determine the content that needs to be actually recommended to the user on the recommendation system.

[0123] Optionally, the type of feedback action that a user can take may be related to the recommendation scenario. For example, the feedback action taken by the user may include click, download, or consumption, etc., which are not specifically limited here. In the case where the feedback action taken by the user is consumption, the user's feedback action may also include the user's consumption amount.

[0124] Step 402: Input the interaction history and the content to be recommended into a first feature extraction network to obtain a first feature.

[0125] The interaction history and the content to be recommended are input into the first feature extraction network together, and the first feature extracted by the first feature extraction network for the input interaction history and the content to be recommended is obtained.

[0126] In this embodiment, the first feature extraction network is a neural network, and the structure of the first feature extraction network can be implemented in a variety of ways, as long as it can perform feature extraction on a large amount of input content. For example, the first feature extraction network is a deep neural network, a product-based neural network (PNN), a deep interest network (DIN), or a deep interest evolution network (DIEN).

[0127] Step 403: Input the user's interested content in the historical recommended content into the second feature extraction network to obtain the second feature.

[0128] In this embodiment, the second feature extraction network is independent of the first feature extraction network and is specifically used to extract features of content that the user is interested in. The content of interest to the user is selected from the historical recommended content in the interaction history, that is, the content of interest to the user is a portion of the content that the recommendation system has recommended to the user.

[0129] For example, user-interested content refers to content in historically recommended content that corresponds to positive feedback actions. Positive feedback actions include one or more of the following feedback actions: click, download, or consume. In other words, user-interested content can be one or more of the following types of content: content that the user has clicked on, downloaded, or consumed.

[0130] It is understandable that, compared to the first feature extraction network, the input of the second feature extraction network is only a small part of the input of the first feature extraction network. Moreover, the input type of the second feature extraction network is relatively simple, and is only the content of interest to the user. Therefore, the structure of the second feature extraction network can be different from that of the first feature extraction network, that is, the second feature extraction network can be implemented using a simpler network structure. For example, the second feature extraction network can be implemented using a multilayer perceptron (MLP).

[0131] Step 404 : predicting the user's feedback action on the content to be recommended based on the first feature and the second feature.

[0132] In this embodiment, based on the first and second features, a prediction network can be used to predict the user's feedback action regarding the recommended content, thereby outputting a predicted feedback action. Specifically, the prediction network can output the probability that the user will take a certain feedback action regarding the recommended content, such as the probability of the user clicking on or downloading the recommended content. Alternatively, the prediction network can output the degree to which the user will take a feedback action regarding the recommended content, such as the amount of money the user will spend on the recommended content. In short, the output of the prediction network can be determined based on actual scenarios and is not specifically limited in this embodiment.

[0133] Step 405: Recommend content to the user based on the predicted feedback action.

[0134] Specifically, for different content to be recommended, the above steps 401-404 will be used to predict the user's feedback action for each content to be recommended, so that content can be recommended to the user based on the feedback action corresponding to each content to be recommended, for example, content that the user is most likely to take positive feedback action (such as click, download or pay) can be recommended to the user.

[0135] For example, please refer to Figure 5, which is a schematic diagram of the execution of a content recommendation method provided in an embodiment of the present application. As shown in Figure 5, the interaction history between the user and the recommendation system and the content to be recommended are first input into the first feature extraction network, and at the same time, the user's content of interest extracted from the interaction history is input into the second feature extraction network to obtain the first feature output by the first feature extraction network and the second feature output by the second feature extraction network. Among them, the network structure of the first feature extraction network is different from the network structure of the second feature extraction network, and the parameter amount of the network structure of the first feature extraction network is greater than the parameter amount of the network structure of the second feature extraction network. Then, by inputting the first feature and the second feature into the prediction network, the user's feedback action on the content to be recommended predicted by the prediction network is obtained.

[0136] Optionally, in some embodiments, features of the content to be recommended may be obtained before predicting the feedback action. Then, based on the first and second features, and the features of the content to be recommended, the user's feedback action for the content to be recommended is predicted. Specifically, features of the content to be recommended are extracted separately as content that requires emphasis, and these features are used together with the first and second features to predict the user's feedback action, thereby improving the accuracy of the model's prediction of the user's feedback action.

[0137] Since the content to be recommended itself is a direct factor affecting whether the user will take feedback action, combining the content to be recommended with the content that the user is interested in as prior knowledge that the model needs to focus on can enable the model to pay more attention to the content to be recommended and the content that the user is interested in, thereby improving the accuracy of the model's prediction of the user's feedback action on the content to be recommended.

[0138] It should be noted that in this embodiment, there are multiple ways to obtain the features of the content to be recommended. For example, the features can be obtained by converting the identifier of the content to be recommended into features. Another example is by combining the identifier of the content to be recommended with the type of the content to be recommended and converting these features into features. In general, different content to be recommended can be converted into different features.

[0139] Optionally, when predicting a user's feedback action regarding the content to be recommended based on the first feature, the second feature, and the features of the content to be recommended, the second feature may be first fused with the features of the content to be recommended to obtain a fused feature. The fusion of the second feature with the features of the content to be recommended may be performed, for example, by performing a dot product operation on the second feature and the features of the content to be recommended.

[0140] Then, the first feature and the fused feature are input into the prediction network to predict the user's feedback action for the content to be recommended. For example, the first feature and the fused feature are concatenated and input into the prediction network to obtain the feedback action output by the prediction network.

[0141] For example, please refer to Figure 6, which is another execution diagram of the content recommendation method provided in an embodiment of the present application. As shown in Figure 6, the interaction history between the user and the recommendation system and the content to be recommended are first input into the first feature extraction network, and at the same time, the user's content of interest extracted from the interaction history is input into the second feature extraction network to obtain the first feature output by the first feature extraction network and the second feature output by the second feature extraction network. Then, the second feature is multiplied by the feature of the content to be recommended to obtain a fused feature. Finally, by inputting the first feature and the fused feature into the prediction network, the user's feedback action on the content to be recommended predicted by the prediction network is obtained.

[0142] In this embodiment, the first feature extraction network, the second feature extraction network and the prediction network are trained uniformly based on the same training objective, and the prediction network is used to predict feedback actions based on input features (such as features output by the first feature extraction network and the second feature extraction network).

[0143] Specifically, the first feature extraction network, the second feature extraction network, and the prediction network are used to form a target model for predicting feedback actions. During each iteration of the training phase, the parameters of the first feature extraction network, the second feature extraction network, and the prediction network are updated using a backpropagation algorithm. In other words, the first feature extraction network, the second feature extraction network, and the prediction network are trained as a unified whole with the same training objective, ensuring that the final prediction network can accurately predict feedback actions based on the outputs of the first and second feature extraction networks.

[0144] The above describes how to predict user feedback actions for recommended content based on a trained network. The following describes the detailed process of determining the network structure used to predict feedback actions.

[0145] Specifically, the network structures of the first feature extraction network, the second feature extraction network, and the prediction network described above are all fixed, meaning the network structure of the target model used to implement content recommendation is well-defined. However, in actual applications, researchers often need to first build multiple content recommendation models with different network structures and then select the one with the best performance from these multiple content recommendation models to be used as the final model for online application, in order to ensure the accuracy of model-based content recommendations.

[0146] Generally speaking, in the traditional process of selecting content recommendation models, multiple candidate content recommendation models are trained to output user feedback actions in response to input content during the selection phase. The performance of each content recommendation model is then determined by evaluating the accuracy of the feedback actions predicted by each trained content recommendation model. However, this traditional method of selecting content recommendation models does not effectively evaluate them, and the selected content recommendation models often fail to recommend content that users are actually interested in when deployed online.

[0147] Therefore, in this embodiment, a new method for selecting a content recommendation model is proposed, which can ensure the effectiveness of the process of selecting a content recommendation model and ensure that the selected content recommendation model is the model with the best performance and the best ability to recommend content of interest to users.

[0148] Specifically, the target model, comprising the first feature extraction network, the second feature extraction network, and the prediction network, is determined from a plurality of content recommendation models. Furthermore, during the selection phase of the plurality of content recommendation models, the plurality of content recommendation models are trained to output the user's level of interest in the input content. That is, when selecting the plurality of content recommendation models, the plurality of content recommendation models are not trained to output the user's feedback actions, but rather to output the user's level of interest in the input content. The user's level of interest in the input content can be a degree value measured using a unified standard, which can represent the user's level of interest in the input content.

[0149] Specifically, when a content recommendation model is ultimately used to predict the probability of a user clicking on content, interest level can be used to measure the user's willingness to click on the content. When a content recommendation model is ultimately used to predict the amount of money a user will spend on content, interest level can be used to measure the user's willingness to pay for the content. In general, by using a standardized interest level as the output of a content recommendation model, the model can learn from the user's willingness to provide feedback on the content, thereby ensuring that the model's evaluation metrics are more aligned with the user's actual thoughts and ultimately selecting the optimal content recommendation model.

[0150] For example, when the prediction network in the target model is used to predict the amount of money a user will spend on content to be recommended, during the training process of multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user in the content already consumed.

[0151] Generally speaking, the training process of a content recommendation model refers to the process of constructing a loss function based on the output of the content recommendation model and the true value corresponding to the input content (i.e., the label of the input content), and performing training through the loss function. Therefore, the true value corresponding to the input content used to construct the loss function is an important factor for constraining the output of the content recommendation model. In this embodiment, by using the true value of the user's interest in the input content as the label of the input content of the content recommendation model, the content recommendation model can be constrained to be trained to output the user's interest in the input content, thereby completing the training of multiple content recommendation models.

[0152] The true value of a user's interest in the input content is constructed based on a unified standard, specifically based on the user's value assessment of the input content and the percentage of the input content consumed by the user. For example, the average of the user's value assessment of the input content and the percentage of the input content consumed by the user. Specifically, the higher the user's value assessment of the input content, the more worthy the user believes the input content is, and the stronger the user's willingness to consume the input content. The higher the percentage of the input content consumed by the user, the more money the user has spent on the input content, and the stronger the user's willingness to consume the input content.

[0153] Optionally, the user's value measurement for the input content is derived based on the user's spending amount on the input content, the mean spending amount for the input content, and the standard deviation of the spending amount for the input content. For example, the user's value measurement for the input content is calculated by dividing the difference between the user's spending amount for the input content and the mean spending amount for the input content by the standard deviation of the spending amount for the input content. Thus, by introducing the mean and standard deviation of the spending amount for the input content, combined with the user's spending amount for the input content, a unified standard can be used to assess the value of the input content in the user's mind.

[0154] In general, due to the uncertainty of user consumption time and amount, when collecting the user's consumption amount for various contents, it is easy for the collected user consumption amount labels to have a large number of zero values, large variance, and a small number of maximum values. In this way, in the model evaluation stage, when the consumption amount is used as a label to train and evaluate the model, it is easy for the uncertainty of the consumption amount value to cause the effectiveness of the model to be unable to be fairly evaluated. Therefore, in this embodiment, the user's value measurement value for the input content and the proportion of the input content in the user's consumed content are introduced as the true value of the user's interest in the input content. This can standardize the labels of the input content and avoid the problem of a large number of zero values, large variance, and a small number of maximum values ​​in the labels of the training data during the training process, thereby improving the effectiveness of the final evaluation model and helping researchers find the content recommendation model with the best performance.

[0155] It should be noted that the above uses the consumption amount as an example to introduce how to determine the user's interest level, but in fact, some feedback actions that can be expressed by quantitative values ​​can also refer to the above method to determine the user's interest level. For example, replacing the above consumption amount with feedback actions such as video viewing time, song listening time, and number of consumptions can achieve the determination of the user's interest level in different scenarios. I will not go into details here.

[0156] In general, in this embodiment, the process of selecting a content recommendation model involves training multiple candidate content recommendation models to output the user's interest level in the input content. The optimally performing model is then selected from the trained models as the target model for online application. After the target model is selected, its network structure remains fixed, but the training objectives must be reset and the model trained again until it is ultimately trained to output the user's feedback actions in response to the input content.

[0157] The following describes the process of training the target model after selecting it so that it can be put into online application.

[0158] For example, please refer to Figure 7, which is a flow chart of a model training method provided in an embodiment of the present application. As shown in Figure 7, the model training method includes the following steps 701-705.

[0159] Step 701: Acquire training data. The training data includes interaction history, recommended content, and user's target feedback actions for the recommended content. The interaction history includes the recommendation system's historical recommended content and user's feedback actions for the historical recommended content.

[0160] In this embodiment, before training the target model, training data is first obtained based on the historical interactions between users and the recommendation system. This training data includes interaction history, recommended content, and the user's targeted feedback actions in response to the recommended content. The user's targeted feedback actions in response to the recommended content can be understood as labels for the training data.

[0161] Step 702: Input the interaction history and recommended content into a first feature extraction network to obtain a first feature.

[0162] Step 703: Input the user's interested content in the historical recommended content into the second feature extraction network to obtain the second feature.

[0163] Step 704 : Based on the first feature and the second feature, predict the user's feedback action for the content to be recommended through a prediction network to obtain a predicted feedback action.

[0164] Among them, steps 702-704 are similar to the above steps 402-404. For details, please refer to the above steps 402-404 and will not be repeated here.

[0165] Step 705 , training the target model by using a loss function constructed based on the prediction feedback action and the target feedback action to obtain an updated target model, where the target model includes a first feature extraction network, a second feature extraction network, and a prediction network.

[0166] After obtaining the predicted feedback action output by the target model, a loss function can be constructed based on the predicted feedback action and the label of the training data (i.e., the user's target feedback action for the recommended content), so as to further update the parameters of the target model based on the loss function to realize the training of the target model.

[0167] In the process of updating the parameters of the target model based on the loss function, the backpropagation algorithm and the gradient descent method can be used until the value of the loss function reaches a preset threshold, or the number of iterations of training the target model reaches a preset number. The training termination conditions of the target model are not limited here.

[0168] Optionally, based on the first feature and the second feature, predicting the user's feedback action on the content to be recommended includes: obtaining the features of the content to be recommended; and predicting the user's feedback action on the content to be recommended based on the first feature, the second feature and the features of the content to be recommended.

[0169] Optionally, based on the first feature, the second feature and the feature of the content to be recommended, predicting the user's feedback action on the content to be recommended, including: fusing the second feature and the feature of the content to be recommended to obtain a fused feature; inputting the first feature and the fused feature into a prediction network to predict the user's feedback action on the content to be recommended.

[0170] To facilitate understanding, the following will introduce in detail the process of selecting a content recommendation model and implementing content recommendation based on the selected content recommendation model with specific examples.

[0171] For example, please refer to Figure 8, which is a schematic diagram of a selection process of a content recommendation model provided in an embodiment of the present application. As shown in Figure 8, the selection process of the content recommendation model includes the following steps 801-804.

[0172] Step 801: Label standardization of training data.

[0173] When selecting a content recommendation model, it is necessary to first standardize the labels of the training data so that the model can be trained based on the standardized labels. Furthermore, the content recommendation model in this embodiment is a game recommendation model, which is used to recommend games that users are most likely to consume and spend the most money on.

[0174] Specifically, label normalization of training data is divided into three stages: Stage 1 to Stage 3. The following will introduce the three stages of label normalization.

[0175] In the first stage, based on the consumption history of the game to be recommended, the user's value measurement of the game is obtained.

[0176] Specifically, for game p, combined with the historical consumption record set corresponding to game p The tag data s corresponding to the game p can be up Normalize the data to get the value of the game to the user (that is, the value of the game in the user's mind). up The normalization process is shown in the following formula.

[0177] in, Indicates the user's value measurement of the game p; s up represents the amount of money spent by user u on game p in the training dataset; represents the mean of all consumption amounts in the consumption history corresponding to game p; Represents the standard deviation of all consumption amounts in the consumption history corresponding to game p.

[0178] From the above formula we can see that Represents the overall consumption level of the game p, s up and The larger the difference between them, the higher the user's willingness to pay for game p, that is, the higher the user's interest in game p.

[0179] Furthermore, the above standardization process effectively expresses user interest in various games, ensuring that interest is consistently measured across all spending amounts. For example, some games are inherently high-paying. If a user spends the same amount on a low-paying game as on a high-paying game, this indicates the low-paying game is more valuable to the user, meaning the user is more interested in the paid game.

[0180] In the second stage, based on the user's consumption history, the proportion of the recommended game in the user's total consumption is calculated.

[0181] Specifically, the proportion of the consumption of the game to be recommended in the amount already spent by the user can be expressed as the following formula.

[0182] in, Indicates the proportion of the recommended game in the user's total spending; up represents the amount of money spent by user u on game p in the training dataset; Indicates the total amount of money spent by user u on games in the past 180 days; Indicates the total consumption times of user u in the past 180 days.

[0183] In the third stage, the user's value measurement of the game and the proportion of the recommended game's consumption in the user's total consumption are combined to calculate the user's interest in the recommended game and obtain the standardized label of the training data.

[0184] Specifically, combined with the standardized consumption values ​​on the user side and the normalized values ​​on the game side The final normalized value can be calculated in, The user's interest in the recommended games, which represents the normalized label of the training data.

[0185] Step 802 : training multiple candidate content recommendation models based on the standardized labels.

[0186] When training multiple candidate content recommendation models, any candidate content recommendation model can be trained using the same training set and the same training method. During the training of the candidate content recommendation model, the loss function can be constructed as shown in the following formula.

[0187] in, The loss function used to train the candidate content recommendation model; The standardized labels for the training data (i.e., the user's interest in the recommended games); The predicted value output by the candidate content recommendation model; is the observation data in the training set, that is, the set of all interactions between users and the game.

[0188] In the specific model training process, by minimizing the above loss function Optimize the trainable parameters in the candidate content recommendation model and finally complete the training of multiple candidate content recommendation models.

[0189] Step 803: Construct evaluation data.

[0190] For any interaction record between user u and the recommended game p in the test set, randomly select the game pool based on all games. Randomly select 100 uninteracted games As negative samples, and take the symbol = represents a negative sample set consisting of 100 negative samples. In this embodiment, negative samples refer to content that is within the user's exposure range but the user has not interacted with it (i.e., games that the user has not consumed).

[0191] Step 804 : Evaluate the trained multiple candidate content recommendation models based on the evaluation data and determine the target model for online application.

[0192] In the model evaluation phase, for any interaction record (u, p) between a user and the game to be recommended in the test set, we first obtain the negative sample set corresponding to the interaction record. Then, multiple trained candidate content recommendation models are used to calculate the user u’s interest in the interactive game p, as well as the user u’s interest in the negative sample set. Each negative sample in and sort all obtained interest levels.

[0193] Then, the hit rate at k (HR@K) and normalized discounted cumulative gain (NDCG@K) of indicator k are used to evaluate whether the level of interest corresponding to the products that users actually interact with ranks high. In this way, the effectiveness of the trained multiple candidate content recommendation models can be evaluated with the help of the corresponding values ​​of the final evaluation indicators HR@K and NDCG@K.

[0194] HR@K is a metric used in the field of information retrieval. HR@K is used to assess whether a ranking system includes relevant documents in its top k results. Specifically, HR@K calculates the proportion of relevant information contained in the top k documents. This metric can help evaluate the performance of ranking systems, especially when finding relevant information quickly.

[0195] NDCG@K is also a metric used in the information retrieval and ranking fields. NDCG@K considers not only whether a document is relevant, but also the degree of its relevance. NDCG@K weights the relevance score of each document in the ranking results and then normalizes the weights to allow comparison of the performance of different ranking systems. This metric is very useful for measuring the effectiveness of ranking systems.

[0196] In this embodiment, by standardizing the labels of training data, we can reduce the large variance and maximum values ​​in the collected labels of user game spending amounts, which can lead to instability in model training and thus affect effective model training. Furthermore, validating the content recommendation model based on user interactions with the recommended games can reduce the negative impact of uncertainty in user spending time and amounts on model effectiveness assessments, thereby improving the effectiveness of model evaluation.

[0197] After determining the target model for the online application, the target model needs to be retrained, and the target model is trained to output the user's spending amount for the recommended game. That is, the target model predicts the user's spending amount for the recommended game during the training process.

[0198] After completing the retraining of the target model, the retrained target model can be put online for application to implement game recommendations for users. For example, please refer to Figures 9 and 10. Figure 9 is a flow chart of a method for implementing game recommendations based on a retrained target model according to an embodiment of the present application; Figure 10 is a schematic diagram of an embodiment of the present application for inputting different data on different networks of the target model to implement game recommendations. As shown in Figures 9 and 10, the process of implementing game recommendations based on the retrained target model includes the following steps 901-905.

[0199] Step 901: Obtain interaction history and games to be recommended. The interaction history includes games recommended by the recommendation system and user feedback actions for the recommended games.

[0200] In this embodiment, the historical recommended content included in the interaction history is the games that the recommendation system has recommended, and no other types of content are included. The games to be recommended refer to games that need to be determined whether to be recommended to the user.

[0201] For example, assuming each user has only 10 historical interactive games, enter the historical interactive game list of user u Initialize a representation table of historical interactive games Here, m represents the number of historical interactive products in the dataset, and D represents the dimensionality of the representation of historically promoted products. So, for game d1, its corresponding representation is e1.

[0202] Step 902: Input the interaction history and the game to be recommended into a first feature extraction network to obtain a first feature.

[0203] Step 903: Input the games consumed by the user into a second feature extraction network to obtain a second feature.

[0204] In this embodiment, the games that the user has consumed are regarded as the content of interest to the user, so that the features of the games that the user has consumed are extracted through the second feature extraction network.

[0205] For example, in this embodiment, a multi-layer perceptron can be used to extract the interest preference representation v of user u based on the user's historical interactive games. u (That is, the second feature mentioned above). The specific function is as follows:

[0206] Step 904: Perform a dot product operation on the second feature and the feature of the game to be recommended to obtain a fused feature.

[0207] In this embodiment, the features of the game to be recommended are first obtained, for example, based on a pre-built game feature table, the features of the game to be recommended are obtained. p Then, based on the dot product operation, the fusion feature corresponding to the second feature and the feature of the game to be recommended is obtained: up =v u *e p .

[0208] In step 905, the first feature and the fused feature are concatenated and input into the prediction network to obtain the predicted consumption amount output by the prediction network.

[0209] For example, suppose the first feature is represented as Then the input of the prediction network can be expressed as That is, the input v of the prediction network concat The first feature And fusion features: v up Obtained after splicing.

[0210] It should be noted that before the target model is put into use, the process of training the target model is similar to the process of steps 901-905 above. The difference is that after step 905, a loss function needs to be constructed to train the target model. For example, suppose that during the training phase, the output of the prediction network is The label corresponding to the training data (that is, the actual consumption amount of the game in the training data) is s up ; Then, the following loss function can be constructed Thus, by minimizing the loss function The target model can be trained, that is, the parameters in the target model can be updated.

[0211] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0212] Please refer to Figure 11, which is a schematic diagram of the structure of a content recommendation device provided in an embodiment of the present application. As shown in Figure 11, the content recommendation device includes: an acquisition module 1101, which is used to obtain interaction history and content to be recommended, where the interaction history includes the historical recommended content of the recommendation system and the user's feedback actions on the historical recommended content; a processing module 1102, which is used to input the interaction history and the content to be recommended into a first feature extraction network to obtain a first feature; the processing module 1102 is also used to input the user's content of interest in the historical recommended content into a second feature extraction network to obtain a second feature; the processing module 1102 is also used to predict the user's feedback action on the content to be recommended based on the first feature and the second feature; and the processing module 1102 is also used to recommend content to the user based on the predicted feedback action.

[0213] In a possible implementation, the user's interested content is content in historical recommended content that corresponds to a positive feedback action, where the positive feedback action includes one or more of the following feedback actions: click, download, or consume.

[0214] In a possible implementation, the acquisition module 1101 is further configured to acquire features of the content to be recommended;

[0215] The processing module 1102 is further configured to predict a user's feedback action on the content to be recommended based on the first feature, the second feature, and the feature of the content to be recommended.

[0216] In a possible implementation, the processing module 1102 is further configured to:

[0217] Fusing the second feature with the feature of the content to be recommended to obtain a fused feature;

[0218] The first feature and the fused feature are input into the prediction network to predict the user's feedback action on the content to be recommended.

[0219] In one possible implementation, the first feature extraction network, the second feature extraction network, and the prediction network are uniformly trained based on the same training objective, and the prediction network is used to predict the feedback action based on the first feature and the second feature.

[0220] In one possible implementation, a target model including a first feature extraction network, a second feature extraction network, and a prediction network is determined from multiple content recommendation models, and the multiple content recommendation models are trained to output a user's interest level in input content.

[0221] In one possible implementation, when the prediction network is used to predict the amount of money a user will spend on content to be recommended, during the training process of multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user in the content already consumed.

[0222] In a possible implementation, the user's value measurement of the input content is obtained based on the user's consumption amount of the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

[0223] In a possible implementation, the content to be recommended is any one of the following: advertisement, application, product, video, song or news.

[0224] Please refer to Figure 12, which is a structural diagram of a model training device provided by an embodiment of the present application. As shown in Figure 12, the model training device includes: an acquisition module 1201, which is used to acquire training data, the training data including interaction history, recommended content and user's target feedback action for recommended content, the interaction history including the recommendation system's historical recommended content and the user's feedback action for the historical recommended content; a recommendation module 1202, which is used to input the interaction history and recommended content into a first feature extraction network to obtain a first feature; the recommendation module 1202 is also used to input the user's interested content in the historical recommended content into a second feature extraction network to obtain a second feature; the recommendation module 1202 is also used to predict the user's feedback action for the recommended content through a prediction network based on the first feature and the second feature to obtain a predicted feedback action; the recommendation module 1202 is also used to train a target model through a loss function constructed based on the predicted feedback action and the target feedback action to obtain an updated target model, the target model including a first feature extraction network, a second feature extraction network and a prediction network.

[0225] In a possible implementation, the user's interested content is content in historical recommended content that corresponds to a positive feedback action, where the positive feedback action includes one or more of the following feedback actions: click, download, or consume.

[0226] In a possible implementation, the acquisition module 1201 is further configured to acquire features of the recommended content; and the recommendation module 1202 is further configured to predict a user's feedback action on the recommended content based on the first feature, the second feature, and the features of the recommended content.

[0227] In a possible implementation, the recommendation module 1202 is further configured to: fuse the second feature with the feature of the recommended content to obtain a fused feature; and input the first feature and the fused feature into a prediction network to predict the user's feedback action on the recommended content.

[0228] In one possible implementation, the target model is determined from multiple content recommendation models, and the multiple content recommendation models are trained to output the user's interest level in input content.

[0229] In one possible implementation, when the prediction network is used to predict the amount of money a user will spend on content to be recommended, during the training process of multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user in the content already consumed.

[0230] In a possible implementation, the user's value measurement of the input content is obtained based on the user's consumption amount of the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

[0231] In a possible implementation, the recommended content is any one of the following: advertisement, application, product, video, song, or news.

[0232] Please refer to Figure 13, which is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1300 can be specifically manifested as a mobile phone, a tablet, a laptop computer, a smart wearable device, a server, etc., which is not limited here. Specifically, the execution device 1300 includes: a receiver 1301, a transmitter 1302, a processor 1303 and a memory 1304 (wherein the number of processors 1303 in the execution device 1300 can be one or more, and Figure 13 takes one processor as an example), wherein the processor 1303 may include an application processor 13031 and a communication processor 13032. In some embodiments of the present application, the receiver 1301, the transmitter 1302, the processor 1303 and the memory 1304 may be connected via a bus or other means.

[0233] Memory 1304 may include read-only memory and random access memory, and provides instructions and data to processor 1303. A portion of memory 1304 may also include non-volatile random access memory (NVRAM). Memory 1304 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0234] Processor 1303 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0235] The method disclosed in the above embodiment of the present application can be applied to the processor 1303, or implemented by the processor 1303. The processor 1303 can be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 1303 or an instruction in the form of software. The above-mentioned processor 1303 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0236] The processor 1303 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1304, and the processor 1303 reads the information in the memory 1304 and completes the steps of the above method in combination with its hardware.

[0237] Receiver 1301 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1302 can be used to output digital or character information through the first interface. Transmitter 1302 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1302 can also include a display device such as a display screen.

[0238] The electronic device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the method for determining the model structure described in the above embodiment, or so that the chip in the training device executes the method for determining the model structure described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0239] Specifically, see Figure 14 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1400. NPU 1400 is mounted on the host CPU (host CPU) as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1403, which is controlled by controller 1404 to extract matrix data from memory and perform multiplication operations.

[0240] In some implementations, the arithmetic circuit 1403 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1403 is a two-dimensional systolic array. The arithmetic circuit 1403 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1403 is a general-purpose matrix processor.

[0241] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1402 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1401 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1408.

[0242] Unified memory 1406 is used to store input and output data. Weight data is directly transferred to weight memory 1402 through the Direct Memory Access Controller (DMAC) 1405. Input data is also transferred to unified memory 1406 through the DMAC.

[0243] BIU stands for Bus Interface Unit, i.e., bus interface unit 1410 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1409 .

[0244] The bus interface unit 1410 (BIU) is used for the instruction fetch memory 1409 to obtain instructions from the external memory, and is also used for the storage unit access controller 1405 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0245] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1406 or transfer weight data to the weight memory 1402 or transfer input data to the input memory 1401.

[0246] The vector calculation unit 1407 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1403, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0247] In some implementations, the vector calculation unit 1407 can store the processed output vector to the unified memory 1406. For example, the vector calculation unit 1407 can apply a linear function or a nonlinear function to the output of the operation circuit 1403, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1407 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1403, for example, for use in subsequent layers in a neural network.

[0248] An instruction fetch buffer 1409 connected to the controller 1404 is used to store instructions used by the controller 1404;

[0249] Unified memory 1406, input memory 1401, weight memory 1402, and instruction fetch memory 1409 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0250] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0251] Please refer to Figure 15, which is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the method disclosed in Figure 4 above can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.

[0252] 15 schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0253] In one embodiment, computer readable storage medium 1500 is provided using signal bearing medium 1501. Signal bearing medium 1501 may include one or more program instructions 1502 that, when executed by one or more processors, may provide the functionality or portions of the functionality described above with respect to FIG.

[0254] In some examples, signal bearing medium 1501 may include computer readable medium 1503 such as, but not limited to, a hard drive, compact disk (CD), digital video disk (DVD), digital tape, memory, ROM or RAM, and the like.

[0255] In some embodiments, the signal-bearing medium 1501 may include a computer-recordable medium 1504, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1501 may include a communication medium 1505, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1501 may be communicated via a wireless form of the communication medium 1505 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocol).

[0256] The one or more program instructions 1502 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1502 communicated to the computing device via one or more of computer-readable media 1503, computer-recordable media 1504, and / or communication media 1505.

[0257] It should also be noted that the device embodiments described above are merely illustrative, in which the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0258] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.

[0259] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0260] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training equipment or data center to another website, computer, training equipment or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training equipment, data center, etc. that includes one or more available media integrations. Available media can be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)), etc.

Claims

1. A content recommendation method, characterized in that: include: Acquire interaction history and content to be recommended, wherein the interaction history includes historical recommended content of the recommendation system and user feedback actions on the historical recommended content; Inputting the interaction history and the content to be recommended into a first feature extraction network to obtain a first feature; Inputting the user's interested content in the historical recommended content into a second feature extraction network to obtain a second feature; predicting, based on the first feature and the second feature, a feedback action of the user with respect to the content to be recommended; Recommending content to the user based on the predicted feedback action.

2. The method according to claim 1, characterized in that The user-interested content is content in the historical recommended content that corresponds to a positive feedback action, where the positive feedback action includes at least one of clicking, downloading, or consuming.

3. The method according to claim 1 or 2, characterized in that The predicting, based on the first feature and the second feature, a feedback action of the user with respect to the content to be recommended includes: Acquiring features of the content to be recommended; Based on the first feature, the second feature, and the feature of the content to be recommended, predict a feedback action of the user with respect to the content to be recommended.

4. The method according to claim 3, characterized in that The predicting, based on the first feature, the second feature, and the feature of the content to be recommended, a feedback action of the user with respect to the content to be recommended, includes: fusing the second feature with the feature of the content to be recommended to obtain a fused feature; The first feature and the fused feature are input into a prediction network to predict the user's feedback action on the content to be recommended.

5. The method according to claim 4, characterized in that The first feature extraction network, the second feature extraction network, and the prediction network are uniformly trained based on the same training objective, and the prediction network is used to predict feedback actions based on the first feature and the second feature.

6. The method according to claim 4 or 5, characterized in that A target model including the first feature extraction network, the second feature extraction network, and the prediction network is determined from a plurality of content recommendation models, and the plurality of content recommendation models are trained to output a user's interest level in input content.

7. The method according to claim 6, characterized in that When the prediction network is used to predict the amount of money the user will spend on the content to be recommended, during the training process of the multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user.

8. The method according to claim 7, characterized in that The user's value measurement value for the input content is obtained based on the user's consumption amount for the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

9. The method according to any one of claims 1 to 8, characterized in that The content to be recommended is any one of the following: advertisement, application, product, video, song or news.

10. A model training method, characterized in that: include: Acquire training data, where the training data includes interaction history, recommended content, and user targeted feedback actions for the recommended content. The interaction history includes historical recommended content by the recommendation system and user feedback actions for the historical recommended content. Inputting the interaction history and the recommended content into a first feature extraction network to obtain a first feature; Inputting the user's interested content in the historical recommended content into a second feature extraction network to obtain a second feature; Based on the first feature and the second feature, predicting, by a prediction network, a feedback action of the user with respect to the content to be recommended, to obtain a predicted feedback action; The target model is trained by a loss function constructed based on the prediction feedback action and the target feedback action to obtain an updated target model, wherein the target model includes the first feature extraction network, the second feature extraction network and the prediction network.

11. A content recommendation device, characterized in that: include: An acquisition module is used to acquire interaction history and content to be recommended, wherein the interaction history includes historical recommended content of the recommendation system and user feedback actions on the historical recommended content; a processing module, configured to input the interaction history and the content to be recommended into a first feature extraction network to obtain a first feature; The processing module is further configured to input the user's interested content in the historical recommended content into a second feature extraction network to obtain a second feature; The processing module is further configured to predict a feedback action of the user with respect to the content to be recommended based on the first feature and the second feature; The processing module is further configured to recommend content to the user based on the predicted feedback action.

12. The device according to claim 11, characterized in that The user-interested content is content in the historical recommended content that corresponds to a positive feedback action, where the positive feedback action includes at least one of clicking, downloading, or consuming.

13. The device according to claim 11 or 12, characterized in that The acquisition module is further configured to acquire features of the content to be recommended; The processing module is further configured to predict the user's feedback action on the content to be recommended based on the first feature, the second feature, and features of the content to be recommended.

14. The device according to claim 13, characterized in that The processing module is further configured to: fusing the second feature with the feature of the content to be recommended to obtain a fused feature; The first feature and the fused feature are input into a prediction network to predict the user's feedback action on the content to be recommended.

15. The device according to claim 14, characterized in that The first feature extraction network, the second feature extraction network, and the prediction network are uniformly trained based on the same training objective, and the prediction network is used to predict feedback actions based on the first feature and the second feature.

16. The device according to claim 14 or 15, characterized in that A target model including the first feature extraction network, the second feature extraction network, and the prediction network is determined from a plurality of content recommendation models, and the plurality of content recommendation models are trained to output a user's interest level in input content.

17. The device according to claim 16, characterized in that When the prediction network is used to predict the amount of money the user will spend on the content to be recommended, during the training process of the multiple content recommendation models, the true value of the user's interest in the input content is obtained based on the user's value measurement of the input content and the proportion of the input content consumed by the user.

18. The device according to claim 17, characterized in that The user's value measurement value for the input content is obtained based on the user's consumption amount for the input content, the mean consumption amount of the input content, and the standard deviation of the consumption amount of the input content.

19. The device according to any one of claims 11 to 18, characterized in that The content to be recommended is any one of the following: advertisement, application, product, video, song or news.

20. A model training device, characterized in that: include: an acquisition module, configured to acquire training data, wherein the training data includes interaction history, recommended content, and user target feedback actions for the recommended content, wherein the interaction history includes historical recommended content of the recommendation system and user feedback actions for the historical recommended content; a processing module, configured to input the interaction history and the recommended content into a first feature extraction network to obtain a first feature; The processing module is further configured to input the user's interested content in the historical recommended content into a second feature extraction network to obtain a second feature; The processing module is further configured to predict, through a prediction network, a feedback action of the user with respect to the recommended content based on the first feature and the second feature, to obtain a predicted feedback action; The processing module is further used to train the target model by using a loss function constructed based on the prediction feedback action and the target feedback action to obtain an updated target model, wherein the target model includes the first feature extraction network, the second feature extraction network and the prediction network.

21. A content recommendation device, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 9.

22. A model training device, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device performs the method according to claim 10.

23. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 10.

24. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Voice packet recommendation method and device, equipment and storage medium

    CN113746875A

  • Commodity recommendation method, prediction model training method and related equipment

    CN115564517A

  • Neural network training method, apparatus and device, and computer readable storage medium

    CN116385075A

  • Content recommendation method and device, equipment, storage medium and computer program product

    CN116521971A

  • Recommendation system using improved neural network

    US10896459B1